Stable Diffusion - Image to Prompts (Ongoing)

A kaggle competition

This an ongoing kaggle competition

The goal of this competition is to reverse the typical direction of a generative text-to-image model: to predict the prompts from the images generated by stable diffusion 2.0.

I am working on trying different models like CLIP and handling with embeddings to solve this question.