Stable Diffusion - Image to Prompts (Ongoing)
A kaggle competition
This an ongoing kaggle competition
The goal of this competition is to reverse the typical direction of a generative text-to-image model: to predict the prompts from the images generated by stable diffusion 2.0.
I am working on trying different models like CLIP and handling with embeddings to solve this question.