(large multimodal model).
What are the developments:
π€ LLaVa: A competitor to the open-source GPT4-V
π Langchain for identifying images: RAG on images
π MiniGPT-v2: Visual-language hybrid tasks
π¨ SEED-LLaMA: Simulates human seeing, reading, and imagining
