(50-4) 13 * << * >> * Russian * English * Content * All Issues

Improve Image Text Descriptions using Large Language Models
N. Andriyanov1, A. Kim1

1Financial University, 125167, Russia, Moscow, Leningradsky pr-t, b. 49

  Full text (PDF)

DOI: 10.18287/COJ1860

Article ID: 1860

Language: English

Abstract:
In this paper, a multi-stage approach to improve text queries (prompts) for image generation is proposed and comprehensively investigated. First, the GPT-2 model, pre-trained on 18,000 raw query-quality query pairs from the Lexica.art platform, automatically expands and refines the original prompts to produce more detailed and semantically accurate images when generated by the diffusion network Stable Diffusion 1.5. Next, the Image Captioning task (BLIP2) and four large language models (DeepSeek, Grok, ChatGPT, YandexGPT) are used to compare the quality of signature expansion, demonstrating different stylistic strategies for augmenting initial descriptions. The proposed "Captioning → Prompt Enhancer (Mistral) → Stable Diffusion+LoRA" Pipeline additionally includes a pre-training of Mistral's own model on BLEU, METEOR and CIDEr metrics, providing a steady increase in the quality of textual descriptions (BLEU from 0.12 to 0.435 after 500 epochs) and a significant reduction in the FID metric for image generation (from 0.3482 to 0.1873). In the final stage, the LoRA modules embedded in UNet and the Stable Diffusion text encoder allow efficient learning of the generation of previously "unknown" objects (rare or fictional), reducing FID to 0.172 at rank = 64. An expert survey (134 respondents) confirmed the visual preference of images generated by optimized queries, demonstrating the potential of the proposed technique to improve the quality of multimodal systems.

Keywords:
image generation, large language models, prompt, text enhancement, multimodal models, LoRA.

Citation:
Andriyanov N, Kim A. Improve Image Text Descriptions using Large Language Models. Computer Optics 2026; 50(4): 1860. doi: 10.18287/COJ1860.

References:

  1. Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners. OpenAI technical report; 2019. Available at: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf (accessed 30 March 2026).
  2. Li J, Li D, Savarese S, Hoi S. BLIP-2: bootstrapping language-image pre-training with frozen pre-trained models. arXiv 2023; arXiv:2301.12597.
  3. Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image synthesis with latent diffusion models. In: Proc IEEE/CVF Conf Comput Vis Pattern Recognit (CVPR); 2022. pp. 10684-10695.
  4. Sahoo P, Mondal S, Sarkar D. Systematic survey of prompt engineering in large language models: techniques and applications. arXiv 2024; arXiv:2402.07927.
  5. Liu C, Zhou M, Zhang Y, Yang K. Analyzing the impact of prompt tokens on text-to-image generation. In: Proc ACM Int Conf Multimedia (MM); 2023. pp. 1234-1243.
  6. Chang M, Rong L, Xu Y, Zhang Q. Efficient prompting methods for large language models: a survey. arXiv 2024; arXiv:2404.01077.
  7. Li J, Galley M, Brockett C, Spithourakis G, Dolan B. Improving language understanding by generative pre-training. OpenAI Blog; 2018. Available at: https://openai.com/research/language-unsupervised (accessed 30 March 2026).
  8. Liu C, Singh K, Saha S, et al. Prompt engineering on vision-language models: a survey. arXiv 2023; arXiv:2307.12980.
  9. Ilieva G. Extension of interval-valued hesitant Fermatean fuzzy TOPSIS for evaluating and benchmarking of generative AI chatbots. Electronics 2025; 14(3): 555. doi:10.3390/electronics14030555.
  10. Kaswan KS, Dhatterwal JS, Malik K, Baliyan A. Generative AI: a review on models and applications. In: Proc 2023 Int Conf Commun Secur Artif Intell (ICCSAI); 2023. pp. 699-704. doi:10.1109/ICCSAI59793.2023.10421601.
  11. Hagos DH, Battle R, Rawat DB. Recent advances in generative AI and large language models: current status, challenges, and perspective. IEEE Trans Artif Intell 2024; 5(12): 5873-5893. doi:10.1109/TAI.2024.3444742.
  12. Andriyanov N. Multimodal data processing based on text classifiers and image recognition. In: Rousseau JJ, Kapralos B, eds. Pattern Recognition, Computer Vision, Image Processing (ICPR 2022). Lecture Notes in Computer Science, vol. 13644. Cham: Springer; 2023. pp. 414-423. doi:10.1007/978-3-031-37742-6_31.
  13. Roumeliotis KI, Tselikas ND, Nasiopoulos DK. Leveraging large language models in tourism: comparative study of latest GPT Omni models and BERT NLP for customer review classification and sentiment analysis. Information 2024; 15(12): 792. doi:10.3390/info15120792.
  14. Kanimozhiselvi CS. Image captioning using deep learning. In: Proc 2022 Int Conf Comput Commun Informatics (ICCCI); 2022. pp. 1-7. doi:10.1109/ICCCI54379.2022.9740788.
  15. Andriyanov N, Kim A, Fao X. Using generative models to improve fire detection efficiency. In: Proc 2024 X Int Conf Inf Technol Nanotechnol (ITNT); 2024. pp. 1-4. doi:10.1109/ITNT60778.2024.10582386.
  16. Andriyanov NA. Combining text and image analysis methods for solving multimodal classification problems. Pattern Recognit Image Anal 2022; 32(3): 489-494. doi:10.1134/S1054661822030026.
  17. LexicaArt. Lexica image database. Available at: https://lexica.art/ (accessed 31 March 2026).
  18. Andriyanov N, Dementiev V. Image caption extending using LLM and style transfer. In: Palaiahnakote S, Schuckers S, Ogier JM, Bhattacharya P, Pal U, Bhattacharya S, eds. Pattern Recognition (ICPR 2024). Lecture Notes in Computer Science, vol. 15616. Cham: Springer; 2025. pp. 108-118. doi:10.1007/978-3-031-87663-9_9.
  19. Andriyanov NA, Burankina PV, Dementyiev VE. Using artificial neural networks for reinforced concrete structures condition monitoring. In: Proc 2024 Int Russian Smart Industry Conf (SmartIndustryCon); 2024. pp. 450-454. doi:10.1109/SmartIndustryCon61328.2024.10515891.
  20. Veselov D, Andriyanov N, Trung L. Detection of intracranial hemorrhage by artificial intelligence and deep learning methods. In: Proc 2024 X Int Conf Inf Technol Nanotechnol (ITNT); 2024. pp. 1-6. doi:10.1109/ITNT60778.2024.10582393.
  21. Hessel J, Holtzman A, Forbes M, Le Bras R, Choi Y. CLIPScore: a reference-free evaluation metric for image captioning. arXiv 2021; arXiv:2104.08718.
  22. Schuhmann C. LAION-Aesthetics: high-quality image subset and aesthetic predictor. LAION blog; 16 Aug 2022. Available at: https://laion.ai/blog/laion-aesthetics/ (accessed 31 March 2026).
  23. Hao Y, Chi Z, Dong L, Wei F. Optimizing prompts for text-to-image generation. arXiv 2023; arXiv:2212.09611.
  24. Hertz A, Mokady R, Tenenbaum J, Aberman K, Pritch Y, Cohen-Or D. Prompt-to-Prompt image editing with cross attention control. arXiv 2022; arXiv:2208.01626.
  25. Mudgal P. REFLEX: reference-free evaluation of log summarization via large language model judgment. arXiv 2025; arXiv:2511.07458. Available at: https://arxiv.org/pdf/2511.07458.pdf (accessed 30 March 2026).
  26. Peng Y, Cui Y, Tang H, et al. DreamBench++: human-aligned benchmark for personalized image generation. arXiv 2024; arXiv:2406.16855. Available at: https://arxiv.org/pdf/2406.16855v2.pdf (accessed 30 March 2026).
  27. Zhang J, Huang WX, Lu MX, Li LW, Wang X, Shen YP, Wang YF. Efficient U-shaped transformer network for low-light power image denoising. Computer Optics 2025; 49(5): 775-784. doi:10.18287/2412-6179-CO-1629.
  28. Li MX, Xu CJ. Convolutional neural network-based low light image enhancement method. Computer Optics 2025; 49(2): 334-343. doi:10.18287/2412-6179-CO-1499.
  29. Nikonorov AV, Petrov MV, Bibikov SA, Kutikova VV, Morozov AA, Kazanskiy NL. Image restoration in diffractive optical systems using deep learning and deconvolution. Computer Optics 2017; 41(6): 875-887. doi:10.18287/2412-6179-2017-41-6-875-887.
  30. Khonina SN, Kazanskiy NL, Oseledets IV, Nikonorov AV, Butt MA. Synergy between artificial intelligence and hyperspectral imaging: a review. Technologies 2024; 12(9): 163. doi:10.3390/technologies12090163.
  31. Pavlov VA, Belov AA, Nguen VT, Jovanovski N, Ovsyannikova AS. Comparison of neural networks for suppression of multiplicative noise in images. Computer Optics 2024; 48(3): 425-431. doi:10.18287/2412-6179-CO-1400.
  32. Kalai AT, Nachum O, Vempala SS, Zhang E. Why language models hallucinate. arXiv 2025; arXiv:2509.04664.
  33. Pagan N, Baumann J, Elokda E, De Pasquale G, Bolognani S, Hannák A. Classification of feedback loops and their relation to biases in automated decision-making systems. arXiv 2023; arXiv:2305.06055. Available at: https://arxiv.org/pdf/2305.06055.pdf (accessed 31 March 2026).

151, Molodogvardeiskaya str., Samara, 443001, Russia; E-mail: journal@computeroptics.ru; Tel: +7 (846) 242-41-24 (Executive secretary), +7 (846) 332-56-22 (Issuing editor), Fax: +7 (846) 332-56-20