(50-4) 13 *
<<
*
>>
* Русский *
English
*
Содержание *
Все выпуски
Improve Image Text Descriptions using Large Language Models
N. Andriyanov1, A. Kim1
1Financial University, 125167, Russia, Moscow,
Leningradsky pr-t, b. 49
Полный текст (PDF)
DOI: 10.18287/COJ1860
ID статьи: 1860
Аннотация:
In this paper, a multi-stage approach to improve text queries (prompts)
for image generation is proposed and comprehensively investigated.
First, the GPT-2 model, pre-trained on 18,000 raw query-quality query
pairs from the Lexica.art platform, automatically expands and refines
the original prompts to produce more detailed and semantically accurate
images when generated by the diffusion network Stable Diffusion 1.5.
Next, the Image Captioning task (BLIP2) and four large language models
(DeepSeek, Grok, ChatGPT, YandexGPT) are used to compare the quality of
signature expansion, demonstrating different stylistic strategies for
augmenting initial descriptions. The proposed "Captioning → Prompt
Enhancer (Mistral) → Stable Diffusion+LoRA" Pipeline additionally
includes a pre-training of Mistral's own model on BLEU, METEOR and CIDEr
metrics, providing a steady increase in the quality of textual
descriptions (BLEU from 0.12 to 0.435 after 500 epochs) and a significant
reduction in the FID metric for image generation (from 0.3482 to 0.1873).
In the final stage, the LoRA modules embedded in UNet and the Stable
Diffusion text encoder allow efficient learning of the generation of
previously "unknown" objects (rare or fictional), reducing FID to 0.172
at rank = 64. An expert survey (134 respondents) confirmed the visual
preference of images generated by optimized queries, demonstrating the
potential of the proposed technique to improve the quality of multimodal
systems.
Ключевые слова:
image generation, large language models, prompt, text enhancement,
multimodal models, LoRA.
Citation:
Andriyanov N, Kim A. Improve Image Text Descriptions using Large Language
Models. Computer Optics 2026; 50(4): 1860. doi: 10.18287/COJ1860.
References:
-
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models
are unsupervised multitask learners. OpenAI technical report; 2019.
Available at:
https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
(accessed 30 March 2026).
-
Li J, Li D, Savarese S, Hoi S. BLIP-2: bootstrapping language-image
pre-training with frozen pre-trained models. arXiv 2023;
arXiv:2301.12597.
-
Rombach R, Blattmann A, Lorenz D, Esser P, Ommer B. High-resolution image
synthesis with latent diffusion models. In: Proc IEEE/CVF Conf Comput Vis
Pattern Recognit (CVPR); 2022. pp. 10684-10695.
-
Sahoo P, Mondal S, Sarkar D. Systematic survey of prompt engineering in
large language models: techniques and applications. arXiv 2024;
arXiv:2402.07927.
-
Liu C, Zhou M, Zhang Y, Yang K. Analyzing the impact of prompt tokens on
text-to-image generation. In: Proc ACM Int Conf Multimedia (MM); 2023.
pp. 1234-1243.
-
Chang M, Rong L, Xu Y, Zhang Q. Efficient prompting methods for large
language models: a survey. arXiv 2024; arXiv:2404.01077.
-
Li J, Galley M, Brockett C, Spithourakis G, Dolan B. Improving language
understanding by generative pre-training. OpenAI Blog; 2018.
Available at: https://openai.com/research/language-unsupervised
(accessed 30 March 2026).
-
Liu C, Singh K, Saha S, et al. Prompt engineering on vision-language
models: a survey. arXiv 2023; arXiv:2307.12980.
-
Ilieva G. Extension of interval-valued hesitant Fermatean fuzzy TOPSIS
for evaluating and benchmarking of generative AI chatbots. Electronics
2025; 14(3): 555. doi:10.3390/electronics14030555.
-
Kaswan KS, Dhatterwal JS, Malik K, Baliyan A. Generative AI: a review on
models and applications. In: Proc 2023 Int Conf Commun Secur Artif Intell
(ICCSAI); 2023. pp. 699-704. doi:10.1109/ICCSAI59793.2023.10421601.
-
Hagos DH, Battle R, Rawat DB. Recent advances in generative AI and large
language models: current status, challenges, and perspective. IEEE Trans
Artif Intell 2024; 5(12): 5873-5893. doi:10.1109/TAI.2024.3444742.
-
Andriyanov N. Multimodal data processing based on text classifiers and
image recognition. In: Rousseau JJ, Kapralos B, eds. Pattern Recognition,
Computer Vision, Image Processing (ICPR 2022). Lecture Notes in Computer
Science, vol. 13644. Cham: Springer; 2023. pp. 414-423.
doi:10.1007/978-3-031-37742-6_31.
-
Roumeliotis KI, Tselikas ND, Nasiopoulos DK. Leveraging large language
models in tourism: comparative study of latest GPT Omni models and BERT
NLP for customer review classification and sentiment analysis.
Information 2024; 15(12): 792. doi:10.3390/info15120792.
-
Kanimozhiselvi CS. Image captioning using deep learning. In: Proc 2022
Int Conf Comput Commun Informatics (ICCCI); 2022. pp. 1-7.
doi:10.1109/ICCCI54379.2022.9740788.
-
Andriyanov N, Kim A, Fao X. Using generative models to improve fire
detection efficiency. In: Proc 2024 X Int Conf Inf Technol Nanotechnol
(ITNT); 2024. pp. 1-4. doi:10.1109/ITNT60778.2024.10582386.
-
Andriyanov NA. Combining text and image analysis methods for solving
multimodal classification problems. Pattern Recognit Image Anal 2022;
32(3): 489-494. doi:10.1134/S1054661822030026.
-
LexicaArt. Lexica image database. Available at: https://lexica.art/
(accessed 31 March 2026).
-
Andriyanov N, Dementiev V. Image caption extending using LLM and style
transfer. In: Palaiahnakote S, Schuckers S, Ogier JM, Bhattacharya P,
Pal U, Bhattacharya S, eds. Pattern Recognition (ICPR 2024). Lecture
Notes in Computer Science, vol. 15616. Cham: Springer; 2025. pp. 108-118.
doi:10.1007/978-3-031-87663-9_9.
-
Andriyanov NA, Burankina PV, Dementyiev VE. Using artificial neural
networks for reinforced concrete structures condition monitoring. In:
Proc 2024 Int Russian Smart Industry Conf (SmartIndustryCon); 2024.
pp. 450-454. doi:10.1109/SmartIndustryCon61328.2024.10515891.
-
Veselov D, Andriyanov N, Trung L. Detection of intracranial hemorrhage
by artificial intelligence and deep learning methods. In: Proc 2024 X
Int Conf Inf Technol Nanotechnol (ITNT); 2024. pp. 1-6.
doi:10.1109/ITNT60778.2024.10582393.
-
Hessel J, Holtzman A, Forbes M, Le Bras R, Choi Y. CLIPScore: a
reference-free evaluation metric for image captioning. arXiv 2021;
arXiv:2104.08718.
-
Schuhmann C. LAION-Aesthetics: high-quality image subset and aesthetic
predictor. LAION blog; 16 Aug 2022. Available at:
https://laion.ai/blog/laion-aesthetics/ (accessed 31 March 2026).
-
Hao Y, Chi Z, Dong L, Wei F. Optimizing prompts for text-to-image
generation. arXiv 2023; arXiv:2212.09611.
-
Hertz A, Mokady R, Tenenbaum J, Aberman K, Pritch Y, Cohen-Or D.
Prompt-to-Prompt image editing with cross attention control. arXiv 2022;
arXiv:2208.01626.
-
Mudgal P. REFLEX: reference-free evaluation of log summarization via
large language model judgment. arXiv 2025; arXiv:2511.07458.
Available at: https://arxiv.org/pdf/2511.07458.pdf
(accessed 30 March 2026).
-
Peng Y, Cui Y, Tang H, et al. DreamBench++: human-aligned benchmark for
personalized image generation. arXiv 2024; arXiv:2406.16855.
Available at: https://arxiv.org/pdf/2406.16855v2.pdf
(accessed 30 March 2026).
-
Zhang J, Huang WX, Lu MX, Li LW, Wang X, Shen YP, Wang YF. Efficient
U-shaped transformer network for low-light power image denoising.
Computer Optics 2025; 49(5): 775-784.
doi:10.18287/2412-6179-CO-1629.
-
Li MX, Xu CJ. Convolutional neural network-based low light image
enhancement method. Computer Optics 2025; 49(2): 334-343.
doi:10.18287/2412-6179-CO-1499.
-
Nikonorov AV, Petrov MV, Bibikov SA, Kutikova VV, Morozov AA,
Kazanskiy NL. Image restoration in diffractive optical systems using
deep learning and deconvolution. Computer Optics 2017; 41(6): 875-887.
doi:10.18287/2412-6179-2017-41-6-875-887.
-
Khonina SN, Kazanskiy NL, Oseledets IV, Nikonorov AV, Butt MA. Synergy
between artificial intelligence and hyperspectral imaging: a review.
Technologies 2024; 12(9): 163.
doi:10.3390/technologies12090163.
-
Pavlov VA, Belov AA, Nguen VT, Jovanovski N, Ovsyannikova AS.
Comparison of neural networks for suppression of multiplicative noise in
images. Computer Optics 2024; 48(3): 425-431.
doi:10.18287/2412-6179-CO-1400.
-
Kalai AT, Nachum O, Vempala SS, Zhang E. Why language models
hallucinate. arXiv 2025; arXiv:2509.04664.
-
Pagan N, Baumann J, Elokda E, De Pasquale G, Bolognani S, Hannák A.
Classification of feedback loops and their relation to biases in
automated decision-making systems. arXiv 2023; arXiv:2305.06055.
Available at: https://arxiv.org/pdf/2305.06055.pdf
(accessed 31 March 2026).
Россия, 443001, Самара, ул. Молодогвардейская, 151; электронная почта:
journal@computeroptics.ru;
тел: +7 (846) 242-41-24 (ответственный секретарь),
+7 (846) 332-56-22 (технический редактор),
факс: +7 (846) 332-56-20