
MeL-ReHS2LSTM: A Multimodal Semantic Retrieval Framework with Fuzzy Ranking and Lightweight Caption Parsing
Multimodal retrieval systems are essential for deriving semantic insights from visual and textual data. Achieving precise alignment between image captions and text queries remains a non -trivial challenge. Conventional methods such as Long Short -Term Memory ( LSTM )-based or transformer -based encoders often struggle to balance efficiency, semantic accuracy, and interpretability. These approaches typically lack robust mechanisms for handling the fuzzy nature of human language. This stu dy proposes a lightweight and int erpretable framework that combines Bootstrapped Language Image Pretraining ( BLIP )-generated captions with embeddings using Meta -Learning Rectified Hyperbolic Secant Group Norm -Lasso -based Long Short -Term Memory networks (MeL -ReHS 2LSTM) and uses Neutrosophi c Cosine Similarity -based Fuzzy Inference System (NCS -FIS) to improve semantic retrieval. The architecture uses Parentheses Parsing Abstract Syntax Tree (PP -AST) for the symbolic structuring of captions and Dilated Skip -Connection Graph Neural Network (DSC -GNN) for the enhanced context embeddings to provide the ability to perform both neural and rule -based analysis in parallel. A custom dataset of 20 real -world video clips was created for evaluation, including a variety of scenes such as interviews, team me etings, and sports events. Empirical results show that the accuracy of top -1 is 92.0%, top-5 recall is 96.7%, Bilingual Evaluation Understudy ( BLEU ) score is 0.7521, and the average matching time is 0.58ms, which overcomes the traditional retrieval baselines. The proposed methodology provides a good trade -off between accuracy, runtime performance and interpretability. Future work will explore the extension of the model to zero -shot domains and the combination of adaptive caption refinement strategies .
[1] Kumar R., “Visual Linguistic Model and its Applications in Image Captioning,” SN Computer Science , vol. 1, pp. 1 -12, 2024. https://osf.io/preprints/osf/zy9m4_v1
[2] Lee B., Bang J., Song H., and Kang B., “Alzheimer’s Disease Recognition Using Graph Neural Network by Leveraging Image -Text Similarity from Vision Language Model,” Scientific Reports , vol. 15, pp. 1 -14, 2025. https://doi.org/10.1080/1448837X.2024.2448376
[3] Li Q., Zhao F., Zhao L., Liu M., and et al., “A System of Multimodal Image -Text Retrieval Based on Pre -Trained Models Fusion,” Concurrency and Computation: Practice and Experience , vol. 37, no. 3, pp. 1 -11, 2024. DOI: 10.1002/cpe.8345
[4] Li X., Wang J., and Yan Z., “Can Graph Neural Networks be Adequately Explained? A Survey,” ACM Computing Surveys , vol. 57, no. 5 , pp. 1 -36, 2025. https://doi.org/10.1145/3711122
[5] Liu Z., Liu J., and Ma F., “Improving Cross -Modal Alignment with Synthetic Pairs for Text -Only Image Captioning,” in Proceedings of the 38 th AAAI Conference on Artificial Intelligence , Vancouver, pp. 3864 -3872, 2024. https://ojs.aaai.org/index.php/AAAI/article/view/ 28178
[6] Long Z., Killick G., McCreadie R., and Camarasa G., “Multiway -Adapter: Adapting Multimodal Large Language Models for Scalable Image -Text Retrieval,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing , Seoul, pp. 6580 -6584, 2024. https://ieeexplore.ieee.org/document/10446792
[7] Long Z., Liang K., Aragon -Camarasa G., Henderson P., and Mccreadie R., “Zero -Shot Interactive Text -to-Image Retrieval via Diffusion - Augmented Representations,” arXiv Preprint , vol. arXiv:2501.15379, pp. 1 -11, 2025. https://arxiv.org/abs/2501.15379v1
[8] Luo B., Luo B., Wang J., Zewen W., and et al., “Graph -Based Cross -Domain Knowledge Distillation for Cross -Dataset Text -to-Image Person Retrieval,” arXiv Preprint , vol. arXiv:2501.15052, pp. 568 -576, 2025. https://arxiv.org/abs/2501.15052
[9] Lyu S., Tian Z., Ou Z., Zhu Y., and et al., “TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image -Text Retrieval,” arXiv Preprint , vol. arXiv:2501.10935, pp. 19269 - 19277, 2025. https://arxiv.org/abs/2501.10935
[10] Osman A., Shalaby M., Soliman M., and Elsayed K., “Ar -CM -ViMETA: Arabic Image Captioning Based on Concept Model and Vision -Based Multi - Encoder Transformer Architecture,” The International Arab Journal of Information Technology , vol. 21, no. 3, pp. 458 -465, 2024. https://doi.org/10.34028/iajit/21/3/9
[11] Patel V., Modi A., Mistry H., Mishra A.., and et al., “From Alt -Text to Real Context: Revolutionizing Image Captioning Using the Potential of LLM,” International Journal of Scientific Research in Computer Science, Engineering and Information Technology , vol. 11, no. 1, pp. 379 -387, 2025. DOI: 10.32628/CSEIT25111238
[12] Patil D., “Multimodal Artificial Intelligence in Industry: Integrating Text, Image, and Audio for Enhanced Applications Across Sectors,” SSRN , pp. 1 -9, 2025. DOI: 10.2139/ssrn.5057428
[13] Qian X. and Liu B., “Enhanced Image -Text Retrieval Based on CLIP with YOLOv10 and Next -ViT,” in Proceedings of the 4 th International Conference on Computer Vision, Application, and Algorithm , Chengdu, pp. 134862K -134862K -6, 2024. https://doi.org/10.1117/12.3055876
[14] Qin X., Li L., Pang G., and Hao F., “Heterogeneous Graph Fusion Network for Cross - Modal Image -Text Retrieval,” Expert Systems with Applications , vol. 249, no. 1, pp. 123842, 2024. DOI: 10.1016/j.eswa.2024.123842
[15] Tan C., Zhang W., Qi Z., Shih K., and et al., “Generating Multimodal Images with GAN: Integrating Text, Image, and Style,” arXiv Preprint , vol. arXiv:2501.02167v1, pp. 1 -8, 2025. DOI: 10.48550/arxiv.2501.02167
[16] Wang J., Zhang H., Zhong Y., Liang Y., and et al., “Advanced Multimodal Deep Learning Architecture for Image -Text Matching,” in Proceedings of the IEEE 4 th International Conference on Electronic Technology, Communication and Information , Changchun, pp. 1185 -1191, 2024. https://ieeexplore.ieee.org/document/10594167
[17] Wei S., Tan A., and Wang Y., “A Multiscale Hybrid Model for Natural and Remote Sensing Image - Text Retrieval,” in Proceedings of the 6 th International Conference on Machine Learning, Big Data and Business Intelligence , Hangzhou, pp. 41 -45, 2024. DOI: 10.1109/MLBDBI63974.2024.10823800
[18] Yang X. and Tan L., “Multimodal Knowledge Graph Construction for Intelligent Question Answering Systems,” Australian Journal of Electrical and Electronics Engineering , vol. 22, no. 4, pp. 755 -764, 2025. DOI: 10.1080/1448837x.2024.2448376
[19] Yuan Y. and Xue H., “Multimodal Information Integration and Retrieval Framework Based on Graph Neural Networks,” Preprints , vol. 202501.0405.v1, pp. 1 -8, 2025. DOI: 10.20944/preprints202501.0405.v1
[20] Zhang Q., Zhou Y., Yang Z., and Chen A., “Cross - Modal Prominent Fragments Enhancement Aligning Network for Image -Text Retrieval,” in Proceedings of the IEEE International Conference on Multimedia and Expo , Ontario, pp. 1 -6, 2024. https://ieeexplore.ieee.org/document/10687706
[21] Zheng W., Han N., Hu Y., Kang P., and et al., “Cross -Modal Image Text Retrieval Based on Cross -Optimization and Joint Strategy,” in Proceedings of the International Conference on Image, Signal Processing, and Pattern Recognition , Guangzhou, pp.131802I -131802I - 11, 2024. https://doi.org/10.1117/12.3034096