Image Caption Extraction to Aid Visual Learning

Shailesh Sangle*, Palak Kabra**, Mihir Gharat***, Dhiraj Jha****
*-**** Department of Computer Engineering, Thakur College of Engineering and Technology, Mumbai, India.
Periodicity:April - June'2023
DOI : https://doi.org/10.26634/jip.10.2.19404

Abstract

An image caption generator is essential for social media enthusiasts or visually impaired individuals. It can be used as a plugin in popular social media platforms to recommend suitable captions or to assist visually impaired people in comprehending the image content on the web, thereby eliminating ambiguity in image meaning and ensuring accurate knowledge acquisition. This research describes an image caption generator that utilizes a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) model to generate natural language descriptions of images. The CNN was employed to extract features from the input image, which were then fed into the LSTM to generate the corresponding caption. The model is trained on a large dataset of image-caption pairs, using a combination of supervised and reinforcement learning techniques. The model's performance is evaluated using several metrics, and the results demonstrate that the proposed CNN LSTM model outperforms existing state-of-the-art approaches in generating accurate and diverse image captions. This model has the potential to be used in various applications, including image retrieval, content-based image search, and assisting visually impaired individuals to understanding their surroundings. It also discusses about the structure and functions of the various neural networks involved.

Keywords

CNN, LSTM, BLEU, Encoder, ResNet50.

How to Cite this Article?

Sangle, S., Kabra, P., Gharat, M., and Jha, D. (2023). Image Caption Extraction to Aid Visual Learning. i-manager’s Journal on Image Processing, 10(2), 14-25. https://doi.org/10.26634/jip.10.2.19404

References

[1]. EmptyAnderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2018). Bottom-up and topdown attention for image captioning and visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 6077-6086).
[4]. EmptyKrishnakumar, B., Kousalya, K., Gokul, S., Karthikeyan, R., & Kaviyarasu, D. (2020). Image caption generator using deep learning. International Journal of Advanced Science and Technology, 29(3), 975-980.
[6]. EmptyMiyazaki, T., & Shimizu, N. (2016, August). Crosslingual image caption generation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1, 1780-1790.
[8]. EmptyNaidu, P. A., Vats, S., Anand, G., & Nalina, V. (2020). A deep learning model for image caption generation. International Journal of Computer Sciences and Engineering, 8(6), 10-17.
[10]. EmptyPapineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002, July). Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311-318).
[11]. EmptyPawar, R., Jadhav, O., & Nalage, R. (2022). Image caption generator using CNN and LSTM (GUI application). International Journal of Advance Research and Innovative Ideas in Education, 8(3).
[15]. EmptyTripathi, S., & Sharma, R. (2019). Image Caption Generator Using CNN and LSTM. International Journal of Creative Research Thoughts (IJCRT), 1-6.
[16]. EmptyVinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3156-3164).
[19]. EmptyYou, Q., Jin, H., Wang, Z., Fang, C., & Luo, J. (2016). Image captioning with semantic attention. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4651-4659).
If you have access to this article please login to view the article or kindly login to purchase the article

Purchase Instant Access

Single Article

North Americas,UK,
Middle East,Europe
India Rest of world
USD EUR INR USD-ROW
Online 15 15

Options for accessing this content:
  • If you would like institutional access to this content, please recommend the title to your librarian.
    Library Recommendation Form
  • If you already have i-manager's user account: Login above and proceed to purchase the article.
  • New Users: Please register, then proceed to purchase the article.