i-manager Publications

Image Caption Extraction to Aid Visual Learning

Shailesh Sangle*, Palak Kabra**, Mihir Gharat***, Dhiraj Jha****

*-**** Department of Computer Engineering, Thakur College of Engineering and Technology, Mumbai, India.

Periodicity:April - June'2023
DOI : https://doi.org/10.26634/jip.10.2.19404

Abstract

An image caption generator is essential for social media enthusiasts or visually impaired individuals. It can be used as a plugin in popular social media platforms to recommend suitable captions or to assist visually impaired people in comprehending the image content on the web, thereby eliminating ambiguity in image meaning and ensuring accurate knowledge acquisition. This research describes an image caption generator that utilizes a Convolutional Neural Network (CNN) and a Long Short-Term Memory (LSTM) model to generate natural language descriptions of images. The CNN was employed to extract features from the input image, which were then fed into the LSTM to generate the corresponding caption. The model is trained on a large dataset of image-caption pairs, using a combination of supervised and reinforcement learning techniques. The model's performance is evaluated using several metrics, and the results demonstrate that the proposed CNN LSTM model outperforms existing state-of-the-art approaches in generating accurate and diverse image captions. This model has the potential to be used in various applications, including image retrieval, content-based image search, and assisting visually impaired individuals to understanding their surroundings. It also discusses about the structure and functions of the various neural networks involved.

Keywords

CNN, LSTM, BLEU, Encoder, ResNet50.

How to Cite this Article?

Sangle, S., Kabra, P., Gharat, M., and Jha, D. (2023). Image Caption Extraction to Aid Visual Learning. i-manager’s Journal on Image Processing, 10(2), 14-25. https://doi.org/10.26634/jip.10.2.19404

References

[1]. EmptyAnderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2018). Bottom-up and topdown attention for image captioning and visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 6077-6086).

[2]. Chen, J., Dong, W., & Li, M. (2014). Image Caption Generator based on Deep Neural Networks.

[3]. Hochreiter, S. (1998). The vanishing gradient problem during learning recurrent neural nets and problem solutions. International Journal of Uncertainty, Fuzziness and Knowledge- Based Systems, 6(2), 107-116.

[4]. EmptyKrishnakumar, B., Kousalya, K., Gokul, S., Karthikeyan, R., & Kaviyarasu, D. (2020). Image caption generator using deep learning. International Journal of Advanced Science and Technology, 29(3), 975-980.

[5]. Mathur, P., Gill, A., Yadav, A., Mishra, A., & Bansode, N. K. (2017, June). Camera2Caption: A real-time image caption generator. In 2017 International Conference on Computational Intelligence in Data Science (ICCIDS) (pp. 1-6). IEEE.

[6]. EmptyMiyazaki, T., & Shimizu, N. (2016, August). Crosslingual image caption generation. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 1, 1780-1790.

[7]. Mounika, S., & Vijaybabu, P. (2022). Image caption generator using CNN and LSTM. South Asian Journal of Engineering and Technology, 12(3), 78-86.

[8]. EmptyNaidu, P. A., Vats, S., Anand, G., & Nalina, V. (2020). A deep learning model for image caption generation. International Journal of Computer Sciences and Engineering, 8(6), 10-17.

[9]. Panicker, M. J., Upadhayay, V., Sethi, G., & Mathur, V. (2021). Image caption generator. International Journal of Innovative Technology and Exploring Engineering (IJITEE), 10(3), 87-92.

[10]. EmptyPapineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002, July). Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311-318).

[11]. EmptyPawar, R., Jadhav, O., & Nalage, R. (2022). Image caption generator using CNN and LSTM (GUI application). International Journal of Advance Research and Innovative Ideas in Education, 8(3).

[12]. Sairam, G., Mandha, M., Prashanth, P., & Swetha, P. (2021, November). Image Captioning using CNN and LSTM. In 4th Smart Cities Symposium (SCS 2021), 274-277. IET.

[13]. Tang, Z., Yi, Y., & Sheng, H. (2021). Attention-Guided Image Captioning through Word Information. Sensors, 21(23), 7982.

[14]. Tanti, M., Gatt, A., & Camilleri, K. P. (2018). Where to put the image in an image caption generator. Natural Language Engineering, 24(3), 467-489.

[15]. EmptyTripathi, S., & Sharma, R. (2019). Image Caption Generator Using CNN and LSTM. International Journal of Creative Research Thoughts (IJCRT), 1-6.

[16]. EmptyVinyals, O., Toshev, A., Bengio, S., & Erhan, D. (2015). Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 3156-3164).

[17]. Waghmare, P. M., & Shinde, S. V. (2022). Image caption generation using neural network models and lstm hierarchical structure. In Computational Intelligence in Pattern Recognition: Proceedings of CIPR 2021 (pp. 109- 117). Springer Singapore.

[18]. Wang, H., Zhang, Y., & Yu, X. (2020). An overview of image caption generation methods. Computational Intelligence and Neuroscience.

[19]. EmptyYou, Q., Jin, H., Wang, Z., Fang, C., & Luo, J. (2016). Image captioning with semantic attention. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4651-4659).

	North Americas,UK, Middle East,Europe		India	Rest of world
	USD	EUR	INR	USD-ROW
Pdf	35	35	200	20
Online	15	15	200	15
Pdf & Online	35	35	400	25

Image Caption Extraction to Aid Visual Learning

Abstract

Keywords

How to Cite this Article?

References

If you have access to this article please login to view the article or kindly login to purchase the article

Purchase Instant Access

Options for accessing this content: