Skip navigation
Please use this identifier to cite or link to this item: http://repository.iitr.ac.in/handle/123456789/21710
Title: Caption-Based Region Extraction in Images
Authors: Agrawal P.
Yadav R.
Yadav V.
De K.
Pratim Roy, Partha
Chaudhuri B.B.
Nakagawa M.
Khanna P.
Kumar S.
Published in: Advances in Intelligent Systems and Computing
3rd International Conference on Computer Vision and Image Processing, CVIP 2018
Abstract: Image captioning and object detection are some of the most growing and popular research areas in the field of computer vision. Almost every upcoming technology uses vision in some way, and with various people researching in the field of object detection, many vision problems which seemed intractable seem close to solved now. But there has been less research in identifying regions associating actions with objects. Dense Image Captioning [8] is one such application, which localizes all the important regions in an image along with their description. Something very similar to normal image captioning, but repeated for every salient region in the image. In this paper, we address the aforementioned problem of detecting regions explaining the query caption. We use edge boxes for efficient object proposals, which we further filter down using a score measure. The object proposals are then captioned using a pretrained Inception [19] model. The captions of each of these regions are checked for similarity with the query caption using the skip-thought vectors [9]. This proposed framework produces interesting and efficient results. We provide a quantitative measure of our experiment by taking the intersection over union (IoU) with the ground truth on the visual genome [10] dataset. By combining the above techniques in an orderly manner, we have been able to achieve encouraging results. © 2020, Springer Nature Singapore Pte Ltd.
Citation: Advances in Intelligent Systems and Computing (2020), 1024: 27-38
URI: https://doi.org/10.1007/978-981-32-9291-8_3
http://repository.iitr.ac.in/handle/123456789/21710
Issue Date: 2020
Publisher: Springer Science and Business Media Deutschland GmbH
Keywords: Image captioning
Inception networks
Long short-term memory
Region proposal network
Skip thought vectors
ISBN: 9.78981E+12
ISSN: 21945357
Author Scopus IDs: 57190276702
57211823201
57212997299
57204533373
56880478500
Author Affiliations: Agrawal, P., Indian Institute of Technology Roorkee, Roorkee, India
Yadav, R., Indian Institute of Technology Roorkee, Roorkee, India
Yadav, V., Indian Institute of Technology Roorkee, Roorkee, India
De, K., Indian Institute of Technology Roorkee, Roorkee, India
Pratim Roy, P., Indian Institute of Technology Roorkee, Roorkee, India
Corresponding Author: De, K.; Indian Institute of Technology RoorkeeIndia; email: kanjar.cspdf2017@iitr.ac.in
Appears in Collections:Conference Publications [CS]

Files in This Item:
There are no files associated with this item.
Show full item record


Items in Repository are protected by copyright, with all rights reserved, unless otherwise indicated.