SIFT Meets CNN: A Decade Survey of Instance Retrieval

Liang Zheng, Yi Yang*, Qi Tian

*Corresponding author for this work

Research output: Contribution to journalReview articlepeer-review

551 Citations (Scopus)

Abstract

In the early days, content-based image retrieval (CBIR) was studied with global features. Since 2003, image retrieval based on local descriptors (de facto SIFT) has been extensively studied for over a decade due to the advantage of SIFT in dealing with image transformations. Recently, image representations based on the convolutional neural network (CNN) have attracted increasing interest in the community and demonstrated impressive performance. Given this time of rapid evolution, this article provides a comprehensive survey of instance retrieval over the last decade. Two broad categories, SIFT-based and CNN-based methods, are presented. For the former, according to the codebook size, we organize the literature into using large/medium-sized/small codebooks. For the latter, we discuss three lines of methods, i.e., using pre-trained or fine-tuned CNN models, and hybrid methods. The first two perform a single-pass of an image to the network, while the last category employs a patch-based feature extraction scheme. This survey presents milestones in modern instance retrieval, reviews a broad selection of previous works in different categories, and provides insights on the connection between SIFT and CNN-based methods. After analyzing and comparing retrieval performance of different categories on several datasets, we discuss promising directions towards generic and specialized instance retrieval.

Original languageEnglish
Pages (from-to)1224-1244
Number of pages21
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
Volume40
Issue number5
DOIs
Publication statusPublished - 1 May 2018
Externally publishedYes

Fingerprint

Dive into the research topics of 'SIFT Meets CNN: A Decade Survey of Instance Retrieval'. Together they form a unique fingerprint.

Cite this