Search preprocessing program, device and method using target point cloud, and image search program

By generating and classifying point cloud data based on image feature points and camera information, image grouping and reclassifying are solved, and the problems of high image search calculation cost and unreliable clustering results in the prior art are solved, and efficient and accurate image search and ranking are achieved.

JP7674302B2Active Publication Date: 2025-05-09KDDI CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022082909
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-05-09
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

When processing a large number of similar or reverse symmetric images, the calculation cost is high and it is difficult to effectively search for images. Traditional preprocessing methods mainly rely on the visual information of the images, resulting in the clustering results being unreliable.

Method used

By generating point cloud data based on image feature points and camera information, point cloud classification is performed to generate target-related point cloud groups, then image grouping is performed, image group candidates are determined, and image reclassification is performed through feature quantities and representative values, and finally search and ranking of target images are performed.

Benefits of technology

Efficient preprocessing of image search is realized, and similar images can be more accurately identified and grouped, reducing computing costs, and improving the reliability and efficiency of image search.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007674302000001
    Figure 0007674302000001
  • Figure 0007674302000002
    Figure 0007674302000002
  • Figure 0007674302000003
    Figure 0007674302000003
Patent Text Reader

Abstract

To provide a search preprocessing program independent of only physical appearance of an image with respect to an image group from which images close to a target image including one of a plurality of targets are searched and capable of performing search preprocessing.SOLUTION: In an image search device, a program causes a computer to function as: a point group classification unit which classifies a point group generated on the basis of the feature points of an image included in an image group and relating to a target included in the image group into a plurality of point groups each of which is assumed to relate to one target on the basis of the closeness between points; a classification preprocessing unit which determines an image group candidate, which is a set of images corresponding to the point group, about each point group; and an image classification unit which reclassifies each image included in the image group into a plurality of image groups each of which is assumed to relate to one target on the basis of a feature amount of the image and a representative value of a feature amount related to the image group candidate. The point group is preferably generated on the basis of a feature point of each image and camera information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an image retrieval technique for retrieving images from an image database. [Background technology]

[0002] In recent years, there has been an increasing demand for image retrieval technology that retrieves images similar to a query image from an image database. In the early stages of this technology, a method was used to determine the similarity of images using image features related to texture and color that were generated manually. However, such a method required a large number of images, had a huge computational cost, and lacked generalization capabilities.

[0003] As a technique for dealing with such a problem, for example, Patent Document 1 discloses a technique for generating features for each group in an image database using a test signature in order to deal with the increase in computational cost. Also, for example, Patent Document 2 discloses a technique for extracting statistically salient features in order to improve generalization ability.

[0004] Furthermore, for example, Non-Patent Document 1 proposes a deep learning-based image feature. This feature is not a local feature that reflects local features of an image, but a global feature that reflects information on the entire image.

[0005] By adopting such global features, it is possible to adjust the generalization ability of features and the calculation time. In addition, as disclosed in Non-Patent Documents 2 and 4, it is also possible to perform re-ranking on image search results. Furthermore, as disclosed in Non-Patent Documents 5 and 6, it is also possible to improve the robustness of image search by re-ranking image search results using collaborative representation.

[0006] Meanwhile, unlike these methods, a technology has been proposed that applies a clustering method to the original image database to perform preprocessing (cleaning) and prepare for more appropriate image retrieval. For example, in Non-Patent Document 3, a clustering method is used to label images in an image database in order to obtain training data for a deep learning model used in image retrieval. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] U.S. Pat. No. 5,819,288 [Patent Document 2] US Patent Application Publication No. 2002 / 0178149 [Non-patent literature]

[0008] [Non-Patent Document 1] Jerome Revaud, Jon Almazan, Rafael S. Rezende, and Cesar de Souza, “Learning with Average Precision: Training Image Retrieval with a Listwise Loss”, Proceedings of 2019 IEEE / CVF International Conference on Computer Vision (ICCV), pp 5106-5115, 2019 [Non-Patent Document 2] Bingyi Cao, Andre Araujo, and Jack Sim, “Unifying Deep Local and Global Features for Image Search”, Proceedings of European Conference on Computer Vision (ECCV), pp 726-743, 2020 [Non-Patent Document 3] Ke Mei, Lei li, Jinchang Xu, Yanhua Cheng, and Yugeng Lin, “3rd Place Solution to “Google Landmark Retrieval 2020””, Instance-Level Recognition workshop in European Conference on Computer Vision (ECCV) 2020, <https: / / doi.org / 10.48550 / arXiv.2008.10480>, 2020 [Non-Patent Document 4] Noura Bouhlel, Ghada Feki, Anis Ben Ammar, and Chokri Ben Amar, “Hypergraph learning with collaborative representation for image search reranking”, International Journal of Multimedia Information Retrieval vol. 9, pp 205-214, 2020

Non-Patent Document 5

Non-Patent Document 6

Summary of the Invention

Problems to be Solved by the Invention

[0009] In general, image databases contain a large number of images that are similar in appearance because they are taken from different positions, and also images that are different in appearance from the same position, for example, images that are antisymmetrical (left-right mirrored). Such images are particularly common in camera images taken in close-range areas such as a city, and they pose a major problem when performing image retrieval processing.

[0010] For image search results from an image database that contains many images that are problematic in such image search, the techniques disclosed in, for example, Non-Patent Documents 2 and 4 described above perform a reranking process. However, these techniques perform the reranking process on all pairs of the query image and images in the image database, and therefore the calculation time required to perform this process becomes enormous.

[0011] In response to this, for example, the technology disclosed in Non-Patent Document 3 performs clustering processing as preprocessing on images in an image database that contain problematic images for image retrieval, as described above, and thereby prepares learning data for a deep learning model to be used for image retrieval.

[0012] However, most of the conventional preprocessing techniques for such image databases perform clustering based only on visual information of the appearance of the images. As a result, it is still difficult to reliably obtain clustering results that are effective for image retrieval. Furthermore, it is necessary to perform labeling processing on each generated cluster.

[0013] Therefore, the present invention aims to provide a search preprocessing program, device, and method capable of performing preprocessing for a search that does not depend on the mere appearance of an image on a group of images from which an image similar to a query image (target image) is searched, and also aims to provide an image search program capable of performing such preprocessing to perform a more suitable image search. [Means for solving the problem]

[0014] According to the present invention, there is provided a search pre-processing program for performing pre-processing for a search on an image group from which an image similar to a target image including one of a plurality of targets is searched, the program comprising: a point cloud classification means for classifying a point cloud related to the target included in the image set, the point cloud being generated based on feature points of each image included in the image set, into a plurality of point cloud groups, each of which is regarded as related to one of the targets, based on the proximity between points of the point cloud; A classification pre-processing means for determining, for each of the point cloud groups, an image group candidate that is a set of the images corresponding to the point cloud group; an image classification means for reclassifying each image included in the image group into a plurality of image groups each of which is regarded as relating to one of the targets based on a feature value of the image and a representative value of the feature values ​​related to the candidate image group; A search preprocessing program is provided that causes a computer to function as follows:

[0015] As one embodiment of the search pre-processing program according to the present invention, it is also preferable that the point cloud relating to the target included in the image group is generated based on feature points of each of the images and camera information relating to each image determined from the feature points.

[0016] In another embodiment of the search pre-processing program according to the present invention, it is also preferable that the classification pre-processing means determines a central image corresponding to the target for each of the determined image group candidates based on the number of feature points that are the basis of the point cloud related to the target, sets other images as peripheral images, extracts peripheral images that differ from the central image by a predetermined amount or more in terms of image scale and / or orientation, removes them from the image group candidate, and determines a new image group candidate that includes the extracted peripheral images.

[0017] Furthermore, in the above embodiment for determining image group candidates, it is also preferable that the pre-classification processing means extracts peripheral images that satisfy at least one of the following conditions: the ratio of their own scale to the scale of the central image is outside a specified range; and the difference between their own orientation and the orientation of the central image is greater than or equal to a specified value, and determines new image group candidates.

[0018] According to the present invention, there is also provided an image retrieval program for retrieving an image similar to a target image including one of a plurality of targets from a group of images, the program comprising: a point cloud classification means for classifying a point cloud related to the target included in the image set, the point cloud being generated based on feature points of each image included in the image set, into a plurality of point cloud groups, each of which is regarded as related to one of the targets, based on the proximity between points of the point cloud; A classification pre-processing means for determining, for each of the point cloud groups, an image group candidate that is a set of the images corresponding to the point cloud group; an image classification means for reclassifying each image included in the image group into a plurality of image groups each of which is regarded as relating to one of the targets based on a feature value of the image and a representative value of the feature values ​​related to the candidate image group; an image group search means for searching for an image group that is similar to the target image from among the plurality of image groups, based on the feature amounts of each image belonging to the image group and the feature amount of the target image; An image search program is provided that causes a computer to function as follows.

[0019] As one embodiment of the image retrieval program according to the present invention, it is also preferable that the image group retrieval means determines a coefficient vector that reduces the difference between the amount obtained by multiplying a feature matrix representing the overall features of the image group by a coefficient vector and the feature of the target image, and searches for an image group that is close to the target image based on the determined coefficient vector, the feature of each image belonging to the image group, and the feature of the target image.

[0020] In addition, as another embodiment of the image retrieval program according to the present invention, it is preferable that the image retrieval program further causes the computer to function as an image ranking means for determining an image from among the images included in the image group that is closest to the target image based on the features of each image included in the searched image group and the features of the target image, or for assigning ranking information to the images included in the image group in order of proximity to the target image.

[0021] According to the present invention, there is further provided a search pre-processing device for performing pre-processing for a search on an image group from which an image similar to a target image including one of a plurality of targets is searched, the pre-processing comprising: a point cloud classification means for classifying a point cloud related to the target included in the image set, the point cloud being generated based on feature points of each image included in the image set, into a plurality of point cloud groups, each of which is regarded as related to one of the targets, based on the proximity between points of the point cloud; A classification pre-processing means for determining, for each of the point cloud groups, an image group candidate that is a set of the images corresponding to the point cloud group; an image classification means for reclassifying each image included in the image group into a plurality of image groups each of which is regarded as relating to one of the targets based on a feature value of the image and a representative value of the feature values ​​related to the candidate image group; A search preprocessing device is provided having the following:

[0022] According to the present invention, there is still further provided a search pre-processing method for performing pre-processing for a search on an image group from which an image similar to a target image including one of a plurality of targets is searched, the method comprising the steps of: classifying point clouds associated with the targets in the set of images, generated based on feature points of each image in the set of images, into a number of point cloud groups, each of which is deemed to be associated with one of the targets, based on the proximity between points in the point clouds; determining, for each of the point cloud groups, a candidate image group that is a set of the images corresponding to the point cloud group; reclassifying each image included in the image group into a plurality of image groups each of which is considered to be related to one of the targets based on the feature value of the image and a representative value of the feature values ​​related to the candidate image group; A computer-implemented search pre-processing method is provided, comprising: Effect of the Invention

[0023] According to the search preprocessing program, device, and method of the present invention, a preprocessing for a search that does not depend on the mere appearance of an image can be performed on a group of images from which an image similar to a query image (target image) is searched. Also, according to the image search program of the present invention, it is possible to perform a more suitable image search by performing such preprocessing. [Brief description of the drawings]

[0024] [Figure 1] FIG. 2 is a functional block diagram showing a functional configuration of an embodiment of a search preprocessing device according to the present invention. [Diagram 2] FIG. 2 is a schematic diagram showing a specific example of an image stored and managed in an image database according to the present invention. [Diagram 3] FIG. 2 is a schematic diagram showing a specific example of the classification pre-processing unit according to the present invention selecting images related to landmarks (landmarks). [Figure 4] 5A and 5B are schematic diagrams showing specific examples of image classification processing results by the image classification unit according to the present invention. [Diagram 5] 10 is a schematic diagram for explaining a specific example of image group search processing in an image group search section according to the present invention; FIG. [Figure 6] 10A and 10B are schematic diagrams for explaining a specific example of image ranking processing in the image ranking section according to the present invention; DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.

[0026] [Pre-search processing device, image search device] FIG. 1 is a functional block diagram showing the functional configuration of an embodiment of a search preprocessing device according to the present invention.

[0027] The image search device 1 of this embodiment shown in FIG. (a) Image search processing for searching for an image close to (similar to) a query image that includes a target image containing one of a number of "targets," such as a landmark (geographical target) in this embodiment, from a group of images stored and managed in an image database (DB) 2. In addition, as an embodiment of the search pre-processing device according to the present invention, (b) Pre-processing the above images to more appropriately perform the search (pre-search processing) 1, the image database 2 is installed outside the image search device 1, but it may be provided within the image search device 1 as a component thereof.

[0028] Here, the pre-search processing of (b) above can be performed on the image group in advance, for example, offline, as will be described later in detail, so that the more appropriate image search processing of (a) above can be performed, for example, online, as needed, after obtaining the query image.

[0029] In order to realize the above-described pre-search processing, the image search device 1 specifically performs the following steps: (A) a point cloud classification unit 111 that classifies a point cloud related to "targets" (e.g., landmarks) included in the image group, which is generated based on feature points of each image included in the image group, into a plurality of "point cloud groups" each of which is considered to be related to one of the "targets" based on the proximity between points of the point cloud (e.g., distance in three-dimensional space); (B) a classification pre-processing unit 112 that determines, for each of the “point cloud groups”, “image group candidates” that are a set of images corresponding to the point cloud group; (C) an image classification unit 113 that reclassifies each image included in the image group into a plurality of "image groups" each of which is considered to relate to one of the "targets" based on the feature value of the image and a representative value of the feature values ​​related to the "image group candidates"; It has.

[0030] In order to explain the technical significance of these functional components, a specific example of images stored and managed in the image database 2 will be described below with reference to FIG. 2. In general, the image database 2 may include images that may not be searched as close (similar) images due to differences in viewpoint positions, such as the two images shown in FIG. 2(A) even though they contain the same "target". Also, as in the two images shown in FIG. 2(B), even though they contain the same "target", one image contains almost the entire "target", while the other image contains a part of it (the upper part in FIG. 2(B)) outside the image, and therefore may not be searched as close (similar) images. Furthermore, as in the two images shown in FIG. 2(C), although they contain the same "target", they also contain a considerable amount of differences within the angle of view, and as a result, may not be searched as close (similar) images.

[0031] Moreover, image database 2 may also contain images that are antisymmetrical (left-right reversed) to each other. Conventionally, such images have often not been searched for as close (similar) images, even if they contain the same "target." In addition, it is quite possible that images that contain different "targets" but happen to look similar may be included and thus be searched for as close (similar) images. Furthermore, it is not uncommon that images that contain the same "target" but are not searched for as close (similar) images due to differences in illuminance (for example, the difference between day and night) may be included.

[0032] For a group of images that may contain images that are problematic in such image retrieval, the image retrieval device 1 (Fig. 1) obtains a point cloud related to a "target" as shown in (A) above, and generates from the point cloud a number of "point cloud groups," each of which is considered to relate to one of the "targets." Here, one "point cloud group" includes (points from) a point cloud generated from a number of camera images including, for example, a building A as a landmark, and is classified not simply based on appearance, but taking into consideration the characteristics of, for example, its three-dimensional shape (the distance between points in three-dimensional space).

[0033] Therefore, in this case, this one "point cloud group" can be regarded with a high degree of certainty as being "a group related to building A." As a result, the "image group" generated in (C) above and corresponding to this point cloud group can be understood as a collection of images including building A. As described above, according to the image search device (pre-search processing device) 1, it is possible to perform pre-search processing on the above image group, such as image classification by target, which does not depend solely on the appearance of the images.

[0034] The functional configuration of the image search device (search pre-processing device) 1 of this embodiment will be described in more detail below.

[0035] [Device functional configuration, search preprocessing program / method, image search program / method] According to the functional block diagram of Fig. 1, an image search device (search pre-processing device) 1 according to one embodiment of the present invention has an input / output interface (IF) unit 101 and a processor / memory (a processing system with a memory function). This processor / memory stores an embodiment of an image search program including a search pre-processing program according to the present invention, has a computer function, and executes the image search program to perform image search processing (search pre-processing).

[0036] For this reason, the image search device (search pre-processing device) 1 may be a device dedicated to image search processing (search pre-processing), but it can also be a cloud server, a non-cloud server, a personal computer (PC), a notebook or tablet computer, a mobile terminal such as a smartphone, or even a wearable terminal such as an HMD (Head Mounted Display), all equipped with the image search program (search pre-processing program) according to the present invention.

[0037] In addition, the processor memory (A) A classification pre-processing unit 112 including a point cloud classification unit 111, an image group candidate determination unit 112a, a center / peripheral image determination unit 112b, a scale / orientation judgment unit 112c, and a new image group candidate determination unit 112d, and an image classification unit 113. (a) An image group search unit 121 and an image ranking unit 122 In other words, these functional components can be considered as functions realized by executing an image search program (search preprocessing program) stored in a processor memory. In addition, the process flow shown by connecting the functional components of the image search device (search preprocessing device) 1 in Fig. 1 with arrows can be understood as one embodiment of the image search method (search preprocessing method) according to the present invention.

[0038] Incidentally, a device equipped with a search preprocessing program that embodies the functional configuration unit of (A) above can be regarded as a search preprocessing device according to the present invention even if it does not include the functional configuration of (B) above. In this case, this search preprocessing device, together with a device equipped with a program that embodies the functional configuration unit of (B) above, constitutes an image search system.

[0039] (Point cloud generation and classification processing) Similarly, in the functional block diagram of FIG. 1, the point cloud classification unit 111 in this embodiment is as follows: (a) based on feature points of each image included in an image group retrieved from the image database 2 via the input / output interface unit 101, generate three-dimensional point cloud data relating to targets included in the image group, i.e., a plurality of landmarks (geographical targets) in this embodiment; (b) The generated point cloud data (each point data) is classified into multiple point cloud groups (groups of point data), each of which is considered to relate to one of the landmarks.

[0040] First, regarding (a) above, the point cloud classification unit 111 specifically calculates feature points in each image included in the extracted image group using a known method, and generates a point cloud related to landmarks included in the original image group from the calculated feature points of each image using SfM (Structure from Motion). Here, SfM is a technology that is widely used to obtain 3D point cloud data from multiple aerial camera images taken by a drone, for example, and is described in Non-Patent Document: Johannes L. Schonberger, Jan-Michael Frahm, “Structure-from-Motion Revisited”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR),<https: / / doi.org / 10.1109 / CVPR.2016.445> , 2016, for a detailed explanation.

[0041] Furthermore, various methods have been proposed for this SfM, but in this embodiment, an SfM can be adopted in which the position and orientation of the camera for each image is determined from the correspondence between feature points between the images, the three-dimensional positions of these feature points are determined based on triangulation from the determined camera position and orientation, and this process is repeated each time an image is added, to generate point cloud data corresponding to the final determined three-dimensional positions.

[0042] That is, in this case, the point cloud classification unit 111 generates point cloud data related to landmarks included in the original image group based on the feature points of each image included in the image group and the camera information related to each image determined from the feature points. Incidentally, such point cloud data may be obtained from an external database that stores the results of similar processing, such as the image database 2.

[0043] Next, for the above (b), the point cloud classification unit 111 classifies the point cloud data P(={p|p:=(x, y, z)∈R 3}) is divided into K_pred (≧2) point cloud groups (P1, P2, . . . , P), each of which is considered to be related to one of the landmarks, based on the proximity of the points (3D points) p contained in the point cloud data P, for example by using a K-nearest neighbors classification algorithm or a k-dimensional tree space partitioning algorithm. K_pred , where P = P1 ∪ P2 ∪ · · · ∪ P K_pred ) are classified as follows.

[0044] In this embodiment, the K_pred point cloud groups correspond to the K_pred landmarks present in the area related to the image group. For example, the K_pred landmarks present in the area may be specified in advance together with their positions (obtained by a positioning means such as a GPS or from map data), and the K_pred point cloud groups including the positions of the landmarks (for example, near the center) may be selected from the point cloud groups generated as a result of classifying the point cloud data P. Of course, if only K_pred point cloud groups are generated as a result of classification (by appropriately specifying the K_pred landmarks or by appropriately setting the original image group), such selection is not necessary.

[0045] (Pre-classification processing) Also in the functional block diagram of FIG. 1, the classification pre-processing unit 112 of this embodiment determines, for each point cloud group, an image group candidate that is a set of images corresponding to the point cloud group, and further updates the determined candidates as described below. First, the classification pre-processing unit 112: (A) (As the image group candidate determining unit 112a) divides the original image group I into K_pred point cloud groups (P1, P2, . . . , P K_pred ) corresponding to the K_pred image group candidates (I1, I2, . . . , I K_pred , where I = I1∪I2∪···∪I K_pred ) Note that each image group candidate (I1, I2, . . . , I K_pred ) may be linked to the index of the feature amount x (=φ(f)) of image f generated by equation (3) described later in order to identify images included in the image group candidate.

[0046] Here, one method of dividing the images into candidate image groups will be described with reference to FIG. 3, which shows a specific example of selecting images related to a landmark based on a point cloud group related to the landmark. In the specific example shown in FIG. 3, first, a point cloud group P i For this point cloud group P i The position and orientation information of the multiple (many) cameras involved in the generation of the image (information on the camera positions and the directions of the front faces of the cameras) is identified as shown in FIG.

[0047] In other words, it is possible to identify camera position and orientation information that is determined to be related to an image that includes this landmark based on the content (position and orientation) of the information. Next, the images related to the identified multiple (large number) camera position and orientation information are grouped into one image group candidate I, as shown in FIG. 3(C). i This is what we say.

[0048] Returning to the functional block diagram of FIG. 1, the classification pre-processing unit 112 (a) (As the center / peripheral image determining unit 112b) i In the corresponding landmark point group P i Based on the number of feature points on which the image is based, the “center image” corresponding to this landmark is determined, and the other images are designated as “peripheral images.” (c) (As the scale / orientation determining unit 112c,) each image group candidate I i In the image group candidate I, "peripheral images" that are different from the "center image" in terms of image scale and / or orientation by a predetermined amount are extracted, and the extracted "peripheral images" are classified into the image group candidate I. i Remove from (iv) (as the new image group candidate determining unit 112d) determines new image group candidates including the “peripheral images” extracted in (iii) above, Thereafter, the same processes as those in (a) to (d) above are repeated until the number (K) of image group candidates becomes constant.

[0049] Specifically, the center / peripheral image determining unit 112b in (a) determines each image group candidate I i (i=1, 2, , K_pred), the corresponding point cloud group P i The image that contains the most number of feature points on which the i In this case, the central image c i is the point cloud group P i The image that contains the most points (3D points) that make up the landmarks in (has the most correspondences) is the image in the candidate image group I. i Center image c in i Images other than the surrounding image n i (∈I i \c i )

[0050] Next, the scale / orientation determination unit 112c in (c) performs a scale / orientation determination on the image group candidate I using, for example, SIFT (Scale-Invariant Feature Transform). iThen, the scale and orientation of the image are calculated from the keypoints of the image contained in the center image c. i The condition that the ratio of the scale of the center image c to the scale of the center image c is outside the specified range, and the self-orientation and the center image c i The difference between the orientation of the surrounding image n and the orientation of the surrounding image n is a predetermined value or more. In this embodiment, both of the conditions are satisfied. i The above is extracted.

[0051] Specifically, the central image c i The scale and orientation of ci and ci Let n be the surrounding image. i The scale and orientation of ni and ni For example, the following two conditional expressions (1) 1 / θ < s ci / s ni < θ (2) |o ci -o ni | < β Surrounding images n that satisfy both i The surrounding image n i As it is, image group candidate I i On the other hand, the neighboring image n i Other surrounding images n i Let I be the candidate image group. i It extracts and excludes from.

[0052] Next, the new image group candidate determining unit 112d in (d) above selects the neighboring images n i If this is the first time, the new image group candidate is generated as the (K_pred+1)th group, i.e., I K_pred+1 It becomes.

[0053] After that, the center / peripheral image determining unit 112b in (a) generates the new image group candidate I K_pred+1 For image group candidate I, the scale and orientation of K_pred+1 The image closest to the average value in the center image c K_pred+1 The remaining images are the surrounding images. K_pred+1 Next, the image group candidate I K_pred+1 Regarding (c), the scale / orientation determining unit 112c described above in (c) and the new image group candidate determining unit 112d described above in (d) carry out the same processing as described above.

[0054] This also results in the (K_pred+2)th new image group candidate I K_pred+2 If generated, this image group candidate I K_pred+2 For each of the image group candidates, the same processes as those in (a) to (d) above are repeated, and this cycle is carried out until the generation of new image group candidates stops. In this embodiment, of the image group candidates finally generated and determined, image group candidates that contain less than a predetermined number NL of images are deleted. Furthermore, for image group candidates that contain more than the predetermined number NL of images, NL images (including the central image) are selected in order of proximity to the central image with respect to the camera position, and this image group candidate is designated as the image group candidate that contains the selected NL images.

[0055] In this manner, the classification pre-processing unit 112 finally obtains K image group candidates (I1, I2, . . . , I K ) is determined for each image group candidate (I1, I2, . . . , I K ) is robust to changes in illumination and contains elements (images) that are within a given threshold range for both scale and orientation. In other words, the pre-classification process described above is an ideal preparation for the image group determination process (image classification process) that will be performed next.

[0056] Furthermore, in this embodiment, the above-mentioned pre-classification processing is a clustering process based on three-dimensional geometric information related to landmarks (targets) and image feature representation (scale and orientation) (Geometry and Feature Representation-based Clustering). Therefore, as can be seen from the content, in the pre-classification processing of this embodiment, unlike the conventional method, it is not necessary to perform labeling processing on the original image group when performing clustering processing as pre-classification processing.

[0057] (Image classification processing) Also in the functional block diagram of FIG. 1, the image classification unit 113 classifies an image group I (=I1∪I2∪ ∪ ∪I K ) are compared with the feature vectors of the image and the candidate image groups (I1, I2, . . . , I K ) and the representative values ​​of the features related to the landmarks, the images are "re"classified into multiple image groups, each of which is considered to relate to one of the landmarks.

[0058] First, the image classification unit 113 classifies the above image group I (=I1∪I2∪ ∪ ∪I K ) from each image f included in the (3) x = φ(f) A feature amount x, which is vector information, is extracted by the above. Various neural networks φ can be used as long as they are suitable for extracting image features, but it is preferable to use the one disclosed in Non-Patent Document 1, for example. Of course, instead of extracting features at this stage, feature amounts x may be extracted in advance from each image included in the original image group retrieved from the image database 2, and these may be used appropriately thereafter.

[0059] Next, in this embodiment, the image classification unit 113 performs clustering processing of the images (feature quantities) using k-means clustering. This processing is disclosed, for example, in Non-Patent Document: David MacKay, "Chapter 20. An Example Inference Task: Clustering". Information Theory, Inference and Learning Algorithms. Cambridge University Press, pp. 284-292, 2003. Note that this clustering processing can also be performed by other methods, for example, DBSCAN (Density-Based Spatial Clustering of Applications with Noise).

[0060] Specifically, first, the following formula is used as a "seed" for the clustering process. (4) C k (0) = {x|x=φ(f), f∈I k} Here, k = 1, 2, …, K (K is the number of image group candidates above). By this, the cluster (C1 (0) , C2 (0) , , C K (0) ) is set. Then, (5) m k (0) =(Σ i∈Ck(0) x i ) / |C k (0) | Using the cluster average (m1 (0) , m2 (0) ,...,m K (0) ) is calculated, where Σ i∈Ck(0) x i is cluster C k (0) Feature x (a vector quantity) contained in i is the sum (total) of |C k (0)| is cluster C k (0) Feature x contained in i is the number of.

[0061] Next, an allocation process is (repeatedly) performed in which the feature quantity that is closest to the cluster average quantity of a certain cluster among the cluster average quantities of all the clusters is newly allocated to the certain cluster. Specifically, the following equation is used: (6) m k (t-1) =(Σ i∈Ck(t-1) x i ) / |C k (t-1) |As C k (t) = {x| j is any integer from 1 to K (except k) |xm k (t-1) | 2 <|xm j (t-1) | 2 is satisfied} Then, starting with t=1, we increment t by 1 and create a cluster (C1 (t) , C2 (t) , , C K (t) ) is calculated. In other words, the assignment process of the above formula (6) is performed as the assignment of feature quantity x (x → C k (t) This is repeated T(≧1) times until the equation stabilizes (no longer changes).

[0062] Here, t (=1, 2, . . . , T) is a parameter indicating the number of times the allocation process is performed. 2 is the square of the Euclidean distance between x and m. Incidentally, according to the above formula (6), each feature x is always assigned to one C k (t) will be assigned only to

[0063] Next, the image classification unit 113 classifies the last cluster (C1 (T) , C2(T) , , C K (T) ) for each cluster C k (T) The feature value x of the image contained in each feature matrix X k A feature matrix group X included in (k=1, 2, . . . , K), that is, (7) X = [X 1 , X 2 , , X K ] Here, the feature matrix X k (∈R D×Nk ,D is the number of dimensions of feature x, and Nk is X k The number of features in is expressed as the number of features x(∈C k (T) ), that is, it is a matrix that represents the overall features of the corresponding image group (cluster).

[0064] Each feature matrix X determined in this way k Image group I corresponding to k (k=1, 2, . . . , K) is a suitable search pre-processing result for the images extracted from the image database 2. That is, the image classification unit 113 is capable of generating a classification result in which images including the same "target (such as a landmark)" generally belong to the same image group, regardless of the appearance of the images, such as differences in the camera position and orientation, the presence or absence of image inversion, and even differences in lighting. Of course, such classification results into image groups are very useful in performing image search processing.

[0065] 4 shows a specific example of the result of the image classification process described above. In this specific example, the images in image group I are "re" classified into clusters (image groups) 1, 2, 3, 4, .... Incidentally, such a search pre-processing result, specifically the classification result into image groups (a collection of image data linked with the index k of the image group to which it belongs), is returned to the image database 2 via the input / output interface unit 101 and may be used separately (by the image group search unit 121 or another device), or may be output directly to the image group search unit 121, which will be described later.

[0066] (Image group search processing) Returning to the functional block diagram of FIG. 1, the image group search unit 121 searches the plurality of feature quantity matrices (X 1 , X 2 , , X K ) corresponding to multiple image groups (I1, I2, . . . , I K ), an image group I close to (similar to) the query image fq is selected based on the feature amount x (=φ(f)) of each image f belonging to the image group and the feature amount y (=φ(fq)) of the query image (target image) fq (input via the input / output interface unit 101). k* In other words, we determine the class k* associated with the query image fq.

[0067] Specifically, the image group search unit 121 searches for an image group I such that the difference (average distance) between each column of the feature matrix corresponding to the image group I and the feature y of the query image fq is the smallest. k* That is, the following equation can be used to determine (8) k*=argmin k ((Nk) -1 ×Σ i |X i k -y|2 2 ) Using the above, we extract a group of images I close to the query image fq. k* It is possible to determine the class k* of k(·) is a function that outputs the k value (k=1, 2, ···, K) that minimizes the contents of the parentheses ·. Also, Σ i is the feature matrix X k Column X of i k Furthermore, |·|2 represents the L2-norm disclosed in Non-Patent Documents 5 and 6.

[0068] FIG. 5 is a schematic diagram for explaining a specific example of the image group search process in the image group search section 121. As shown in FIG.

[0069] Fig. 5(A) shows the distribution of feature amount x of image f in image feature amount space. The image group search unit 121 performs image group search based on the distance (in feature amount space) between the feature amount (circle, cross, diamond) of each image included in each image group (A, B, C) and the feature amount of the query image. Specifically, as shown in Fig. 5(B), the sum of the squares of the distances in each image group (Σ i |X i k -y|2 2 ) based on the above equation (8), the class of the query image is determined.

[0070] Returning to the functional block diagram of FIG. 1, the image group search unit 121 calculates the following equation in order to reduce the calculation time in such an image group search process: (9) k*=argmin k (|av_X k -y|2 2 ) Using the above, we extract a group of images I close to the query image fq. k* where av_X may be determined. i k is Image Group I k is the average value of the feature x of all images f in the image group I k This is the average feature value of

[0071] As a further modification of the image group search (determination of a class related to a query image), the image group search unit 121 determines (a) a feature matrix (X 1 , X 2 , , X K ) into the expression coefficient vector (α 1 , α 2 , , α K ) and the feature quantity y of the query image fq. (b) The representation coefficient vector α* is determined so as to minimize the difference between the amount multiplied by the representation coefficient vector α* and the feature quantity y of the query image fq. K ) and the feature value y of the query image fq, an image group I close to (similar to) the query image fq is created. k* Incidentally, such an image group search process, which will be specifically described below, is a technique that has been used in the field of facial recognition, as disclosed in, for example, Non-Patent Documents 5 and 6.

[0072] Specifically, we define the expression coefficient vector group as α(=[α 1 , α 2 , , α K ]) as follows: (10) α*=argmin α (|Xα-y|2 2 +λ×|α|2 2 ) Based on this, we select a linear combination of features (matrices) as a collaborative representation of the query image fq, that is, Xα(≒fq), and determine the representation coefficient vector α*. Note that the argmin α The first term in ensures that this linear combination represents the query image fq, and the second term suppresses the contribution of image groups that are far from the query image fq.

[0073] Here, the following equation is used as a feasible closed-form solution of the above equation (10): (11) α*=Qy Q=(X T X+λI) -1 X T In the above equation (11), I is the unit matrix group, and (·) -1 teeth(·) -1 (·)=I. Furthermore, X T is the transpose of X. Next, by using α* calculated by the above formula (11), a group of images I that are close (similar) to the query image fq is k* The class k* of (12) k*=argmin k (|X k α* k -y|2 / |α* k |2) It is determined by.

[0074] In comparison with the first formula (8), the above formula (12) adjusts the distance between the feature x and feature y for the class k using the representation coefficient vector group α* as the weight, and therefore the class k* can be determined with higher accuracy. In addition, it is true that an additional calculation process is required to calculate this representation coefficient vector group α*. However, this is basically a multiplication process of a vector and a matrix.

[0075] Moreover, the coefficient matrix group Q in the above formula (11) does not depend on the feature y of the query image fq, and therefore the feature matrix group X(=[X 1 , X 2 , , X K ]) can be calculated in advance, for example, offline, when a database of the query image fq is generated. As a result, according to the above formula (12), when a query image fq is given, for example, an image group I close to the query image fq can be calculated online. k* It also becomes possible to determine the class k* with greater accuracy.

[0076] As a further modification for improving the robustness of the image group search process, the image group search unit 121 may use the following formula: (13) α*=argmin α (|Xα-y|2 2 +λ×|α|2 2 +(γ / K)×Σ k=1 K |Xα-X k α k |2 2 ) It is also preferable to determine the expression coefficient vector α* based on the above formula (13). Compared to the above formula (10), this formula (13) further includes a third term (Σ k=1 K |Xα-X k α k |2 2 ) is added. This third term is as close as possible to the final linear combination Xα (for each image group I k (related to feature matrix X k and the expression coefficient vector α k It works to realize the product of and.

[0077] Here, the following equation is used as a feasible closed-form solution of the above equation (13). (14) α*=Q'y Q'=(X T X+(γ / K)×Σ k=1 K (X^ k ) T X^ k +λI) -1 X T Here, X^ k =XX' k ,X' k = [0, , X k , , 0] In the above formula (14), 0 is a zero matrix. Next, by using α* calculated by the above formula (14), a group of images I close to the query image fq can be obtained. k* The class k* of (15) k*=argmin k (|X k α* k -Xα|2 2 ) It is determined by.

[0078] Similarly to the above Q, the coefficient matrix group Q′ does not depend on the feature quantity y of the query image fq, and therefore can be calculated in advance, for example, offline. As a result, according to the above formula (15), when a query image fq is given, for example, an image group I close to the query image fq can be calculated online. k* It also becomes possible to determine the class k* with greater accuracy.

[0079] (Image ranking process) Also in the functional block diagram of FIG. 1, the image ranking unit 122 ranks the retrieved image group I k* Based on the feature value x of each image f included in the image group I and the feature value y of the query image fq, k* Determine the image f that is closest (similar) to the query image fq from among the images f in the image group I k* Ranking information r (e.g., information obtained by sorting image indexes in ascending order of ranking) in order of proximity (similarity) to the query image fq is assigned to the images f included in the query image fq. Here, for example, the ranking information may be information in which each image f is labeled with the ranking value (1, 2, ...) of the image, or table information in which the index of each image f corresponds to the ranking value of the image. Furthermore, the ranking information generated in this manner may be transmitted via the input / output interface unit 101 to an external information processing device, for example, a device that has transmitted a request including the query image fq, and may be used therein.

[0080] Specifically, in this embodiment, the image ranking unit 122 uses the following formula: (16) r=argsort i (|X i k* -y|2) Using this, we can generate ranking information r, where X i k* (∈X k* ) by definition, Image Group I k* These are the features of each image contained in argsort.i (·) is a function that returns the result of sorting the i values ​​for feature index i in order of decreasing parentheses, starting with the i value that minimizes the parentheses' contents ·.

[0081] FIG. 6 is a schematic diagram for explaining a specific example of the image ranking process in the image ranking section 122. As shown in FIG.

[0082] 6, the image ranking unit 122 generates ranking information r indicating the ranking of images (Da, Db, Dc, Dd) included in the searched image group A that is close (similar) to the query image fq using the above formula (16). In addition, based on the generated ranking information r, the image ranking unit 122 assigns the corresponding ranking value (1, 2, 3, 4) to each image (Da, Db, Dc, Dd).

[0083] As a result, image Da is determined to be the image with the first rank that is closest (similar) to query image fq, and then images Dd, Dc, and Db are identified in order of closest (similar) to query image fq.

[0084] As described above in detail, according to the present invention, a point cloud related to a target (e.g., a landmark) is obtained from an image group that may include images that may be problematic in image retrieval, and then a plurality of point cloud groups each regarded as related to one of the targets can be generated from the point cloud. Here, one point cloud group is classified not simply based on appearance, but taking into consideration the characteristics of, for example, its three-dimensional shape (the distance between points in three-dimensional space), and is regarded with a high degree of certainty as "a group related to the corresponding target."

[0085] As a result, the image group corresponding to this point cloud group can be regarded as a collection of images including the target. As described above, according to the present invention, it is possible to perform pre-search processing such as image classification for each target, which does not depend on the mere appearance of the images, on the above image group.

[0086] Furthermore, the pre-search processing according to the present invention or the image search processing using this processing can be utilized in the analysis of a huge amount of image data collected by a large number of users taking pictures of various landmarks in a city and uploading the camera images, and a huge amount of image data generated periodically by a large number of street cameras, to facilitate prediction of pedestrian flow in a city, detection of installed objects and suspicious objects and confirmation of their trends, and prediction and detection of trouble and crime occurrence, etc. In other words, the present invention can also contribute to Goal 11 "Make cities inclusive, safe, resilient and sustainable" of the Sustainable Development Goals (SDGs) led by the United Nations.

[0087] Furthermore, the pre-search processing according to the present invention or the image search processing using this processing can be used to analyze a huge amount of on-site image data collected by many users taking photos of a target area or a target sea area and uploading the camera images, and various conditions in such areas and sea areas, such as crop growth conditions, the current state of the ecosystem, and the impact of climate change, can be investigated. In other words, the present invention can contribute to Goal 13 "Take urgent action to combat climate change and its impacts," Goal 14 "Conserve and sustainably use the oceans and marine resources," and Goal 15 "Sustainably manage forests, combat desertification, halt and reverse land degradation, and halt biodiversity loss" in the Sustainable Development Goals (SDGs) led by the United Nations.

[0088] Regarding the various embodiments of the present invention described above, various changes, modifications, and omissions within the scope of the technical idea and perspective of the present invention can be easily made by those skilled in the art. The above description is merely an example and is not intended to be restrictive in any way. The present invention is limited only by the scope of the claims and their equivalents. [Explanation of symbols]

[0089] 1. Image search device (pre-search processing device) 101 Input / Output Interface (IF) Section 111 Point cloud classification section 112 Classification preprocessing section 112a Image group candidate determination unit 112b Center / periphery image determination unit 112c Scale orientation judgement section 112d New Image Group Candidate Determination Department 113 Image Classification Unit 121 Image Group Search Section 122 Image Ranking Unit 2. Image Database (DB)

Claims

1. A search pre-processing program for performing pre-processing for a search on an image group from which an image similar to a target image including one of a plurality of targets is to be searched, the program comprising: a point cloud classification means for classifying a point cloud related to the target included in the image group, the point cloud being generated based on feature points of each image included in the image group, into a plurality of point cloud groups, each of which is regarded as relating to one of the targets, based on the proximity between points of the point cloud; A classification pre-processing means for determining, for each of the point cloud groups, an image group candidate that is a set of the images corresponding to the point cloud group; an image classification means for reclassifying each image included in the image group into a plurality of image groups each of which is regarded as relating to one of the targets based on a feature value of the image and a representative value of the feature values ​​related to the candidate image group; A search preprocessing program for causing a computer to function as described above.

2. The search preprocessing program according to claim 1, characterized in that the point cloud relating to the target included in the image group is generated based on feature points of each of the images and camera information relating to each of the images determined from the feature points.

3. The search pre-processing program according to claim 1, characterized in that, for each of the determined image group candidates, the classification pre-processing means determines a central image corresponding to the target based on the number of feature points that are the basis of the point cloud related to the target, sets other images as peripheral images, extracts peripheral images that differ from the central image by a predetermined amount or more in terms of image scale and / or orientation, removes them from the image group candidate, and determines a new image group candidate including the extracted peripheral images.

4. The search pre-processing program according to claim 3, characterized in that the classification pre-processing means extracts peripheral images that satisfy at least one of the following conditions: a ratio of their own scale to the scale of the central image is outside a specified range, and a difference between their own orientation and the orientation of the central image is equal to or greater than a specified value, and determines new image group candidates.

5. An image search program for searching an image group for an image similar to a target image including one of a plurality of targets, comprising: a point cloud classification means for classifying a point cloud related to the target included in the image group, the point cloud being generated based on feature points of each image included in the image group, into a plurality of point cloud groups, each of which is regarded as relating to one of the targets, based on the proximity between points of the point cloud; A classification pre-processing means for determining, for each of the point cloud groups, an image group candidate that is a set of the images corresponding to the point cloud group; an image classification means for reclassifying each image included in the image group into a plurality of image groups each of which is regarded as relating to one of the targets based on a feature value of the image and a representative value of the feature values ​​related to the candidate image group; an image group search means for searching for an image group that is similar to the target image from among the plurality of image groups, based on the feature amounts of each image belonging to the image group and the feature amount of the target image; and causing a computer to function as an image search program.

6. The image search program according to claim 5, characterized in that the image group search means determines a coefficient vector that reduces the difference between the amount obtained by multiplying a feature matrix representing the overall features of the image group by a coefficient vector and the feature of the target image, and searches for an image group that is close to the target image based on the determined coefficient vector, the feature of each image belonging to the image group, and the feature of the target image.

7. 7. The image retrieval program according to claim 5 or 6, further comprising a computer that functions as an image ranking means for determining an image that is closest to the target image from among the images included in the image group based on the features of each image included in the searched image group and the features of the target image, or for assigning ranking information to the images included in the image group in order of closeness to the target image.

8. A search pre-processing device that performs pre-processing for a search on an image group from which an image similar to a target image including one of a plurality of targets is searched, the search pre-processing comprising: a point cloud classification means for classifying a point cloud related to the target included in the image group, the point cloud being generated based on feature points of each image included in the image group, into a plurality of point cloud groups, each of which is regarded as relating to one of the targets, based on the proximity between points of the point cloud; A classification pre-processing means for determining, for each of the point cloud groups, an image group candidate that is a set of the images corresponding to the point cloud group; an image classification means for reclassifying each image included in the image group into a plurality of image groups each of which is regarded as relating to one of the targets based on a feature value of the image and a representative value of the feature values ​​related to the candidate image group; A search preprocessing device comprising:

9. A search pre-processing method for performing pre-processing for a search on an image group from which an image similar to a target image including one of a plurality of targets is to be searched, the method comprising: classifying a point cloud associated with the target in the set of images, the point cloud being generated based on feature points of each image in the set of images, into a plurality of point cloud groups, each of which is deemed to be associated with one of the targets, based on the proximity between points in the point cloud; determining, for each of the point cloud groups, a candidate image group that is a set of the images corresponding to the point cloud group; reclassifying each image included in the image group into a plurality of image groups each of which is considered to be related to one of the targets based on the feature value of the image and a representative value of the feature values ​​related to the candidate image group; 13. A computer-implemented search pre-processing method comprising:

Citation Information

Patent Citations

  • System and method for retrieving data

    JP2002342374A

  • Data set creation method, data set creation device, and data set creation program

    JP2020135679A

  • Content -based similarity retrieval system for image data

    US20020178149A1

  • 3D view model generation of an object utilizing geometrically diverse image clusters

    US20210019937A1

  • Statistically based image group descriptor particularly suited for use in an image classification and retrieval system

    US5819288A