Information processing device, information processing method, and recording medium

The information processing apparatus uses depth and object relationships to enhance image similarity searches, addressing inaccuracies in existing methods by accurately determining similar images based on object features.

WO2025150467A1PCT designated stage expired Publication Date: 2025-07-17NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/000016
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2025-01-06
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing image retrieval techniques may incorrectly determine similarity between images due to clustering feature amounts based on depth values, leading to false positives or negatives in identifying similar images.

Method used

An information processing apparatus and method that utilizes depth information of objects within images to accurately determine similarity by extracting and comparing feature amounts, including depth, posture, and relationships between objects in query and target images.

Benefits of technology

Enables accurate retrieval of similar images by considering the depth and positional relationships of objects, reducing false detections and improving the accuracy of image similarity searches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000016_17072025_PF_FP_ABST
    Figure JP2025000016_17072025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing device according to the present invention includes a query acquisition unit that acquires a query image including a plurality of objects, a search unit that, on the basis of features extracted regarding the plurality of objects included in each of a plurality of target images and the query image, searches for a similar image, which is an image that satisfies a predetermined similarity condition with respect to the query image, from among the plurality of target images, and an output control unit that causes an output unit to output search results. The features include depth information regarding depth of a plurality of objects included in each image of the plurality of target images and the query image.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and recording medium

[0001] The present invention relates to an information processing device, an information processing method, and a recording medium.

[0002] Various techniques have been disclosed for searching for images similar to a query image from images stored in a database, etc. For example, Patent Document 1 discloses an image processing device that includes an extraction unit, a specification unit, and a determination unit.

[0003] The extraction means described in Patent Document 1 extracts first local features for each of a plurality of feature points in a first image and second local features for each of a plurality of feature points in a second image obtained by capturing an image of the image capture target. The storage means stores depth information of the image capture target corresponding to each of the feature points in the second image. The identification means identifies at least one distance range based on the depth information. The determination means determines that the second image is similar to the first image if the similarity between the second local feature and the first local feature obtained for the distance range is equal to or greater than a predetermined value.

[0004] According to Patent Document 1, a feature amount identification unit, which corresponds to an identification means, first identifies multiple distance ranges based on depth values ​​of the image capture target corresponding to the positions of feature points. For example, if the depth values ​​of the image capture target are distributed over a distance range of 10 to 100 meters, the distance range of 10 to 100 meters is divided into multiple distance ranges with 10-meter intervals, such as 10 to 20 meters, 20 to 30 meters, and 30 to 40 meters. This process is referred to as depth value gradation in Patent Document 1. After gradating the depth values, the feature amount identification unit clusters the features based on the depth values.

[0005] According to Patent Document 1, a similarity determination unit, which corresponds to a determination means, identifies clusters having a predetermined number of pairs based on the clusters of the comparison target image, and generates a similarity between the comparison source image and the comparison target image from the number of pairs of the cluster. According to Patent Document 1, the number of cluster pairs is determined by comparing the feature amounts of the comparison source image with the feature amounts of the comparison target image, finding pairs of feature amounts of the comparison target image that are similar to the feature amounts of the comparison source image, and counting the number of pairs for each cluster of the comparison target image.

[0006] JP 2016-181181 A

[0007] However, with the technology described in Patent Document 1, there is a risk that similar images that are similar to the query image will not be included in the search results.

[0008] For example, suppose a query image (comparison source image) depicts a person holding a cup, and a comparison target image includes a person holding the cup. In this case, according to the technology described in Patent Literature 1, when clustering feature amounts based on depth values, the cup and the person may be clustered into different clusters. As a result, even though the comparison target image includes a person holding a cup, the comparison target image may be determined to be dissimilar to the query image.

[0009] For example, suppose a query image (comparison source image) depicts a person holding a cup, and a comparison target image includes both the cup and a person not holding the cup. In this case, according to the technology described in Patent Literature 1, when clustering feature amounts based on depth values, the cup and the person may be clustered in the same cluster. As a result, even though the person in the comparison target image is not holding the cup, the comparison target image may be determined to be similar to the query image.

[0010] One of the objectives of the present disclosure is to accurately search for similar images that are similar to a query image.

[0011] The information processing device according to the present disclosure includes a query acquisition means for acquiring a query image including a plurality of objects; a search means for searching, from among a plurality of target images, for similar images that are images that satisfy a predetermined similarity condition with the query image, based on feature amounts extracted for a plurality of objects included in each of the plurality of target images and the query image; and an output control means for causing an output means to output the search results, wherein the feature amounts include depth information regarding the depths of the plurality of objects included in each of the plurality of target images and the query image.

[0012] The information processing method of the present disclosure includes one or more computers acquiring a query image including a plurality of objects, searching for similar images from among the plurality of target images that satisfy predetermined similarity conditions with the query image based on features extracted for the plurality of objects included in each of the plurality of target images and the query image, and outputting the search results to an output means, wherein the features include depth information regarding the depths of the plurality of objects included in each of the plurality of target images and the query image.

[0013] The recording medium of the present disclosure has recorded thereon a program that causes one or more computers to acquire a query image including a plurality of objects, search for similar images from among the plurality of target images that satisfy predetermined similarity conditions with the query image based on feature amounts extracted for the plurality of objects included in each of the plurality of target images and the query image, and output the search results to an output means, wherein the feature amounts include depth information regarding the depths of the plurality of objects included in each of the plurality of target images and the query image.

[0014] According to the present disclosure, it is possible to accurately search for similar images that are similar to a query image.

[0015] 5 is a block diagram showing an example configuration of a first information processing device according to the present disclosure. FIG. 6 is a flowchart showing an example processing operation of a first information processing device according to the present disclosure. FIG. 7 is a block diagram showing a detailed example configuration of a first information processing device according to the present disclosure. FIG. 8 is a flowchart showing a detailed example processing operation of a first information processing device according to the present disclosure. FIG. 9 is a diagram showing an example query image. FIG. 10 is a diagram showing an example target image. FIG. 11 is an example diagram showing, in monochrome shading, the depth calculated for the query image exemplified in FIG. 5. FIG. 12 is an example diagram showing, in monochrome shading, the depth calculated for the target image exemplified in FIG. 6. FIG. 13 is a block diagram showing an example physical configuration of a first information processing device according to the present disclosure. FIG. 14 is a block diagram showing an example configuration of a first information processing device according to the present disclosure. FIG. 15 is a flowchart showing an example processing operation of a first information processing device according to the present disclosure.

[0016] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all drawings, similar components are designated by similar reference numerals, and descriptions thereof will be omitted as appropriate. In addition, in this disclosure, the drawings relate to one or more embodiments.

[0017] First Embodiment (Configuration Example of Information Processing Apparatus 100) As shown in FIG. 1, the information processing apparatus 100 includes a query acquisition unit 110, a search unit 140, and an output control unit 150.

[0018] The query acquisition unit 110 acquires a query image including a plurality of objects.

[0019] The search unit 140 searches for similar images, which are images that satisfy predetermined similarity conditions between the target images and the query image, from among the target images based on feature amounts extracted for multiple objects contained in each of the target images and the query image.

[0020] The output control unit 150 causes the output unit to output the search results.

[0021] The feature amount includes depth information relating to the depths of a plurality of objects included in each of the plurality of target images and the query image.

[0022] According to this information processing device 100, depth information relating to the depths of multiple objects contained in multiple target images and the query image is used to search for similar images that are similar to the query image. Therefore, it is possible to search for target images that include multiple objects that have similar depth relationships to the multiple objects contained in the query image as similar images. Therefore, it is possible to accurately search for similar images that are similar to the query image.

[0023] (Example of Processing Operation of Information Processing Device 100) The information processing device 100 executes information processing as shown in FIG.

[0024] The query acquisition unit 110 acquires a query image including a plurality of objects (step S110).

[0025] The search unit 140 searches for similar images from among the plurality of target images that satisfy predetermined similarity conditions between the plurality of target images and the query image based on features extracted for the plurality of objects contained in each of the plurality of target images and the query image (step S140).

[0026] The output control unit 150 causes the output unit to output the search results (step S150).

[0027] The feature amount includes depth information relating to the depths of a plurality of objects included in each of the plurality of target images and the query image.

[0028] According to this information processing, depth information relating to the depths of multiple objects contained in multiple target images and the query image is used to search for similar images that are similar to a query image. Therefore, target images that include multiple objects that have similar depth relationships to the multiple objects contained in the query image can be searched for as similar images. Therefore, it is possible to accurately search for similar images that are similar to the query image.

[0029] Second Embodiment A detailed example of the information processing device 100 and its processing operation will be described.

[0030] (Detailed configuration example of information processing device 100) As shown in FIG. 3, the information processing device 100 includes a query acquisition unit 110, a search unit 140, and an output control unit 150 similar to those in embodiment 1, a memory unit 120, an extraction unit 130, and an output unit 160.

[0031] The storage unit 120 stores a plurality of target images in advance.

[0032] The extraction unit 130 extracts feature amounts of a plurality of objects included in each of a plurality of target images and a query image.

[0033] (Detailed Example of Processing Operation of Information Processing Device 100) The information processing device 100 executes information processing as shown in Fig. 4. Step S110, similar to that of the first embodiment, is executed.

[0034] The extraction unit 130 extracts feature amounts of a plurality of objects included in each of a plurality of target images and a query image (step S130).

[0035] Steps S140 and S150 are executed in the same manner as in the first embodiment.

[0036] Below, detailed examples of each function, each process, etc. will be described.

[0037] (Regarding the query acquisition unit 110) The query acquisition unit 110 acquires a query image including multiple objects, for example, based on a user instruction. For example, the storage unit 120 may further store the query image in advance, and the query acquisition unit 110 may acquire the query image from the storage unit 120 based on a user instruction. Also, for example, the query acquisition unit 110 may acquire the query image from a camera based on a user instruction. Note that the method of acquiring the query image is not limited to the example given here.

[0038] (Query Image) The query image includes a plurality of objects. The objects may be people or objects. The plurality of objects may include at least one person. Note that the plurality of objects may all be people or all be objects.

[0039] 5 is a diagram showing an example of a query image. The query image shown in the figure is the rightmost person among three people included in the image shown in the figure, i.e., the person holding a cup. This query image is an example in which the multiple objects included in the query image include at least one person.

[0040] Although an example in which the query image is a still image will be described, the query image may also be a video composed of a plurality of frame images.

[0041] (Regarding the Storage Unit 120) The storage unit 120 stores a plurality of target images in advance.

[0042] 6 is a diagram showing an example of a target image. The target image shown in the figure is an example that includes a person P, who is third from the right among four people surrounding a table in the foreground included in the image shown in the figure, and a cup Q. As shown in the figure, person P is not holding cup Q.

[0043] However, the cup Q is shown close to the person P. Therefore, for example, with the technology described in Patent Document 1, there is a risk that the person P may be erroneously detected as holding the cup Q.

[0044] According to the present disclosure, similar images are searched for using depth information relating to the depths of multiple objects contained in a target image and a query image. This reduces the possibility that a person P, as shown in FIG. 5 , is mistakenly detected as a person holding a cup Q, and that a target image including the person P and the cup Q is mistakenly detected as an image similar to the query image exemplified in FIG. 5 (a similar image). This will be more clearly understood from the following explanation.

[0045] Although an example in which the target image is a still image will be described, the target image may also be a video composed of a plurality of frame images. Furthermore, both the query image and the target image may be still images or videos. Furthermore, the query image and the target image may differ in whether they are still images or videos, such as when the query image is a still image and the target image is a video, or when the query image is a video and the target image is a still image.

[0046] (Regarding the Extraction Unit 130) As described above, the extraction unit 130 extracts feature amounts of a plurality of objects included in each of a plurality of target images and a query image. For example, the extraction unit 130 detects a plurality of objects included in each of a plurality of target images and a query image, and extracts feature amounts of the detected plurality of objects.

[0047] As described above, the feature amount includes depth information relating to the depth of a plurality of objects included in each of the plurality of target images and the query image. The depth may be, for example, a value corresponding to the distance from a camera that captured each of the query image and the plurality of target images to each of the objects included in each of the images.

[0048] The feature amount may further include at least one of the orientation of each of a plurality of objects included in each of the plurality of target images and the query image, and the relationship between the plurality of objects.

[0049] The relationship may be, for example, a relationship regarding the actions, states, positions, etc. of multiple objects. In detail, the relationship may be, for example, an action performed by a person using an object, a state relationship between a person and an object, a positional relationship between a person and an object, etc., but is not limited to these. The positional relationship may be, for example, the distance between a person and an object. The position here may be a position in an image, and may be expressed, for example, by the number of pixels. For example, when objects overlap and the positional relationship is expressed by the distance between them, the positional relationship may be zero.

[0050] (Feature Amounts of Query Image) For example, for a query image, the extraction unit 130 detects multiple objects included in the query image acquired by the query acquisition unit 110 and extracts feature amounts of the detected multiple objects. The feature amounts may include depth information related to the depths of the multiple objects included in the query image. The feature amounts may further include at least one of the orientation of each object and the relationship between the multiple objects included in the query image.

[0051] The depth of a query image may be a value corresponding to the distance from the camera that captured the query image to each object included in the query image, for example. In other words, the depth can be said to be a value corresponding to the distance from the camera that captured the query image to each object, with the camera being used as a reference.

[0052] In general, an object is often detected not as a point but as an object region corresponding to the object. In such cases, for example, when the object is composed of a curved surface or when the object region includes regions other than the object, such as the background and foreground, the distance from the camera to the object may vary within the object region.

[0053] The object region may be, for example, a region that surrounds the object in a predetermined shape such as a rectangle. The object region may also be, for example, a region along the outer edge of the object. The region along the outer edge of the object may be identified, for example, by detecting a rectangular region corresponding to the object and then segmenting the region to remove the background, foreground, etc. The technique for identifying such a region along the outer edge of the object is not limited to the example given here, and any general technique may be used.

[0054] The depth of the query image may be a value obtained by statistically processing the depth within the object region corresponding to each object included in the query image. The statistical processing may be, for example, determining the mode, calculating the average, determining the maximum, etc. In these examples, the depth may be the mode, average, maximum, etc. of the depth within the object region corresponding to each object included in the query image.

[0055] Using this method, the extraction unit 130 may calculate the depth of the person as "50" and the depth of the cup as "48" for the query image shown in Fig. 5. Fig. 7 is an example of a diagram showing the depth calculated for the query image shown in Fig. 5 in monochrome shading.

[0056] Furthermore, the extraction unit 130 may calculate a feature amount including a value indicating that the person is sitting and a value indicating that the cup Q is standing, for example, in the query image exemplified in Fig. 5. The extraction unit 130 may calculate a feature amount including a value indicating that the cup is near the person's hand, as the relationship between the person and the cup, for example, in the query image exemplified in Fig. 5.

[0057] Here, an object being close means, for example, that the object is included in a range of pixel values ​​in the image that is equal to or less than a predetermined value, and the same applies hereinafter.

[0058] (Regarding feature amounts of target images) For example, regarding target images, the extraction unit 130 acquires a plurality of target images from the storage unit 120. Then, the extraction unit 130 detects a plurality of objects included in each of the acquired plurality of target images, and extracts feature amounts of the detected plurality of objects.

[0059] The feature amount may include, for each target image, depth information relating to the depths of a plurality of objects included in the target image. The feature amount may further include, for each target image, at least one of the orientation of each of the plurality of objects included in the target image and the relationship between the plurality of objects.

[0060] The depth, posture, and relationship of the target image may correspond to those explained in the example of the query image.

[0061] That is, the depth of the target image may be a value corresponding to the distance from the camera that captured the target image to each object included in the target image. In other words, the depth may be a value corresponding to the distance from the camera that captured the target image to each object, using the camera that captured the image as a reference, as in the case of the query image.

[0062] The depth of the target image may be a value obtained by statistically processing the depth within an object region corresponding to each object included in the target image. As with the query image, the statistical processing may be, for example, determining the mode, calculating the average, or determining the maximum value. In these examples, the depth of the target image may be the mode, average, or maximum value of the depth within an object region corresponding to each object included in the target image.

[0063] Using this method, the extraction unit 130 may calculate, for example, the depth of person P as "60" and the depth of cup Q as "30" for person P and cup Q included in the target image illustrated in Fig. 6. Fig. 8 is an example of a diagram in which the depth calculated for the target image illustrated in Fig. 6 is represented by monochrome shading.

[0064] Furthermore, the extraction unit 130 may calculate a feature amount including a value indicating that the posture of the person P is sitting and a value indicating that the posture of the cup Q is standing in the target image illustrated in Fig. 6, for example. The extraction unit 130 may calculate a feature amount including a value indicating that the cup Q is near the hand of the person P, as the relationship between the person P and the cup Q in the target image illustrated in Fig. 6, for example.

[0065] The extraction unit 130 may use a general technique to extract the feature amount, and one example thereof is to use a machine learning model configured with a neural network.

[0066] The extraction unit 130 may include, for example, an object detection model, a depth estimation model, a relationship estimation model, and a posture estimation model.

[0067] An object detection model is a machine learning model that detects objects from an input image.

[0068] The depth estimation model is a machine learning model that estimates the depth of each of multiple objects detected in an input image.

[0069] The relationship estimation model is a machine learning model that estimates the relationship between objects in each pair of multiple objects detected in an input image. Each pair of multiple objects is a pair consisting of any two objects from the multiple detected objects.

[0070] The pose estimation model is a model that estimates the pose of each of multiple objects detected in an input image.

[0071] Each of these models may take an input of information indicating an input image or a plurality of objects detected in the input image, and output a feature quantity according to its function. Furthermore, each of these models may be a model trained using training data including information indicating an input image or a plurality of objects detected in the input image, and ground truth data.

[0072] These multiple models may be composed of individual neural networks, or some or all of them may be composed of a series of neural networks. Furthermore, the feature quantities may be represented by one or more feature vectors indicating, for example, depth, relationship, and posture, and the same applies hereinafter. Each of the one or more feature vectors may be a vector composed of one or more components.

[0073] The depth may be a value converted into a real-world distance based on an object whose distance from the camera is known (for example, a pole fixed to the ground, a lamp, etc.).

[0074] (Regarding the Search Unit 140) The search unit 140 searches for similar images from among the plurality of target images based on feature amounts extracted from a plurality of objects included in each of the plurality of target images and the query image. A similar image is an image (target image) among the plurality of target images that satisfies a predetermined similarity condition with respect to the query image. The search unit 140 may acquire the feature amounts from the extraction unit 130. Note that the information processing device 100 may not include the extraction unit 130, and the search unit 140 may acquire the feature amounts from an external device (not shown) that has the function of the extraction unit 130.

[0075] The predetermined similarity condition may be defined using, for example, the similarity between images formed by combining each of the plurality of target images with the query image (image similarity). In detail, for example, the predetermined similarity condition may include, but is not limited to, that the image similarity is equal to or greater than a predetermined threshold.

[0076] For example, the search unit 140 may calculate the image similarity for each combination of a plurality of target images and a query image based on the feature amounts extracted from each of the target images and the query image.

[0077] In more detail, for example, the search unit 140 calculates the similarity (e.g., cosine similarity) of the extracted feature quantities related to each of the depth, relationship, and posture for each combination of each of the multiple target images and the query image. Then, the search unit 140 may integrate the similarity of the feature quantities related to each of the depth, relationship, and posture to calculate the similarity between the images (image similarity). This integration may use, for example, the average value of the similarity of each feature quantity, or a weighted average in which the similarity of each feature quantity is weighted using a predetermined weight. Note that the integration method is not limited to these.

[0078] Here, the depth similarity may be expressed, for example, as an absolute value or as a value including a sign such as positive or negative. The depth similarity may also be normalized to a value within a predetermined range, for example, from 0 to 100.

[0079] (When depth similarity is expressed as an absolute value) In detail, for example, the depth similarity may be calculated based on the absolute value of the difference in relative distance between two objects constituting each pair of multiple objects included in each image. That is, the similarity between images (image similarity) may be calculated based on the absolute value of the difference in relative distance between two objects constituting each pair of multiple objects included in each image.

[0080] For example, in the query image shown in Fig. 5, the depth of the person is "50" and the depth of the cup is "48". In this case, the depth similarity for the pair of objects consisting of the person and the cup included in the query image may be calculated as 0.98 (= 1 - |48 - 50| / 100). Here, |X - Y| represents the absolute value of the difference, and this also applies hereinafter.

[0081] 6, for example, suppose that the depth of person P is "60" and the depth of cup Q is "30." In this case, the similarity in terms of depth for the pair of objects formed by person P and cup Q included in the query image may be calculated as 0.7 (=1-|30-60| / 100).

[0082] (When depth similarity is expressed as a signed value) In detail, for example, the depth similarity may be calculated for each pair of multiple objects included in each image based on a value including a sign indicating whether one object constituting the pair is farther or closer to the camera with respect to the other object constituting the pair. In other words, the similarity between images (image similarity) may be calculated for each pair of multiple objects included in each image based on a value including a sign indicating whether the other object constituting the pair is farther or closer to the camera with respect to the other object constituting the pair.

[0083] One of the reference objects may be an object of a predetermined type, such as a person, etc. For example, when the objects constituting each pair of multiple objects are of the same type (e.g., people), the one of the reference objects may be an object belonging to a predetermined attribute (e.g., either male or female).

[0084] The depth may also be an absolute value of the relative distance between two objects constituting each pair of objects included in the same image (i.e., the query image and each image of the multiple target images).The depth may also be a value including a code corresponding to whether one object constituting each pair of objects included in the same image (i.e., the query image and each image of the multiple target images) is farther or closer to the camera than the other object constituting the pair.

[0085] For example, in the query image illustrated in Fig. 5, the depth of the person is "50" and the depth of the cup is "48". In this case, for the object pair consisting of the person and the cup included in the query image, the similarity in terms of depth may be calculated as -0.02 (=(48-50) / 100) with the person as the reference.

[0086] 6, for example, suppose that the depth of person P is "60" and the depth of cup Q is "30." In this case, for the pair of objects formed by person P and cup Q included in the query image, the similarity in terms of depth may be calculated as -0.3 (=(30-60) / 100) using person P as the reference, as in the query image.

[0087] (Regarding the Output Control Unit 150 and the Output Unit 160) The output control unit 150 outputs the search results to the output unit 160. The output unit 160 may be a display or a communication interface that transmits the search results to another device (not shown). Note that examples of the output unit are not limited to these.

[0088] The search results may include, for example, at least one of similar images, the number of similar images, etc. Note that the search results are not limited to those exemplified here. The output control unit 150 may display such search results on a display or transmit them to another device via a communication interface.

[0089] (Example of Physical Configuration of Information Processing Device 100) As shown in FIG. 9, the information processing device 100 physically includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, a network interface 1050, an input interface 1060, and an output interface 1070.

[0090] The bus 1010 is a data transmission path for transmitting and receiving data among the processor 1020, memory 1030, storage device 1040, network interface 1050, input interface 1060, and output interface 1070. However, the method of connecting the processor 1020 and the like to each other is not limited to bus connection.

[0091] The processor 1020 is implemented as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit).

[0092] The memory 1030 is a main storage device realized by a RAM (Random Access Memory) or the like.

[0093] The storage device 1040 is an auxiliary storage device realized by a hard disk drive (HDD), a solid state drive (SSD), a memory card, a read-only memory (ROM), or the like. The storage device 1040 stores program modules for realizing the functions of the device that includes the storage device 1040. The processor 1020 loads each of these program modules into the memory 1030 and executes them to realize the function corresponding to that program module.

[0094] The network interface 1050 is an interface for connecting a device equipped with it to a communication network.

[0095] The input interface 1060 is an interface for the user to input information, and is configured from, for example, a touch panel, a keyboard, a mouse, and the like.

[0096] The output interface 1070 is an interface for presenting information to the user, and is configured, for example, by a liquid crystal panel, an organic EL (Electro-Luminescence) panel, or the like.

[0097] In this way, the functions of the information processing device 100 can be realized by the physical components cooperating to execute a software program. Therefore, the present invention may be realized as a software program or as a non-transitory storage medium on which the program is recorded.

[0098] (Operations and Effects) As described above, according to this embodiment, the information processing device 100 includes the query acquisition unit 110 , the search unit 140 , and the output control unit 150 .

[0099] The query acquisition unit 110 acquires a query image including multiple objects. The search unit 140 searches the multiple target images for similar images that satisfy predetermined similarity conditions with the query image, based on feature amounts extracted for multiple objects included in each of the multiple target images and the query image. The output control unit 150 outputs the search results to the output unit. The feature amounts include depth information regarding the depths of multiple objects included in each of the multiple target images and the query image.

[0100] According to this method, depth information relating to the depths of multiple objects contained in multiple target images and the query image is used to search for similar images that are similar to a query image. Therefore, target images that include multiple objects that have similar depth relationships to the multiple objects contained in the query image can be searched for as similar images. Therefore, it is possible to accurately search for similar images that are similar to the query image.

[0101] According to this embodiment, the depth is a value corresponding to the distance from the camera that captured each of the query image and the plurality of target images to each object included in each of the images.

[0102] This allows a search for similar images by searching for target images that contain multiple objects that have similar relationships between the multiple objects included in the query image and the values ​​corresponding to the distances from the camera to the objects included in each of the images, thereby enabling a search for similar images that are similar to the query image with high accuracy.

[0103] According to this embodiment, the plurality of objects included in the query image includes at least one person.

[0104] This allows target images containing people and other objects that have similar depth relationships to the people and other objects contained in the query image to be searched for as similar images, thereby enabling accurate search for similar images that are similar to the query image containing people and other objects.

[0105] According to this embodiment, the depth is a value obtained by statistically processing depths within an object region corresponding to each object included in each of a plurality of target images and a query image, where the object region is a region surrounding the object in a predetermined shape or a region along the outer edge of the object.

[0106] This allows a target image containing multiple objects with similar depth relationships to multiple objects contained in a query image to be searched for as similar images, thereby enabling accurate search for similar images that are similar to the query image.

[0107] According to this embodiment, the feature amount further includes at least one of the posture of each of a plurality of objects included in each of the plurality of target images and the query image, and the relationship between the plurality of objects.

[0108] This makes it possible to search for target images containing multiple objects that are similar to the multiple objects contained in the query image in at least one of the orientations of the objects and the relationships between the multiple objects, thereby enabling more accurate searches for similar images that are similar to the query image.

[0109] According to this embodiment, the predetermined similarity condition is defined using the similarity between images formed by combining each of a plurality of target images with a query image. The similarity between images is calculated based on the absolute value of the difference in relative distance between two objects constituting each pair of a plurality of objects included in each image, or, for each pair of a plurality of objects included in each image, a value including a sign indicating whether one object constituting the pair is farther or closer from the camera than the other object constituting the pair.

[0110] According to this method, to search for similar images to a query image, similarity between images based on the difference in depth between multiple objects contained in each of multiple target images and the query image is used. Therefore, target images containing multiple objects with similar depth relationships to the multiple objects contained in the query image can be searched for as similar images. Therefore, it is possible to search for similar images to the query image with high accuracy.

[0111] [Embodiment 3] For each of a plurality of objects included in each of the query image and a plurality of target images, a plurality of feature points such as joint points of a person, a center of gravity of an object, and top, bottom, left, and right corner points may be detected. Then, a feature amount including depth information may be extracted for each feature point. An example will be described in which the feature amount includes depth information for each feature point of each object included in each of the plurality of target images and the query image.

[0112] In this embodiment, for the sake of simplicity, descriptions that overlap with other embodiments will be omitted as appropriate.

[0113] (Configuration example of information processing device 200) As shown in FIG. 10 , the information processing device 200 includes a query acquisition unit 110, a storage unit 120, an output control unit 150, an output unit 160, an extraction unit 230, and a search unit 240, which are similar to those in embodiment 2.

[0114] The extraction unit 230 extracts feature points for each of a plurality of objects included in each of a plurality of target images and a query image, and the feature amounts for each of the feature points.

[0115] The feature amount may include depth information for each feature point of each object included in each of the plurality of target images and the query image. The object included in each image may include at least one person. That is, the feature amount may include depth information for each feature point of at least one person included in each of the plurality of target images and the query image.

[0116] The search unit 240 searches for similar images from among the plurality of target images based on the feature amounts for each feature point extracted for each of the plurality of objects contained in each of the plurality of target images and the query image.

[0117] (Example of processing operation of information processing device 200) The information processing device 200 executes information processing as shown in Fig. 11. Step S110, similar to that of the first embodiment, is executed.

[0118] The extraction unit 230 extracts feature points for each of a plurality of objects included in each of a plurality of target images and a query image, and the feature amounts for each of the feature points (step S230).

[0119] The search unit 240 searches for similar images from among the plurality of target images based on the feature amounts for each feature point extracted for each of the plurality of objects contained in each of the plurality of target images and the query image (step S240).

[0120] Step S150 is executed in the same manner as in the first embodiment.

[0121] Below, detailed examples of the functions and processes of the extraction unit 230 and the search unit 240 will be described.

[0122] (Regarding the extraction unit 230) The extraction unit 230 extracts feature points for each object included in each of a plurality of target images and a query image, and feature amounts for each of the feature points. For example, the extraction unit 230 detects multiple objects included in each of a plurality of target images and a query image, and extracts feature points for each of the detected multiple objects. Then, for example, the extraction unit 230 extracts feature amounts for each of the extracted feature points. The extraction unit 230 may further extract feature amounts similar to those in the first embodiment for each of the extracted objects.

[0123] For example, if the object is a person, the feature points may be joint points of the person. For example, if the object is an object, the feature points may be one or more predetermined points, such as a center of gravity and top, bottom, left, and right corner points.

[0124] The feature amount for each feature point may include depth information relating to the depth of the feature point. The depth of the feature point is the depth described in the first embodiment applied to the feature point. That is, for example, the depth of the feature point may be a value corresponding to the distance from the camera that captured each of the query image and the multiple target images to each feature point of each object included in each of the images.

[0125] The feature may further include, for each pair of objects included in each of the target images and the query image, the positional relationship between each combination of feature points belonging to different objects that make up the pair.

[0126] As described above, each pair of multiple objects is a pair made up of any two objects from among the multiple detected objects.

[0127] For example, suppose that for a pair consisting of a person P and a cup Q, a plurality of feature points P1 to PN related to the person P and a plurality of feature points Q1 to QM related to the cup Q are extracted. Here, N and M are integers equal to or greater than 1, and at least one of them is an integer equal to or greater than 2. In this case, each combination of a plurality of feature points belonging to different objects constituting the pair is a combination of any one of the feature points P1 to PN belonging to the person P and any one of the feature points Q1 to QM belonging to the cup Q. That is, in this case, combinations of a plurality of feature points belonging to different objects constituting the pair include (P1, Q1)...(P1, QM)...(PN, Q1)...(PN, QM), etc.

[0128] (Regarding the Search Unit 240) The search unit 240 searches for similar images from among the plurality of target images and the query image based on feature amounts extracted for each feature point of a plurality of objects included in each of the plurality of target images and the query image. To search for similar images from among the plurality of target images, the search unit 240 may further use feature amounts similar to those in the first embodiment, i.e., feature amounts extracted for a plurality of objects included in each of the plurality of target images and the query image.

[0129] The similar image may be, for example, an image (target image) among a plurality of target images that satisfies a predetermined similarity condition with respect to the query image, as in embodiment 1. The predetermined similarity condition may be defined, for example, using the similarity (image similarity) between an image formed by combining each of the plurality of target images with the query image, as in embodiment 1.

[0130] In this embodiment, for example, in order to calculate image similarity, in addition to the similarity of the features related to depth, relationship, and posture described in embodiment 1, similarity related to the combination of multiple feature points belonging to the objects that make up the pair may also be used.

[0131] The similarity between combinations of multiple feature points may be, for example, a cosine similarity between feature amounts for each combination of multiple feature points. Furthermore, the similarity between combinations of multiple feature points may include a similarity between the positional relationships of each combination of multiple feature points.

[0132] As described above, according to this embodiment, the information processing device 200 includes the extraction unit 230 that extracts feature points and feature amounts for each of a plurality of objects included in each of a plurality of target images and a query image. The feature amounts include depth information for each feature point of at least one person included in each of the plurality of target images and the query image.

[0133] According to this method, in order to search for similar images that are similar to a query image that includes at least one person, depth information regarding the depths of each feature point of a plurality of objects included in each of a plurality of target images and the query image is used. Therefore, it is possible to search for target images that include a plurality of objects that include feature points that have similar depth relationships for the feature points of the plurality of objects included in the query image as similar images. Therefore, it is possible to search for similar images that are similar to the query image with even greater accuracy.

[0134] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0135] In addition, although the flowcharts used in the above description show a sequence of steps (processes), the order of steps executed in each embodiment is not limited to the sequence shown in the flowcharts. In each embodiment, the order of steps shown in the diagrams can be changed as long as it does not cause any problems in terms of the content.

[0136] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes.

[0137] 1. An information processing device comprising: a query acquisition means for acquiring a query image including a plurality of objects; a search means for searching for similar images from a plurality of target images based on feature amounts extracted for a plurality of objects included in each of the plurality of target images and the query image, the similar images being images that satisfy predetermined similarity conditions with the query image; and an output control means for outputting the search results to an output means, the feature amounts including depth information relating to the depths of the plurality of objects included in each of the plurality of target images and the query image. 2. The information processing device described in 1., wherein the depth is a value corresponding to the distance from a camera that captured each of the query image and the plurality of target images to each of the objects included in each image. 3. The information processing device described in 1. or 2., wherein the plurality of objects included in the query image include at least one person. 4. The information processing device described in any one of 1. to 4., further comprising extraction means for extracting feature points and the feature amount for each of a plurality of objects included in each of the plurality of target images and the query image, wherein the feature amount includes the depth information for each feature point of the at least one person included in each of the plurality of target images and the query image. 5. The information processing device described in any one of 1. to 4., wherein the depth is a value obtained by statistically processing depths within an object region corresponding to each object included in each of the plurality of target images and the query image, and the object region is a region surrounding the object in a predetermined shape or a region along the outer edge of the object. 6. The information processing device described in any one of 1. to 5., wherein the feature amount further includes at least one of an orientation of each of a plurality of objects included in each of the plurality of target images and the query image and a relationship between the plurality of objects.7. The information processing device described in any one of 1. to 6., wherein the predetermined similarity condition is defined using a similarity between images formed by a combination of each of the plurality of target images and the query image, and the similarity between the images is calculated based on an absolute value of a difference in relative distance between two objects constituting each pair of the plurality of objects included in each of the images, or a value including a sign indicating whether one object constituting each pair of the plurality of objects included in each of the images is nearer or farther from the camera than the other object constituting the pair. 8. An information processing method including one or more computers acquiring a query image including a plurality of objects, searching for similar images from the plurality of target images that satisfy a predetermined similarity condition with the query image based on feature amounts extracted for the plurality of objects included in each of the plurality of target images and the query image, and outputting the search results to an output means, wherein the feature amounts include depth information regarding the depths of the plurality of objects included in each of the plurality of target images and the query image. The information processing method according to 8., wherein the depth is a value corresponding to a distance from a camera that captured each of the query image and the multiple target images to each of the objects included in each of the images. 10. The information processing method according to 8. or 9., wherein the multiple objects included in the query image include at least one person. 11. The information processing method according to 10., further comprising extracting feature points for each of the multiple objects included in each of the multiple target images and the query image and the feature amounts for each of the feature points, wherein the feature amounts include the depth information for each of the feature points of the at least one person included in each of the multiple target images and the query image. 12. The information processing method according to any one of 8. to 11., wherein the depth is a value obtained by statistically processing depths within object regions corresponding to each of the objects included in each of the multiple target images and the query image, wherein the object regions are regions that surround the object in a predetermined shape or regions along the outer edge of the object.13. The information processing method described in any one of 8. to 12., wherein the feature amount further includes at least one of the posture of each of a plurality of objects included in each of the plurality of target images and the query image and the relationship between the plurality of objects. 14. The information processing method described in any one of 8. to 13., wherein the predetermined similarity condition is specified using the similarity between images formed by combining each of the plurality of target images with the query image, and the similarity between images is calculated based on the absolute value of the difference in relative distance between two objects constituting each pair of the plurality of objects included in each of the images, or a value including a sign indicating whether one object constituting each pair is farther or closer to the camera than the other object constituting the pair, for each pair of the plurality of objects included in each of the images. 15. 16. A program causing one or more computers to acquire a query image including a plurality of objects, search for similar images from a plurality of target images that satisfy predetermined similarity conditions with the query image based on feature amounts extracted for a plurality of objects included in each of the plurality of target images and the query image, and output the search results to an output means, wherein the feature amounts include depth information regarding the depths of the plurality of objects included in each of the plurality of target images and the query image. 16. The program described in 15., wherein the depth is a value corresponding to the distance from a camera that captured each of the query image and the plurality of target images to each of the objects included in each image. 17. The program described in 15. or 16., wherein the plurality of objects included in the query image include at least one person. 18. 17. The program described in Item 17, further comprising: extracting, for a plurality of objects included in each of the plurality of target images and the query image, a feature point for each of the objects and the feature amount for each of the feature points, wherein the feature amount includes the depth information for each of the feature points of the at least one person included in each of the plurality of target images and the query image.19. The program described in any one of 15. to 18., wherein the depth is a value obtained by statistically processing depths within an object region corresponding to each object included in each of a plurality of target images and the query image, and the object region is a region surrounding the object in a predetermined shape or a region along the outer edge of the object. 20. The program described in any one of 15. to 19., wherein the feature amount further includes at least one of the orientation of each of a plurality of objects included in each of the plurality of target images and the query image and the relationship between the plurality of objects. 21. The predetermined similarity condition is defined using a similarity between images formed by combining each of the plurality of target images with the query image, and the similarity between images is calculated based on an absolute value of a difference in relative distance between two objects constituting each pair of a plurality of objects included in each of the images, or a value including a sign indicating whether one object constituting each pair of a plurality of objects included in each image is farther or closer to the camera than the other object constituting the pair. 22. A recording medium having the program according to any one of 15. to 21. recorded thereon.

[0138] This application claims priority based on Japanese Patent Application No. 2024-001092, filed January 9, 2024, the disclosure of which is incorporated herein by reference in its entirety.

[0139] 100, 200 Information processing device 110 Query acquisition unit 120 Storage unit 130, 230 Extraction unit 140, 240 Search unit 150 Output control unit 160 Output unit

Claims

1. An information processing apparatus comprising: query acquisition means for acquiring a query image including a plurality of objects; search means for searching, from among the plurality of target images, for a similar image that satisfies a predetermined similarity condition with the query image based on feature amounts extracted for the plurality of objects included in each of the plurality of target images and the query image; and output control means for causing the output means to output the search result, wherein the feature amount includes depth information regarding the depth of the plurality of objects included in each image of the plurality of target images and the query image.

2. The information processing apparatus according to claim 1, wherein the depth is a value corresponding to the distance from the camera that captured each of the query image and the plurality of target images to each of the objects included in each of the images.

3. The information processing apparatus according to claim 1 or 2, wherein the plurality of objects included in the query image includes at least one person.

4. The information processing apparatus according to claim 3, further comprising extraction means for extracting, for the plurality of objects included in each of the plurality of target images and the query image, feature points for each object and the feature amounts for each feature point, wherein the feature amount includes the depth information for each feature point of the at least one person included in each of the plurality of target images and the query image.

5. The information processing apparatus according to claim 1 or 2, wherein the depth is a value obtained by statistically processing the depth within an object region corresponding to each object included in each of the plurality of target images and the query image, and the object region is a region surrounding the object in a predetermined shape or a region along the outer edge of the object.

6. The information processing apparatus according to claim 1 or 2, wherein the feature amount further includes at least one of the posture of each object and the relationship between the plurality of objects for the plurality of objects included in each of the plurality of target images and the query image.

7. The predetermined similarity condition is defined using the similarity between images constituted by combinations of each of the plurality of target images and the query image. The similarity between images is the absolute value of the difference in the relative distances between each pair of two objects constituting each pair among the plurality of objects included in each image, or for each pair of the plurality of objects included in each image, a value including a sign according to whether the other object is far or near from the camera with respect to one object constituting the pair. The information processing apparatus according to claim 1 or 2.

8. An information processing method, comprising: one or more computers acquiring a query image including a plurality of objects; searching, from among the plurality of target images, for a similar image that satisfies a predetermined similarity condition with the query image based on feature amounts extracted for the plurality of objects included in each of the plurality of target images and the query image; and causing an output means to output the search result, wherein the feature amounts include depth information regarding the depth of the plurality of objects included in each of the plurality of target images and the query image.

9. A recording medium recorded with a program, which causes one or more computers to execute: acquiring a query image including a plurality of objects; searching, from among the plurality of target images, for a similar image that satisfies a predetermined similarity condition with the query image based on feature amounts extracted for the plurality of objects included in each of the plurality of target images and the query image; and causing an output means to output the search result, wherein the feature amounts include depth information regarding the depth of the plurality of objects included in each of the plurality of target images and the query image.

Citation Information

Patent Citations

  • Electronic device

    WO2013114732A1

  • Information processing device, information processing method, and program

    WO2021241166A1

  • Image processing system, image processing method, and non-transitory computer-readable medium

    WO2023112321A1