Image processing apparatus, image processing method, and program
The image processing apparatus improves image similarity calculation accuracy by focusing on keypoints within the same cluster, enhancing precision in image matching tasks.
Patent Information
- Application Number
- JP2023528889
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-06-17
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2041-06-17
AI Technical Summary
Existing techniques for calculating the similarity between two images, while improving accuracy, still face challenges in achieving high precision due to the complexity of associating keypoints from different subjects within the images.
An image processing apparatus and method that perform an extraction process for feature amounts, a detection process for keypoints, and an estimation process for pixel clusters. The similarity calculation is then focused on keypoints within the same estimated cluster, avoiding association between keypoints from different clusters.
This approach enhances the accuracy of similarity calculation between two images by ensuring that only keypoints from the same cluster are associated, thereby improving the precision of image similarity assessment.
Smart Images

Figure 0007694658000005 
Figure 0007694658000006 
Figure 0007694658000007
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.
Background Art
[0002] A technique related to the present invention is disclosed in Non-Patent Document 1. Non-Patent Document 1 discloses a technique (R2D2: Repeatable and Reliable Detector and Descriptor) that generates a feature amount map, a map related to repeatability (a repeatability map), and a map related to reliability (a reliability map) based on an image, and detects characteristic portions (keypoints) of the appearance of a subject included in the image with high accuracy based on them.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] A technique for calculating the similarity between two images with high accuracy is desired. By performing image matching using the keypoints detected using the technique described in Non-Patent Document 1, the accuracy of calculating the similarity between two images is improved. However, further improvement in the accuracy is expected.
[0005] An object of the present invention is to provide a new technique for calculating the similarity between two images with high accuracy.
Means for Solving the Problems
[0006] According to the present invention, image processing means for performing an extraction process for extracting feature amounts from an image and an estimation process for estimating a cluster to which each pixel belongs, similarity calculation means for calculating the similarity between two images by calculating the similarity of the feature amounts between pixels estimated to belong to the same cluster, an image processing apparatus having the above is provided.
[0007] Also, according to the present invention, a computer performs an image processing step of performing an extraction process for extracting feature amounts from an image and an estimation process for estimating a cluster to which each pixel belongs, and a similarity calculation step of calculating the similarity between two images by calculating the similarity of the feature amounts between pixels estimated to belong to the same cluster, an image processing method for executing the above is provided.
[0008] Also, according to the present invention, a computer is made to function as image processing means for performing an extraction process for extracting feature amounts from an image and an estimation process for estimating a cluster to which each pixel belongs, and similarity calculation means for calculating the similarity between two images by calculating the similarity of the feature amounts between pixels estimated to belong to the same cluster, a program for causing the above to function is provided.
Advantages of the Invention
[0009] According to the present invention, a new technique for calculating the similarity between two images with high accuracy is realized.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.
[0012] <First Embodiment> 「Overview」 An image may contain a plurality of subjects. The subjects included vary depending on the shooting location, shooting timing, etc. For example, various objects such as roads, plants, buildings, people, automobiles, buses, the sky, etc. can be subjects. If keypoint matching (feature point matching) between the first image and the second image is performed without considering this point, there may be a problem of associating a keypoint detected from a first subject in the first image with a keypoint detected from a second subject (a subject different from the first subject) in the second image. As a result, the calculation accuracy of the similarity between the two images decreases.
[0013] The image processing apparatus of the present embodiment has a feature for reducing such inconvenience. Specifically, the image processing apparatus of the present embodiment performs an extraction process for extracting feature amounts, a detection process for detecting keypoints, and an estimation process for estimating a cluster to which each pixel belongs on the image to be processed. Then, the image processing apparatus calculates the similarity of the feature amounts between the keypoints estimated to belong to the same cluster, and performs association of the keypoints based on the calculation result. The image processing apparatus does not calculate the similarity of the feature amounts and perform association between the keypoints estimated to belong to different clusters. In this way, only the association between the keypoints estimated to belong to the same cluster is realized, and the association between the keypoints estimated to belong to different clusters can be avoided. As a result, the calculation accuracy of the similarity between the two images is improved.
[0014] "Configuration" Next, the configuration of the image processing apparatus will be described. First, an example of the hardware configuration of the image processing apparatus will be described. Each functional unit of the image processing apparatus is realized by an arbitrary combination of hardware and software centered around the CPU (Central Processing Unit), memory, program loaded into the memory, storage unit such as a hard disk storing the program (in addition to the program stored in advance at the stage of shipping the apparatus, programs downloaded from storage media such as CDs (Compact Discs) and servers on the Internet can also be stored), and network connection interface. And it is understood by those skilled in the art that there are various modifications to the realization method and apparatus.
[0015] FIG. 1 is a block diagram illustrating the hardware configuration of the image processing apparatus. As shown in FIG. 1, the image processing apparatus includes a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The image processing apparatus may not have the peripheral circuit 4A. Note that the image processing apparatus may be composed of a plurality of physically and / or logically separated apparatuses, or may be composed of one physically and / or logically integrated apparatus. When the image processing apparatus is composed of a plurality of physically and / or logically separated apparatuses, each of the plurality of apparatuses can have the above-described hardware configuration.
[0016] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data from each other. The processor 1A is an arithmetic processing device such as a CPU or GPU (Graphics Processing Unit). The memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. The input device is, for example, a keyboard, a mouse, a microphone, a physical button, a touch panel, etc. The output device is, for example, a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform operations based on their operation results.
[0017] Next, the functional configuration of the image processing apparatus will be described. FIG. 2 shows an example of a functional block diagram of the image processing apparatus 10 according to the present embodiment. As shown in the figure, the image processing apparatus 10 includes an acquisition unit 11, an image processing unit 12, and a similarity calculation unit 13.
[0018] The acquisition unit 11 acquires two images. These two images are the objects for calculating the similarity between each other. For example, the acquisition unit 11 may acquire two images specified by a user input, or may acquire two images selected based on a predetermined rule from the images stored in the storage unit (database). Further, the acquisition unit 11 may acquire one image specified by a user input and one image selected based on a predetermined rule from the images stored in the database.
[0019] The image processing unit 12 performs extraction processing, detection processing, and estimation processing on the image acquired by the acquisition unit 11. Note that these processes may be executed in advance on the images stored in the database, and the results of these processes may be stored in the storage unit in association with each image. In this case, it is not necessary to execute the extraction processing, detection processing, and estimation processing again on the images acquired from the database.
[0020] The extraction processing is a process of extracting the feature amounts of an image. For example, as shown in FIG. 3, when an image is input to a learned estimation model, the feature amounts of the image are extracted and data of a feature amount group is created. The data of the feature amount group indicates the feature amounts of the respective pixels. In the case of the illustrated example, the feature amount of each pixel is indicated by C-dimensional data. The estimation model is, for example, a CNN (Convolutional Neural Network), but is not limited thereto. Generation of the data of the feature amount group can be realized using any conventional technique.
[0021] The detection processing is a process of detecting key points from an image. In the present embodiment, it is assumed that the technique described in Non-Patent Document 1 is used to detect key points, but other methods may be adopted to detect key points. non- A detailed description of the technique described in Patent Document 1 is omitted here. When the technique described in Non-Patent Document 1 is used, as shown in FIG. 3, when an image is input to a learned estimation model, a Repeatability map is created. The Repeatability map indicates the weighted value of each pixel. The image processing unit 12 can detect key points using such a Repeatability map.
[0022] The estimation process is a process of estimating the cluster to which each pixel belongs. In the estimation process, the image is divided into a plurality of clusters. Each cluster corresponds to each type of subject. For example, there is one cluster corresponding to a road and one cluster corresponding to a plant. That is, the process of dividing the image into a plurality of clusters is a process of dividing the image into a plurality of areas for each of the plurality of subjects. When an image is input to the learned estimation model as shown in FIG. 3, a Segmentation map is created. The Segmentation map shows the result of dividing the above image into a plurality of clusters, that is, the cluster to which each pixel belongs.
[0023] In this embodiment, a well-known Segmentation technique is used to create a Segmentation map. Examples of well-known Segmentation techniques include, for example, Semantic Segmentation, Instance Segmentation, Panoptic Segmentation, etc. In this embodiment, for example, when focusing on a certain pixel, a teacherless Segmentation method that utilizes the fact that adjacent pixels have a stronger correlation and distant pixels have a weaker correlation is used to create a Segmentation map. Based on such a Segmentation map, although the cluster (cluster identification information) to which each pixel belongs can be specified, the type of subject indicated by each pixel cannot be specified.
[0024] Note that an example of the learning method of the above-described estimation model will be described in the fourth embodiment.
[0025] Returning to FIG. 2, the similarity calculation unit 13 calculates the similarity between two images. First, the similarity calculation unit 13 calculates the similarity of feature amounts between pixels estimated to belong to the same cluster, and determines a combination of pixels to be associated with each other based on the calculation result.
[0026] Specifically, the similarity calculation unit 13 calculates the similarity between a first keypoint, which is a keypoint (pixel) detected from the first image, and a second keypoint, which is a keypoint detected from the second image, and determines a combination of the first keypoint and the second keypoint that are associated with each other based on the calculated similarity. In this process, the similarity calculation unit 13 calculates the similarity between the first keypoint and the second keypoint that are estimated to belong to the same cluster, and determines a combination of the first keypoint and the second keypoint that are associated with each other based on the calculated similarity. Note that the similarity calculation unit 13 does not calculate the similarity between the first keypoint and the second keypoint that are estimated to belong to different clusters. Therefore, the first keypoint and the second keypoint that are estimated to belong to different clusters are not associated with each other.
[0027] This process will be described with reference to FIG. 4. First, assume that from each of the first image and the second image, data of a Segmentation map and a feature amount group as shown in the figure are created. In the shown Segmentation map, it is shown to which cluster among clusters 1, 2, 3,... each pixel belongs. The similarity calculation unit 13 calculates the similarity between a first keypoint estimated to belong to cluster 1 among the keypoints detected from the first image and a second keypoint estimated to belong to cluster 1 among the keypoints detected from the second image, and determines a combination of the first keypoint and the second keypoint that are associated with each other based on the calculated similarity. Similarly, the similarity calculation unit 13 calculates the similarity between a first keypoint estimated to belong to cluster 2 among the keypoints detected from the first image and a second keypoint estimated to belong to cluster 2 among the keypoints detected from the second image, and determines a combination of the first keypoint and the second keypoint that are associated with each other based on the calculated similarity. In such a process, the first keypoint estimated to belong to cluster 1 can be associated only with the second keypoint estimated to belong to cluster 1. The first keypoint estimated to belong to cluster 1 will not be associated with the second keypoint estimated to belong to other clusters.
[0028] Note that the method for calculating the similarity between keypoints and the method for determining the keypoints to be associated with each other based on the calculated similarity between keypoints can be realized by adopting any conventional technique.
[0029] After the similarity calculation unit 13 determines the combination of pixels (combination of keypoints) to be associated with each other, it calculates the similarity between the two images based on the result of the association. The method for calculating the similarity between the two images based on the result of the association can be realized by adopting any conventional technique.
[0030] Next, an example of the processing flow of the image processing apparatus 10 will be described with reference to the flowchart of FIG. 5.
[0031] First, the image processing apparatus 10 acquires two images for which the degree of similarity to each other is to be calculated (S10).
[0032] Next, the image processing apparatus 10 executes, for each image, an extraction process for extracting feature amounts, a detection process for detecting keypoints, and an estimation process for estimating the cluster to which each pixel belongs (S11). Note that if the extraction process, the detection process, and the estimation process have been executed in advance for the images acquired in S10 and the results are stored in the database, the image processing apparatus 10 may acquire the results from the database, and there is no need to execute the extraction process, the detection process, and the estimation process again for those images.
[0033] Next, the image processing apparatus 10 calculates the degree of similarity of the feature amounts between the pixels estimated to belong to the same cluster. Then, the image processing apparatus 10 determines a combination of keypoints to be associated with each other based on the calculated result, and calculates the degree of similarity of the two images based on the result of the association (S12).
[0034] "Operation and Effect" As described above, the image processing apparatus 10 of the present embodiment performs an extraction process for extracting feature amounts, a detection process for detecting keypoints, and an estimation process for estimating the cluster to which each pixel belongs on the image to be processed. Then, the image processing apparatus 10 calculates the degree of similarity of the feature amounts between the keypoints estimated to belong to the same cluster, and performs association of the keypoints based on the calculation result. The image processing apparatus 10 does not calculate the degree of similarity of the feature amounts and perform association between the keypoints estimated to belong to different clusters. In this way, only the association between the keypoints estimated to belong to the same cluster is realized, and the association between the keypoints estimated to belong to different clusters can be avoided. As a result, the calculation accuracy of the degree of similarity of the two images is improved.
[0035] <Second Embodiment> In this embodiment, based on user input, a plurality of clusters are classified into a reference cluster and a non-reference cluster. Then, the image processing apparatus 10 calculates the similarity of feature amounts between the key points estimated to belong to the same cluster described in the first embodiment using the key points estimated to belong to the reference cluster and without using the key points estimated to belong to the non-reference cluster, and performs association of the key points based on the result, thereby calculating the similarity of the two images.
[0036] An example of the functional block diagram of the image processing apparatus 10 of this embodiment is shown in FIG. 2, similar to the first embodiment.
[0037] The similarity calculation unit 13 calculates the similarity between a first key point, which is a key point (pixel) detected from the first image, and a second key point, which is a key point detected from the second image, and determines a combination of the first key point and the second key point that are associated with each other based on the calculated similarity. In this process, the similarity calculation unit 13 uses only the key points estimated to belong to the reference cluster and does not use the key points estimated to belong to the non-reference cluster.
[0038] That is, the similarity calculation unit 13 calculates the similarity between the first key point and the second key point estimated to belong to the same cluster using only the key points estimated to belong to the reference cluster, and determines a combination of the first key point and the second key point that are associated with each other based on the calculated similarity. Note that the similarity calculation unit 13 does not use the key points estimated to belong to the non-reference cluster for this process. Therefore, the key points estimated to belong to the non-reference cluster are not associated with any key points. Also, similar to the first embodiment, the similarity calculation unit 13 does not calculate the similarity between the first key point and the second key point estimated to belong to different clusters. Therefore, the first key point and the second key point estimated to belong to different clusters are not associated with each other.
[0039] Here, an example of a method for classifying clusters into reference clusters and non-reference clusters will be described. The user makes an input to determine whether each of a plurality of clusters (cluster 1, 2, 3, ···) is a reference cluster or a non-reference cluster. The user may make an input for such classification for each combination of images for which similarity is calculated. Alternatively, the content input by the user for such classification may be stored in the image processing apparatus 10, and the classification content may be applied to a plurality of combinations of images. The similarity calculation unit 13 specifies, based on the content of the user input, whether each of the plurality of clusters is classified as a reference cluster or a non-reference cluster.
[0040] Here, an example of an interface screen for receiving the above user input will be described. For example, the image processing apparatus 10 outputs an interface screen that displays a Segmentation map (a Segmentation map created from the first image or a Segmentation map created from the second image) as shown in FIG. 4. The Segmentation map shown in FIG. 4 shows the boundaries of a plurality of clusters using techniques such as contour lines and color separation, and displays the plurality of clusters so that they can be distinguished from each other. The interface screen is configured to receive a user input for specifying whether each of the plurality of clusters shown in the Segmentation map is a reference cluster or a non-reference cluster.
[0041] The user can estimate the type of subject indicated by each cluster based on, for example, the shape of each of the plurality of clusters shown in the Segmentation map (the shape composed of pixels belonging to each cluster). As another example, the image processing apparatus 10 may display, on the interface screen, the image from which the Segmentation map shown in FIG. 4 is derived, together with the Segmentation map. Then, the user may identify the type of subject indicated by each of the plurality of clusters shown in the Segmentation map by comparing the Segmentation map with the image from which the Segmentation map is derived.
[0042] Note that which cluster is to be used as the reference cluster can be freely determined in consideration of the usage scenario of the image processing apparatus 10 or the like. For example, as described in the third embodiment, when specifying the position (the photographed position) indicated by an image using the calculation result of the similarity of images, it is preferable to use, as the reference cluster, a subject whose existing position, such as a building or a road, is fixed, and to use, as non-reference clusters, subjects whose existing positions, such as a person or a car, fluctuate.
[0043] Next, an example of the processing flow of the image processing apparatus 10 will be described using the flowchart of FIG. 6.
[0044] First, the image processing apparatus 10 acquires two images for which the mutual similarity is to be calculated (S20).
[0045] Next, the image processing apparatus 10 executes, for each image, an extraction process for extracting feature amounts, a detection process for detecting key points, and an estimation process for estimating the cluster to which each pixel belongs (S21). Note that if the extraction process, the detection process, and the estimation process have been executed in advance for the images acquired in S20 and the results are stored in the database, the image processing apparatus 10 may acquire the results from the database, and there is no need to execute the extraction process, the detection process, and the estimation process again for those images.
[0046] Next, the image processing apparatus 10 calculates the similarity of the two images by calculating the similarity of the feature amounts between the key points estimated to belong to the same cluster, using the key points estimated to belong to the reference cluster and without using the key points estimated to belong to the non-reference clusters (S22).
[0047] Other configurations of the image processing apparatus 10 of the present embodiment are the same as those of the first embodiment.
[0048] According to the image processing apparatus 10 of the present embodiment, the same operational effects as those of the image processing apparatus 10 of the first embodiment are achieved. Further, according to the image processing apparatus 10 of the present embodiment, since the similarity of an image can be calculated using only appropriate clusters, the calculation accuracy of the similarity is improved.
[0049] <Third Embodiment> The image processing apparatus 10 of the present embodiment has a function of calculating the similarity between a processing target image and each of a plurality of reference images associated with position information, and outputting position information related to the processing target image based on the calculation result. This will be described in detail below.
[0050] FIG. 7 shows an example of a functional block diagram of the image processing apparatus 10 of the present embodiment. As shown in the figure, the image processing apparatus 10 includes an acquisition unit 11, an image processing unit 12, a similarity calculation unit 13, and a result output unit 14.
[0051] The acquisition unit 11 acquires a processing target image. The acquisition unit 11 acquires a processing target image specified / selected / determined, etc. by an input from the user. An image for which position information of the photographed position (the position of the camera when the image was photographed) is required is acquired as the processing target image. For example, an image that is not geotagged and whose photographed position is unknown is acquired as the processing target image.
[0052] The image processing unit 12 performs an extraction process for extracting feature amounts, a detection process for detecting keypoints, and an estimation process for estimating the cluster to which each pixel belongs on the processing target image. The details of each process are the same as those of the first and second embodiments.
[0053] The similarity calculation unit 13 calculates the similarity between the image to be processed and each of the plurality of reference images stored in the database. The image processing apparatus 10 may include the database, or an external apparatus configured to be communicable with the image processing apparatus 10 may include the database. FIG. 8 schematically shows an example of information stored in the database. In the illustrated example, position information is registered in association with each of the plurality of reference images. The position information represents the position indicated by each reference image. The position information may indicate a relatively narrow area using latitude / longitude, an address, etc., or may indicate a relatively wide area using a country name, a prefecture name, a city / town / village name, etc.
[0054] Details of the process of calculating the similarity are the same as those in the first and second embodiments. When adopting the second embodiment, it is preferable to use a reference cluster for a subject whose existing position is fixed, such as a building or a road, and a non-reference cluster for a subject whose existing position fluctuates, such as a person or a car.
[0055] The result output unit 14 outputs the position information associated with the reference image whose similarity to the image to be processed is equal to or greater than the threshold value as the position information related to the image to be processed, that is, the position information indicated by the image to be processed.
[0056] Next, an example of the processing flow of the image processing apparatus 10 will be described using the flowchart of FIG. 9.
[0057] First, the image processing apparatus 10 acquires the image to be processed (S30).
[0058] Next, the image processing apparatus 10 executes an extraction process for extracting feature amounts, a detection process for detecting keypoints, and an estimation process for estimating the cluster to which each pixel belongs on the image to be processed (S31).
[0059] Next, the image processing apparatus 10 calculates the similarity between the image to be processed and each of the plurality of reference images stored in the database (S32).
[0060] When there is a reference image whose similarity to the image to be processed is equal to or greater than the reference value (Yes in S33), the image processing apparatus 10 outputs the position information associated with the reference image as the position information related to the image to be processed, that is, the position information indicated by the image to be processed (S34).
[0061] On the other hand, when there is no reference image whose similarity to the image to be processed is equal to or greater than the reference value (No in S33), the image processing apparatus 10 outputs that the position information indicated by the image to be processed is unknown (S35).
[0062] Other configurations of the image processing apparatus 10 in the present embodiment are the same as those in the first and second embodiments.
[0063] According to the image processing apparatus 10 of the present embodiment, the same operational effects as those of the image processing apparatuses 10 of the first and second embodiments are achieved. Further, according to the image processing apparatus 10 of the present embodiment, since the similarity of the images can be specified with high accuracy, the position indicated by the image can be specified with high accuracy by using the result.
[0064] <Fourth Embodiment> In the present embodiment, the extraction process for extracting features, the detection process for detecting key points, and the estimation model used in the estimation process for estimating the cluster to which each pixel belongs are learned by a characteristic method.
[0065] First, pairs of images including the same subject are used as teacher data. The pairs of images may be pairs of different images generated by photographing the same subject at different timings. In this case, the pairs of different images may have different shooting angles, distances to the subject, lighting conditions, etc., or may be the same. As another example, by performing image processing such as color tone change on a certain one image, a pair of images including the same subject (a pair of the image before editing and the image after editing) may be created.
[0066] The learning device that learns the estimation model inputs each of the pair of images A and B into the estimation model, and executes extraction processing, detection processing, and estimation processing on images A and B. As a result, corresponding to each of images A and B, result objects (data of a feature amount group, Repeatability map, K-dimensional data group, Segmentation map, etc.) as shown in, for example, FIG. 3 are obtained. Then, the learning device optimizes various parameters of the estimation model based on at least one of these result objects and a characteristic loss function. That is, the learning device optimizes various parameters of the estimation model so as to minimize the loss function.
[0067] The loss function L is defined as in the following formula (1). The loss function L is the loss function L seg and the loss function L rep and is created based on. In formula (1), the loss function L seg and the loss function L rep and the added value is the loss function L.
[0068]
Equation
[0069] The loss function L rep is defined as in the following formula (2). The loss function L rep is a loss function regarding the reproducibility of feature amounts. Since the details are as disclosed in Non-Patent Document 1, the description here is omitted.
[0070]
Equation
[0071] The loss function L seg is a loss function regarding the correlation between pixels. The loss function L seg is a statistical value (average value, etc.) of the values calculated for each pixel based on the function L seg,u . The function L seg,u is defined as in the following formula (3).
[0072] [Mathematics]
[0073] Function F u is the K - dimensional data of pixel u(=(i, j)) of image A as shown in FIG. 10.
[0074] Function F' g(u)+t is the K - dimensional data of pixel {g(u)+t} of image B as shown in FIG. 11. Pixel {g(u)+t} of image B is the pixel displaced by displacement amount t from pixel g(u) of image B. Pixel g(u) of image B is the pixel corresponding to pixel u of image A. Corresponding pixels to each other indicate the same part of the same subject.
[0075] Function T is a set of predefined displacement amounts t.
[0076] Function I is defined as in the following formula (4). Function H is an entropy function.
[0077] [Mathematics]
[0078] The image processing unit 12 of the image processing apparatus 10 according to the present embodiment executes an extraction process for extracting feature amounts, a detection process for detecting keypoints, and an estimation process for estimating the cluster to which each pixel belongs based on an estimation model learned by the above - described characteristic method. Other configurations of the image processing apparatus 10 according to the present embodiment are the same as those of the first to third embodiments.
[0079] As described above, according to the image processing apparatus 10 of the present embodiment, the same operational effects as those of the first to third embodiments are achieved. Further, according to the image processing apparatus 10 of the present embodiment, an extraction process for extracting feature amounts, a detection process for detecting keypoints, and an estimation process for estimating a cluster to which each pixel belongs are executed based on an estimation model learned by a characteristic method. Therefore, the accuracy of these processes is improved.
[0080] In this specification, "acquisition" means, based on user input or based on a program instruction, "the self-device goes to obtain data stored in another device or a storage medium (active acquisition)", for example, requesting or inquiring of another device and receiving, accessing another device or a storage medium and reading out, etc., and, based on user input or based on a program instruction, "the self-device inputs data output from another device (passive acquisition)", for example, receiving data distributed (or transmitted, push-notified, etc.), and also selecting and acquiring from the received data or information, and "generating new data by editing (textualizing, rearranging data, extracting some data, changing file format, etc.) the data and acquiring the new data", including at least any one of them.
[0081] Some or all of the above-described embodiments may be described as follows in the appended claims, but are not limited thereto. 1. Image processing means for performing an extraction process for extracting feature amounts from an image and an estimation process for estimating a cluster to which each pixel belongs; Similarity calculation means for calculating the similarity between two images by calculating the similarity of the feature amounts between pixels estimated to belong to the same cluster; An image processing apparatus having the above. 2. The plurality of clusters are classified into a reference cluster and a non-reference cluster, The similarity calculation means calculates the similarity of the feature amounts using pixels estimated to belong to the reference cluster and without using pixels estimated to belong to the non-reference cluster. The image processing apparatus according to 1. 3. The image processing apparatus according to claim 2, wherein, based on a user input, the plurality of clusters are classified into the reference cluster and the non-reference cluster. 4. The similarity calculation means determines pixels serving as keypoints, and calculates the similarity between the two images by calculating the similarity of the feature amounts between the pixels determined as the keypoints. The image processing apparatus according to any one of claims 1 to 3. 5. The similarity calculation means calculates the similarity between the image to be processed and each of a plurality of reference images associated with position information, and further includes result output means for outputting the position information associated with the reference image having a similarity with the image to be processed equal to or greater than a threshold value as the position information related to the image to be processed. The image processing apparatus according to any one of claims 1 to 4. 6. The image processing means performs the extraction process and the estimation process based on an estimation model learned based on a loss function created based on a loss function related to the correlation between pixels and a loss function related to the reproducibility of feature amounts. The image processing apparatus according to any one of claims 1 to 5. 7. A computer performs an image processing step of performing an extraction process for extracting a feature amount and an estimation process for estimating a cluster to which each pixel belongs on an image, and a similarity calculation step of calculating the similarity between the two images by calculating the similarity of the feature amounts between the pixels estimated to belong to the same cluster. An image processing method for executing. 8. A computer is caused to function as image processing means for performing an extraction process for extracting a feature amount and an estimation process for estimating a cluster to which each pixel belongs on an image, and similarity calculation means for calculating the similarity between the two images by calculating the similarity of the feature amounts between the pixels estimated to belong to the same cluster. A program for functioning as.
Description of Signs
[0082] 10 Image processing apparatus 11 Acquisition unit 12 Image processing unit 13 Similarity calculation unit 14 Result output unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus
Claims
1. Image processing means for performing an extraction process for extracting feature amounts from an image, creating a map indicating weight values of respective pixels, performing a detection process for detecting pixels to be keypoints based on the map, and performing an estimation process for estimating clusters to which the respective pixels belong; Similarity calculation means for calculating a similarity between a first keypoint among the keypoints detected from a first image and a second keypoint that belongs to the same cluster as the first keypoint among the keypoints detected from a second image, determining a combination of the first keypoint and the second keypoint to be associated with each other based on the calculated similarity, and calculating a similarity between the first image and the second image based on a result of the association between the first keypoint and the second keypoint; An image processing apparatus having the above.
2. The plurality of clusters are classified into a reference cluster and a non-reference cluster based on a user input for designating which of the plurality of clusters is to be a reference cluster and which is to be a non-reference cluster. The similarity calculation means uses, as the first keypoint, a keypoint belonging to the reference cluster among the keypoints detected from the first image, and does not use a keypoint belonging to the non-reference cluster as the first keypoint, calculates a similarity between the first keypoint and the second keypoint, and determines a combination of the first keypoint and the second keypoint to be associated with each other. The image processing apparatus according to claim 1.
3. The similarity calculation means calculates a similarity between a processing target image and each of a plurality of reference images associated with position information indicating a shooting position. The image processing apparatus according to claim 1 or 2, further comprising result output means for outputting, as information indicating a shooting position of the processing target image, the position information associated with the reference image having a similarity with the processing target image equal to or greater than a threshold value.
4. A computer, An image processing step that performs an extraction process for extracting feature amounts from an image, creates a map indicating weighted values of respective pixels, performs a detection process for detecting pixels to be keypoints based on the map, and performs an estimation process for estimating clusters to which the respective pixels belong, and a similarity calculation step that calculates a similarity between a first keypoint among the keypoints detected from a first image and a second keypoint that belongs to the same cluster as the first keypoint among the keypoints detected from a second image, determines a combination of the first keypoint and the second keypoint to be associated with each other based on the calculated similarity, and calculates a similarity between the first image and the second image based on a result of the association between the first keypoint and the second keypoint, An image processing method for executing the above steps.
5. The plurality of clusters are classified into a reference cluster and a non-reference cluster based on a user input that designates whether each of the plurality of clusters is to be a reference cluster or a non-reference cluster, In the similarity calculation step, the keypoints belonging to the reference cluster among the keypoints detected from the first image are used as the first keypoints, and the similarity between the first keypoints and the second keypoints is calculated and the combination of the first keypoints and the second keypoints to be associated with each other is determined without using the keypoints belonging to the non-reference cluster as the first keypoints. The image processing method according to claim 4.
6. A computer, image processing means for performing an extraction process for extracting feature amounts from an image, creating a map indicating weighted values of respective pixels, performing a detection process for detecting pixels to be keypoints based on the map, and performing an estimation process for estimating clusters to which the respective pixels belong, and Calculate the similarity between a first keypoint among the keypoints detected from the first image and a second keypoint belonging to the same cluster as the first keypoint among the keypoints detected from the second image, determine a combination of the first keypoint and the second keypoint that are associated with each other based on the calculated similarity, and calculate the similarity between the first image and the second image based on the result of the association between the first keypoint and the second keypoint. A similarity calculation means, A program that functions as
7. The plurality of clusters are classified into the reference cluster and the non-reference cluster based on a user input that designates whether each of the plurality of clusters is to be a reference cluster or a non-reference cluster, The similarity calculation means uses the keypoints belonging to the reference cluster among the keypoints detected from the first image as the first keypoints, and does not use the keypoints belonging to the non-reference cluster as the first keypoints. The program according to claim 6, which calculates the similarity between the first keypoints and the second keypoints and determines a combination of the first keypoints and the second keypoints that are associated with each other.
Citation Information
Patent Citations
Mobile unit information processor, mobile unit information processing method, and program
JP2009074995A
Surveillance image retrieval apparatus and surveillance system
JP2011029737A
Image recognition dive, method and program
JP2018194956A
Image processing device, image processing method, program, integrated circuit
WO2013031096A1