Feature extraction device, image matching device and computer program

The feature extraction device addresses the challenge of identifying products in complex backgrounds by aligning features across multiple scales and dimensions, enhancing foreground features, and improving classification accuracy for precise product matching.

JP2025106799APending Publication Date: 2025-07-16FUJITSU LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2024216063
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-04
Filing Date
2024-12-11
Publication Date
2025-07-16

AI Technical Summary

Technical Problem

Existing image matching technologies struggle to accurately identify products in complex backgrounds with similar appearances and varying angles, especially in large supermarkets, due to challenges in extracting relevant features from product images.

Method used

A feature extraction device that aligns features in multiple scales and dimensions, uses a classifier module to minimize prediction differences, and enhances foreground features by increasing background classes, improving accuracy in identifying product categories.

Benefits of technology

Enhances the accuracy of identifying product categories by enriching image information and refining prediction results through feature alignment and classification, enabling precise matching in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106799000001_ABST
    Figure 2025106799000001_ABST
Patent Text Reader

Abstract

To provide a feature extraction device, an image matching device and a program.SOLUTION: A feature extraction device includes a module which: extracts a feature having a plurality of scales and a plurality of feature dimensions from a matching standby image; aligns all or a part of the extracted features in the feature dimensions, aligns the features aligned in the feature dimensions two by two in order of "from small to large" and "from large to small" in a scale dimension, and thereby generates a fusion feature; includes a plurality of classifiers, allocated to the corresponding fusion features, and trains the classifiers so that a difference between a prediction result of each of the fusion features aligned in the order of "from small to large" in the scale dimension, and a prediction result of all of the fusion features aligned in the order of "from large to small" is minimized; and divides the fusion features into a foreground feature and a background feature on the basis of a predetermined rule, and increases one class to the background feature, for other classifiers during training.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural networks, and in particular, to a feature extraction technique for fine-grained verification.

Background Art

[0002] In recent years, techniques for understanding the interests of shoppers' product selections based on images or image frames within videos have received wide attention. This typically requires matching a queried product image with corresponding images in an image library.

[0003] In large supermarkets, there are often hundreds or thousands of products. Products in several different categories are very similar. Even products in the same category may have quite different appearances when viewed from different angles. Also, the backgrounds of product images are usually relatively complex and may include customers' hands, shopping carts, display shelves, etc. This makes it even more difficult to obtain features for identifying (recognizing) products.

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present invention is to provide a feature extraction device, an image matching device, and a computer program in order to solve the above-mentioned problems existing in the prior art.

Means for Solving the Problems

[0005] According to one aspect of the present invention, there is provided a device for extracting features, which is configured to extract features having a plurality of scales (scales) and a plurality of feature dimensions respectively from an image waiting for matching; A feature fusion module configured to generate a fused feature by aligning all or some of the extracted features in a feature dimension and then aligning the features aligned in the feature dimension two by two in ascending and descending order in a scale dimension, where the fused feature includes the aligned features and the features before alignment; A classifier module including a plurality of classifiers, each of the plurality of classifiers being assigned to a corresponding fused feature and trained to minimize the difference between the prediction result of each fused feature aligned in ascending order in the scale dimension and the prediction results of all fused features aligned in descending order in the scale dimension, or to minimize the difference between the prediction result of each fused feature aligned in descending order in the scale dimension and the prediction results of all fused features aligned in ascending order in the scale dimension; and A foreground enhancement module configured to divide the fused feature into a foreground feature and a background feature based on a predetermined rule and, during training, increase the number of classes for the background feature by one compared to the foreground feature for other classifiers, where the foreground feature is a feature associated with the category (identification category / category to be identified) of the object in the image waiting for matching. The foreground enhancement module is included.

[0006] Preferably, the feature extraction module includes N backbone blocks, the plurality of scales includes N scales, and each backbone block is configured to extract a feature having one of the N scales, where N is an integer greater than or equal to 2.

[0007] Preferably, the plurality of classifiers includes N×2 classifiers, the fused feature is divided into N×2 feature sets based on the scale, and the N×2 classifiers are respectively assigned to the N×2 feature sets.

[0008] Preferably, the plurality of classifiers can be realized by a fully connected layer.

[0009] Preferably, the prediction result is a probability distribution in which the object belongs to each of two or more categories.

[0010] Preferably, the predetermined rule is to divide the fusion feature into a foreground feature and a background feature based on the sorting in the correct (ground truth) category of the probability distributions output from all the classifiers.

[0011] Preferably, the plurality of classifiers are trained based on a first loss function and a second loss function, and the other classifiers in the foreground enhancement module are trained based on the second loss function.

[0012] Preferably, the first loss function is a Kullback-Leibler divergence loss function, and the second loss function is a cross-entropy loss function.

[0013] Preferably, the total loss function is a weighted sum of the first loss function and the second loss function, and the training includes minimizing the total loss function.

[0014] Preferably, aligning the features aligned in the feature dimension two by two in the order of "small to large" and "large to small" in the scale dimension is performed in parallel or sequentially.

[0015] Preferably, the feature dimension is the length of the extracted feature, and the scale is obtained based on the resolution of the image waiting for matching.

[0016] Also, according to another aspect of the present invention, an image matching device is provided, which is a feature extraction module configured to extract features having a plurality of scales and a plurality of feature dimensions from an image waiting for matching, wherein the extracted features are related to an object in the image waiting for matching; A feature fusion module configured to generate a fused feature by aligning all or some of the extracted features in feature dimensions, and then aligning the features aligned in feature dimensions two by two in the order of "from small to large" and "from large to small" in scale dimensions, where the fused feature includes the aligned features and the features before alignment; and including a matching module that stitches all or some of the fused features, calculates the distance between the stitched features and all the images in the image library, and rearranges all the images in the image library based on the calculated result; The network parameters of the feature extraction module and the feature fusion module are set by training the device for extracting the above-mentioned features.

[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing a program, which, when executed by a processor, causes a computer to execute a method for extracting features, the method including: extracting features having a plurality of scales and a plurality of feature dimensions from an image to be matched respectively; generating a fused feature by aligning all or some of the extracted features in feature dimensions, and then aligning the features aligned in feature dimensions two by two in the order of "from small to large" and "from large to small" in scale dimensions, where the fused feature includes the aligned features and the features before alignment; training a plurality of classifiers, each of the plurality of classifiers being assigned to a corresponding fused feature and minimizing the difference between the prediction result of each fused feature aligned in the order of "from small to large" in scale dimension and the prediction results of all the fused features aligned in the order of "from large to small", or minimizing the difference between the prediction result of each fused feature aligned in the order of "from large to small" and the prediction results of all the fused features aligned in the order of "from small to large"; and Based on certain rules, the fused features are divided into foreground features and background features, and during training, for other classifiers, one class is increased for the background features compared to the foreground features, where the foreground features are the features associated with the category of the object in the image.

[0018] According to another aspect of the present invention, there is further provided a computer-readable storage medium storing a program, and when the program is executed by a processor, it causes the computer to execute an image matching method, and the image matching method includes: extracting features having a plurality of scales and a plurality of feature dimensions respectively from the image to be matched, and the extracted features are related to the object in the image to be matched; aligning all or some of the extracted features in the feature dimension, and generating fused features by aligning the features aligned in the feature dimension two by two in the order of "from small to large" and "from large to small" in the scale dimension, where the fused features include the aligned features and the features before alignment; and including stitching all or some of the fused features, calculating the distances between the stitched features and all the images in the image library, and sorting all the images in the image library based on the calculated results. The extraction and fusion of features are executed by the feature extraction module and the feature fusion module in the above-mentioned image matching device.

[0019] According to another aspect of the present invention, corresponding computer program code and computer program products are also provided.

Advantages of the Invention

[0020] The feature extraction device and the image matching device of the present invention can improve the accuracy of identifying the category of the object in the image.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Embodiments for Carrying Out the Invention

[0022] Hereinafter, preferred embodiments for carrying out the present invention will be described in detail with reference to the accompanying drawings. Note that such embodiments are merely examples and do not limit the present invention.

[0023] In order to solve the problems existing in the prior art, the present invention provides a feature extraction method for fine-grained verification. By focusing on extracting features highly relevant to the object in the image waiting for matching, it is possible to quickly and accurately find an image containing the object of the same category from the image library.

[0024] FIG. 1 is a diagram showing a feature extraction device 100 in an embodiment of the present invention.

[0025] As shown in FIG. 1, the apparatus 100 for extracting features includes a feature extraction module 10, a feature fusion module 20, a classifier module 30, and a foreground enhancement module 40. The feature extraction module 10 can be implemented, for example, using backbone blocks in a backbone network, and may include, for example, four backbone blocks 1, 2, 3, and 4. The classifier module 30 includes eight classifiers 11 to 18. The apparatus 100 further includes a classifier 19 for the foreground enhancement module 40.

[0026] As can be understood, although four backbone blocks are shown in FIG. 1, the present invention is not limited thereto, and two or more backbone blocks may be used. Each backbone block extracts features in the image waiting for matching at one specific (predetermined) scale.

[0027] Preferably, the number of classifiers included in the classifier module 30 is twice the number of backbone blocks. The number of feature sets after fusion in the feature fusion module 20 is also twice the number of backbone blocks, and each feature set is assigned to one corresponding classifier in the classifier module 30.

[0028] Note that the classifier module 30 can be realized, for example, by a fully connected layer in a neural network.

[0029] Hereinafter, the operations of each module in the apparatus 100 shown in FIG. 1 will be described in detail.

[0030] First, by inputting the image waiting for matching into the four backbone blocks in the feature extraction module 10 respectively, feature sets f1 to f4 having four scales are extracted respectively.

[0031] Next, the feature fusion module 20 receives the feature sets f1 to f4 with four scales from the feature extraction module 10 and performs fusion in both the feature dimension and the scale dimension. First, in the feature dimension, the feature sets f1 to f4 of the four scales are aligned to the same level at once, and then the features aligned in the feature dimension are aligned two by two in the order of "from small to large" and "from large to small" in the scale dimension, thereby generating the fused features (fusion features) F1 to F8.

[0032] For example, when aligning in the order of "from small to large" in the scale dimension, after aligning the feature set f4 with the feature set f3, an aligned feature set f4' is generated. The aligned feature set f4' is added to the feature set f3 to generate a fused feature set (fusion feature set) F3, and also, the initial feature set f4 remains unchanged and becomes the fusion feature set F4. The same operation is performed for the feature sets f3 to f1. As can be understood, when aligning in the order of "from small to large" in the scale dimension, the feature set f1 with the largest scale dimension is no longer aligned.

[0033] For example, when aligning in the order of "from large to small" in the scale dimension, after aligning the feature set f1 with the feature set f2, an aligned feature set f1' is generated. The aligned feature set f1' is added to the feature set f2 to generate the fusion feature set F6, and also, the initial feature set f1 remains unchanged and becomes the fusion feature set F5. The same operation is performed for the feature sets f2 to f4. As can be understood, when aligning in the order of "from large to small" in the scale dimension, the feature set f4 with the smallest scale dimension is no longer aligned.

[0034] Note that according to the needs, the feature sets f1 to f4 with four scales may be aligned to one specific (predetermined) level in the feature dimension C, and this level may be any one of the feature dimensions in these feature sets, or another feature dimension.

[0035] Also, when adding the sorted feature set f1’ to the feature set f2, appropriate weights may be assigned to f1’ according to needs. The same applies to the feature sets f2 to f4.

[0036] Figure 2 is a diagram showing the process of aligning the extracted features. For the sake of convenience in explanation, only the mutual alignment between two feature sets is shown in Figure 2. In Figure 2, C represents the feature dimension (e.g., the length of the feature), and N represents the scale dimension (i.e., the number of elements in the feature map, e.g., the resolution). The larger the value of C, the more semantic information the feature has and the lower the resolution. Also, the larger the value of N, the higher the resolution of the feature and the less semantic information.

[0037] As shown in Figure 2, first, align the feature set f1 with the level of the feature set f2 in the feature dimension C. As can be understood, both the feature sets f3 and f4 are aligned with the level of f2.

[0038] After that, by increasing the feature set f2 to the level of the feature set f1 in the scale dimension N and decreasing the feature set f1 to the level of the feature set f2, the fused feature sets F1, F2, F5, and F6 are obtained.

[0039] As can be understood, similarly, the fused feature sets F3, F4, F7, and F8 can also be obtained.

[0040] Also, as can be understood, the fused feature sets F1 to F4 are the results of fusing the extracted feature sets f1 to f4 in ascending order in the scale dimension, and the fused feature sets F5 to F8 are the results of fusing the extracted feature sets f1 to f4 in descending order in the scale dimension.

[0041] Next, an example of the calculation method of feature fusion will be introduced in conjunction with Figure 1.

[0042] First, to fuse the features output from each backbone block, the features of each backbone block are aligned to the same feature dimension level. For example, the alignment of features may be performed by a combination of non-linear transformation and linear transformation such as an activation function. However, as can be understood, the present invention is not limited thereto, and the transformation of features may be performed by any other appropriate method.

[0043] Next, by performing fusion on the feature sets f1 to f4 in the order of "from small to large" in the scale dimension, feature_td i can be obtained. For example, feature_td i can be calculated by the following formula.

[0044] feature_td i = feature_ i + upsampling(feature i+1 ) (1) Among them, feature_ i represents the features aligned to the same feature dimension level, where i = 1, 2,..., M - 1, and M represents the number of backbone blocks.

[0045] Note that the upsampling transformation of feature i+1 may be realized by, for example, 1D convolution or linear transformation.

[0046] Next, by performing fusion on the feature sets f1 to f4 in the order of "from large to small" in the scale dimension, feature_bu j can be obtained. For example, feature_bu j can be calculated by the following formula.

[0047] feature_bu j = feature_ j + downsampling(feature_ j-1 ) (2) Among them, feature_ jrepresents features aligned at the same feature dimension level, where j = 2, 3, …, M, and M represents the number of backbone blocks.

[0048] For example, downsampling of feature_ j-1 can be realized by 1D convolution or linear transformation.

[0049] Here, for the sake of convenience of explanation, the subscripts i and j are used to represent the feature alignment in dimension N "from small to large" and "from large to small", respectively.

[0050] The above introduces one calculation example of feature fusion, but the present invention is not limited thereto, and feature fusion may be performed in any suitable conventional manner.

[0051] Note that after aligning all the feature sets in the feature dimension C once, they may be aligned two by two in a step-by-step manner in the scale dimension N, or two by two in a cross-step manner. Aligning step by step means aligning between two adjacent feature sets, and cross-step alignment means aligning between two non-adjacent feature sets. Taking the case where four feature sets f1 to f4 are aligned in ascending order in the scale dimension as an example, in the case of step-by-step alignment, f1 may be aligned with f2, f2 with f3, and f3 with f4, or only the alignment from f2 to f3 may be performed, and the alignments from f1 to f2 and from f3 to f4 may not be performed. Also, in the case of cross-step alignment, f1 may be directly aligned with f3, and f2 with f4, or only the alignment from f1 to f3 may be performed, and the alignment from f2 to f4 may not be performed. In short, according to actual needs, the feature sets that need to be aligned may be customized.

[0052] Also, according to actual needs, all the extracted features may be fused, or some of the extracted features may be fused.

[0053] As can be understood, the generated fused features include both the aligned features aligned in the feature dimension and the scale dimension, and the extracted features before alignment.

[0054] By aligning the extracted features in the feature dimension and the scale dimension, the image information can be made richer and the prediction result of the classifier module 30 can be made more accurate.

[0055] Referring to FIG. 1 again. After the feature fusion module 20 obtains eight fused feature sets F1 to F8, the classifiers 11 to 18 of the classifier module 30 respectively perform predictions on the eight fused feature sets F1 to F8 after fusion. For example, the classifier 11 performs a prediction on the feature set F1, the classifier 12 performs a prediction on the feature set F2, the classifier 13 performs a prediction on the feature set F3, and the others can be analogized based on this.

[0056] The classifiers 11 to 18 minimize the difference between the prediction results of each fused feature arranged in the order of "small to large" in the scale dimension and the prediction results of all fused features arranged in the order of "large to small", or are trained to minimize the difference between the prediction results of each fused feature arranged in the order of "small to large" and the prediction results of all fused features arranged in the order of "large to small". For example, the prediction result of the feature set F1 by the classifier 11 is compared with each of the prediction results of the feature sets F5 to F8 by the classifiers 15 to 18, and the difference between the prediction result of the feature set F1 and each of the prediction results of the feature sets F5 to F8 is minimized. For example, the prediction result of the feature set F5 by the classifier 15 is compared with each of the prediction results of the feature sets F1 to F4 by the classifiers 11 to 14, and the difference between the prediction result of the feature set F5 and each of the prediction results of the feature sets F1 to F4 is minimized.

[0057] Note that the results output by the classifiers 11 to 18 are the probability distributions of the objects (e.g., products) in the image waiting for matching belonging to each (product) category. Among them, the number of categories may be defined in advance or may be determined according to the number of categories of the images in the image library. Minimizing the difference between the prediction results means making the probability distributions of the two feature sets match as much as possible.

[0058] For example, the classifiers 11 to 18 can be trained using a loss function based on KL (Kullback-Leibler divergence) divergence. Y_td i represents the classification result of feature_td i , and Y_bu j represents the classification result of feature_bu j . Suppose Y_td i has a dimension of [N i , C gt , and Y_bu j has a dimension of [N j , C gt , where C gt is the number of categories for training.

[0059] By calculating the difference between Y_td i and Y_bu j , it is possible to impose a constraint such that Y_td i approaches Y_bu j . To keep the dimension unchanged, Y_td i and Y_bu j represent the average value of the feature in dimension N by the following formula (3). The first loss function used for the classifiers 11 to 18 can be calculated by the following formula (3).

[0060]

Equation

[0061] As can be understood, the above loss function \(L1\) is just one example, and the present invention is not limited thereto. Any other appropriate loss function may be used as long as the same function can be realized. For example, the loss function \(L1\) may be constructed based on the L-2 norm.

[0062] The feature fusion module 20 further provides the fused features to the foreground enhancement module 40. The foreground enhancement module 40 is configured to divide the fused features into foreground features and background features based on a predetermined rule, and during training, for other classifiers, increase one class for the background features compared to the foreground features, where the foreground features are the features associated with the category of the object in the image waiting for matching.

[0063] Specifically, in this embodiment, considering the fused feature \(feature\_bu\) j and assuming that some elements in \(feature\_bu\) j belong to the foreground in the image waiting for matching, and other elements belong to the background. As described above, \(Y\_bu\) j represents the predicted probability that each element of \(feature\_bu\) j belongs to each category, and this predicted probability should be high for the correct category of the foreground elements and low for the correct category of the background elements.

[0064] First, sort the predicted probabilities of each element in \(feature\_bu\) j in the correct category in the order from "large to small". After sorting, assume that the last \((1 - \gamma1)\) (for example, \(\gamma1 = 1 / 4\)) elements are background elements, and the first \(\gamma2\) (for example, \(1 / 16\)) elements may be foreground elements, where \(\gamma2 < \gamma1\).

[0065] After that, a constraint is imposed such that the element having the highest prediction probability in the correct category is always determined to be correct. Also, the average value of the prediction probabilities of the elements of γ2 (for example, 1 / 16) from the beginning is always determined to be correct, and a constraint is imposed such that the elements of (1 - γ1) (for example, γ1 = 1 / 4) from the end should not be determined to belong to any category of the training data (for example, the images in the image library).

[0066] After identifying all foreground features and background features, add one pseudo label to all background features (increase one class compared to the foreground features), and input them to the classifier 19 to make predictions.

[0067] As can be understood, in the present invention, the foreground feature refers to a feature having a high relevance to the category of the object (product) in the image waiting for matching.

[0068] The second loss function for training the classifier 19 is as follows, and it is based on, for example, the cross-entropy loss function.

[0069]

Number

[0070]

Number

[0071] Therefore, the total loss function may be, for example, the weighted sum of the first loss function and the second loss function as follows.

[0072] L total = w1L1 + w2L2 (5) Among them, w1 and w2 are the weights of the loss functions L1 and L2, respectively, and may be set according to actual needs.

[0073] Total loss function L total By minimizing the total loss function L to train the classifiers 11 to 19, the network parameters of the feature extraction module 10, the feature fusion module 20, the classifier module 30, and the foreground enhancement module 40 are optimized. In particular, the feature extraction module 10 is enabled to extract features highly relevant to the category of the object in the image waiting for matching.

[0074] FIG. 3 is a diagram showing an image matching device 200 according to an embodiment of the present invention. The image matching device 200 includes a feature extraction module 10, a feature fusion module 20, and an image matching module 50. Among them, the network parameters of the feature extraction module 10 and the feature fusion module 20 are set by training the above-described device 100. That is, the features of the image extracted by the feature extraction module 10 have a high correlation with the object of the category waiting to be identified in the image.

[0075] The working principles of the feature extraction module 10 and the feature fusion module 20 in FIG. 3 are the same as those described above in combination with FIGS. 1 and 2, and the detailed description thereof is omitted here.

[0076] After receiving the fused features from the feature fusion module 20, the image matching module 50 stitches all or part of the fused features, calculates the distance between the stitched features and all the images in the image library 60, and sorts all the images in the image library based on the calculated results. The probability that the first image in the sorted result matches the category of the object (product) in the image waiting for matching is the highest.

[0077] Note that before performing the stitching of the features, it is necessary to take the average of the features waiting for stitching. For example, taking the stitching of the feature feature_bu j as an example, what is stitched is the average value of all elements in feature_bu j ​

[0078] In addition, for the calculation of the distance between the features after stitching and the images in the image library 60, any conventional distance calculation method, such as those based on Euclidean distance or cosine distance, etc., may be adopted.

[0079] As can be understood, the image library 60 may include, for example, images of various products in a supermarket, or may include any images related to category distinction.

[0080] Also, as can be understood, according to actual needs, all the fused features may be stitched, or only some of the fused features may be stitched.

[0081] As shown in FIG. 4, based on the calculated distance, the image in the image library 60 with the closest (smallest) distance is used as the first image in the sequence, and analogies can be made based on this. Therefore, it can be determined that the category of the object in the image waiting for matching matches the category of the first image.

[0082] The image matching device 200 of the present invention can accurately identify the category of the object in the image waiting for matching.

[0083] FIG. 5 is a flowchart of the feature extraction method 500 in an embodiment of the present invention. The method 500 is realized by the above-mentioned feature extraction device 100.

[0084] First, in step 501, the feature extraction module 10 in the device 100 is used to extract features having a plurality of scales and a plurality of feature dimensions from the image waiting for matching, respectively.

[0085] Next, in step 502, the feature fusion module 20 is used to align all or some of the extracted features in the feature dimension, and then align the features aligned in the feature dimension two by two in the order of "from small to large" and "from large to small" in the scale dimension to generate fused features, where the fused features include the aligned features and the features before alignment.

[0086] Subsequently, in step 503, a plurality of classifiers in the classifier module 30 are trained, each classifier is assigned to the corresponding fused feature, and the difference between the prediction result of each fused feature aligned in the order of "from large to small" in the scale dimension and the prediction results of all fused features aligned in the order of "from small to large" is minimized, or the difference between the prediction result of each fused feature aligned in the order of "from small to large" and the prediction results of all fused features aligned in the order of "from large to small" is minimized.

[0087] Finally, in step 504, the foreground enhancement module 40 is used to divide the fused features into foreground features and background features, and during training, for other classifiers, one class is added to the background features compared with the foreground features, where the foreground features are the features associated with the category of the object in the image waiting for matching.

[0088] As can be understood, the specific operation methods of each step of the method 500 are described in detail in the above-mentioned apparatus 100 in combination with FIGS. 1 and 2, and the detailed description thereof is omitted here.

[0089] FIG. 6 is a flowchart of an image matching method 600 according to an embodiment of the present invention. The method 600 is implemented by the above-mentioned image matching device 200.

[0090] First, in step 601, the feature extraction module 10 in the image matching device 200 is used to extract features having a plurality of scales and a plurality of feature dimensions from the image waiting for matching respectively. The extracted features have a high correlation with the object in the image waiting for matching.

[0091] Next, in step 602, using the feature fusion module 20, all or some of the extracted features are aligned in the feature dimension, and then the features aligned in the feature dimension are aligned two by two in the scale dimension in the order of "from small to large" and "from large to small", thereby generating fused features, where the fused features include the aligned features and the features before alignment.

[0092] Finally, in step 603, all or some of the fused features are stitched, the distance between the stitched features and all the images in the image library is calculated, and all the images in the image library are rearranged based on the calculated results.

[0093] As can be understood, the specific operation methods of each step of method 600 have been described in detail in the above-mentioned image matching device 200 in conjunction with FIG. 3, and the detailed description thereof is omitted here.

[0094] The above-described method may be fully realized by a computer-executable program, or may be partially or fully realized by hardware and / or firmware. When realized by hardware and / or firmware, or when installing a computer-executable program on a hardware device capable of executing the program, the above-described feature extraction device and image matching device can be realized.

[0095] Each module or unit in the above-described device or equipment may be configured in a software, firmware, hardware, or a combination thereof manner. The specific means and methods that can be used for the configuration are well-known to those skilled in the art, and the detailed description thereof is omitted here. When realized by software or firmware, a program constituting the software is installed from a storage medium or network into a computer having a dedicated hardware structure (for example, the general-purpose computer 700 shown in FIG. 7), and when various programs are installed in the computer, various functions and the like can be realized.

[0096] FIG. 7 is a block diagram showing a hardware configuration (general-purpose computer) 700 that can implement the method, apparatus, and / or equipment in an embodiment of the present invention.

[0097] The general-purpose computer 700 may be, for example, a computer system. It should be noted that the general-purpose computer 700 is merely an example and does not limit the application scope or functions of the method, apparatus, and equipment according to the present invention. Also, the general-purpose computer 700 does not depend on any module, unit, etc. or their combinations in the above-mentioned method, apparatus, and equipment.

[0098] As shown in FIG. 7, the central processing unit (CPU) 701 performs various processes based on a program stored in the ROM 702 or a program loaded from the storage unit 708 to the RAM 703. In the RAM 703, data and the like necessary when the CPU 701 performs various processes can also be stored according to needs. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output interface 705 is also connected to the bus 704.

[0099] In addition, the following components are further connected to the input / output interface 705, that is, an input unit 706 including a keyboard, etc., a display unit including a liquid crystal display (LCD), etc., and an output unit 707 including a speaker, etc., a storage unit 708 including a hard disk, etc., and a communication unit 709 including a network interface card, for example, a LAN card, a modem, etc. The communication unit 709 performs communication processing via a network such as the Internet or a LAN. The drive 710 may be connected to the input / output interface 705 according to needs. A removable medium 711, for example, a semiconductor memory, etc., can be set in the drive 710 as needed, and the computer program read from it can be installed in the storage unit 708.

[0100] Furthermore, the present invention further provides a program product including machine (computer) readable instruction codes. When such instruction codes are read and executed by a machine, the methods in the embodiments of the present invention described above can be executed. Correspondingly, various storage media that carry such a program product, such as magnetic disks (including floppy disks (registered trademark)), optical disks (including CD-ROMs and DVDs), magneto-optical disks (including MD (registered trademark)), and semiconductor memories, are also included in the present invention.

[0101] The above storage media may include, for example, magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc., but are not limited thereto.

[0102] Also, each operation (processing / step) in the above method can also be realized in the form of a computer-executable program stored in various machine-readable storage media.

[0103] Also, regarding the above embodiments, etc., the following is further disclosed as an appendix.

[0104] (Appendix 1) An apparatus for extracting features, a feature extraction module configured to extract features having a plurality of scales and a plurality of feature dimensions from an image waiting for matching respectively; a feature fusion module configured to align all or part of the extracted features in terms of feature dimensions, and generate a fusion feature by aligning the features aligned in terms of feature dimensions two by two in the order of "from small to large" and "from large to small" in terms of scale dimensions, wherein the fusion feature includes the aligned features and the features before alignment; A classifier module including a plurality of classifiers, wherein each of the plurality of classifiers is assigned to a corresponding fusion feature and is trained to minimize the difference between the prediction result of each fusion feature arranged in the order of "from small to large" in the scale dimension and the prediction results of all fusion features arranged in the order of "from large to small", or to minimize the difference between the prediction result of each fusion feature arranged in the order of "from large to small" and the prediction results of all fusion features arranged in the order of "from small to large"; and A foreground enhancement module configured to divide the fusion features into foreground features and background features based on a predetermined rule, and during training, for other classifiers, to increase one class for the background features compared to the foreground features, wherein the foreground features are features associated with the identification of the category of the object in the image waiting for matching. The device includes such a module.

[0105] (Appendix 2) The device according to Appendix 1, wherein The feature extraction module includes N backbone blocks, the plurality of scales include N scales, each backbone block is configured to extract features having one of the N scales, and N is an integer greater than or equal to 2.

[0106] (Appendix 3) The device according to Appendix 2, wherein The plurality of classifiers include N×2 classifiers, the fusion features are divided into N×2 feature sets based on scale, and the N×2 classifiers are respectively assigned to the N×2 feature sets.

[0107] (Appendix 4) The device according to any one of Appendices 1 to 3, wherein The prediction result is a probability distribution of the object belonging to each of two or more categories.

[0108] (Appendix 5) The device according to Appendix 4, wherein The apparatus is to divide the fusion feature into the foreground feature and the background feature based on the sorting in the correct category of the probability distribution output from all classifiers.

[0109] (Appendix 6) The apparatus according to any one of Appendices 1 to 3, wherein the plurality of classifiers are trained based on a first loss function and a second loss function, and other classifiers in the foreground enhancement module are trained based on the second loss function.

[0110] (Appendix 7) The apparatus according to Appendix 6, wherein the first loss function is a Kullback-Leibler divergence loss function, and the second loss function is a cross-entropy loss function.

[0111] (Appendix 8) The apparatus according to Appendix 6, wherein the total loss function is a weighted sum of the first loss function and the second loss function, and the training includes minimizing the total loss function.

[0112] (Appendix 9) The apparatus according to any one of Appendices 1 to 3, wherein aligning two features aligned in the feature dimension in the order of "from small to large" and "from large to small" in the scale dimension is executed in parallel or sequentially.

[0113] (Appendix 10) The apparatus according to Appendix 9, wherein the feature dimension is the length of the extracted feature.

[0114] (Appendix 11) The apparatus according to any one of Appendices 1 to 3, wherein the scale is obtained based on the resolution of the image waiting for matching.

[0115] (Appendix 12) A method for extracting features, comprising: extracting features having a plurality of scales and a plurality of feature dimensions from an image waiting for matching respectively; aligning all or part of the extracted features in terms of feature dimensions, and generating a fused feature by aligning the features aligned in terms of feature dimensions two by two in ascending and descending order in terms of scale dimension, wherein the fused feature includes the aligned features and the features before alignment; training a plurality of classifiers, each of the plurality of classifiers being assigned to a corresponding fused feature and minimizing the difference between the prediction result of each fused feature aligned in ascending order in terms of scale dimension and the prediction results of all fused features aligned in descending order, or minimizing the difference between the prediction result of each fused feature aligned in descending order in terms of scale dimension and the prediction results of all fused features aligned in ascending order; and dividing the fused feature into a foreground feature and a background feature based on a predetermined rule, and during the training, for other classifiers, increasing one class for the background feature compared to the foreground feature, wherein the foreground feature is a feature associated with the category of the object in the image waiting for matching.

[0116] (Appendix 13) The method according to Appendix 12, wherein the prediction result is a probability distribution of the object belonging to each of two or more categories.

[0117] (Appendix 14) The method according to Appendix 12, wherein the predetermined rule is to divide the fused feature into the foreground feature and the background feature based on the rearrangement in the correct category of the probability distributions output from all classifiers.

[0118] (Appendix 15) The method according to any one of Appendices 12 to 14, wherein The method in which the plurality of classifiers are trained based on a first loss function and a second loss function, and the other classifiers are trained based on the second loss function.

[0119] (Appendix 16) The method according to Appendix 15, wherein the first loss function is a Kullback-Leibler divergence loss function, and the second loss function is a cross-entropy loss function, the total loss function is a weighted sum of the first loss function and the second loss function, the training includes minimizing the total loss function.

[0120] (Appendix 17) An image matching device, comprising a feature extraction module configured to extract features having a plurality of scales and a plurality of feature dimensions from a matching waiting image, wherein the extracted features relate to an object in the matching waiting image; a feature fusion module configured to align all or some of the extracted features in the feature dimension, and generate a fusion feature by aligning the features aligned in the feature dimension two by two in the order of "from small to large" and "from large to small" in the scale dimension, wherein the fusion feature includes the aligned features and the features before alignment; and a matching module configured to stitch all or some of the fusion features, calculate the distance between the stitched features and all the images in the image library, and rearrange all the images in the image library based on the calculated result. The network parameters of each of the feature extraction module and the feature fusion module are set by training the device according to any one of Appendices 1 to 11.

[0121] (Appendix 18) An image matching method, comprising Extract features having a plurality of scales and a plurality of feature dimensions from the image waiting for matching, and the extracted features relate to the object in the image waiting for matching; Arrange all or some of the extracted features in the feature dimension, and then arrange the features arranged in the feature dimension two by two in the order of "from small to large" and "from large to small" in the scale dimension to generate a fused feature, and the fused feature includes the arranged features and the features before arrangement; and Include stitching all or some of the fused features, calculating the distance between the stitched features and all the images in the image library, and sorting all the images in the image library based on the calculated results, The method for image matching, wherein the extraction and fusion of the features are executed by the feature extraction module and the feature fusion module in the image matching device described in Appendix 17.

[0122] (Appendix 19) A computer-readable storage medium storing a program, When the program is executed by a processor, the computer-readable storage medium causes a computer to execute the method described in any one of Appendices 12 to 16.

[0123] (Appendix 20) A computer-readable storage medium storing a program, When the program is executed by a processor, the computer-readable storage medium causes a computer to execute the method described in Appendix 18.

[0124] The preferred embodiments of the present invention have been described above. However, the present invention is not limited to this embodiment. Without departing from the spirit of the present invention, any changes to the present invention belong to the technical scope of the present invention.

Claims

1. An apparatus for extracting features, comprising: A feature extraction module configured to extract features having a plurality of scales and a plurality of feature dimensions from an image waiting for matching respectively; A feature fusion module configured to align all or part of the extracted features in the feature dimension, and generate a fusion feature by aligning the features aligned in the feature dimension two by two in the order of "from small to large" and "from large to small" in the scale dimension, wherein the fusion feature includes the aligned features and the features before alignment; A classifier module including a plurality of classifiers, wherein each of the plurality of classifiers is assigned to a corresponding fusion feature, and the difference between the prediction result of each fusion feature aligned in the order of "from small to large" in the scale dimension and the prediction results of all fusion features aligned in the order of "from large to small" is minimized, or the difference between the prediction result of each fusion feature aligned in the order of "from small to large" and the prediction results of all fusion features aligned in the order of "from large to small" is minimized, and the classifier module is trained accordingly; and A foreground enhancement module configured to divide the fusion feature into a foreground feature and a background feature based on a predetermined rule, and increase one class for the background feature compared to the foreground feature for other classifiers during the training, wherein the foreground feature is a feature associated with the category of the object in the image waiting for matching, and the apparatus includes the foreground enhancement module.

2. The apparatus according to claim 1, wherein The feature extraction module includes N backbone blocks, the plurality of scales include N scales, each backbone block is configured to extract features having one of the N scales, and N is an integer greater than or equal to 2.

3. The apparatus according to claim 2, wherein The plurality of classifiers include N×2 classifiers, the fusion feature is divided into N×2 feature sets based on the scale, and the N×2 classifiers are respectively assigned to the N×2 feature sets.

4. The apparatus according to any one of claims 1 to 3, wherein The prediction result is a probability distribution of the object belonging to each of two or more categories.

5. The apparatus according to claim 4, wherein The apparatus is such that the predetermined rule divides the fusion feature into the foreground feature and the background feature based on the sorting in the correct category of the probability distribution output by all the classifiers.

6. The apparatus according to any one of claims 1 to 3, wherein the plurality of classifiers are trained based on a first loss function and a second loss function, and other classifiers in the foreground enhancement module are trained based on the second loss function, and the total loss function is a weighted sum of the first loss function and the second loss function, wherein the training includes minimizing the total loss function.

7. The apparatus according to any one of claims 1 to 3, wherein the feature dimension is the length of the feature to be extracted, and the scale is obtained based on the resolution of the image waiting for matching.

8. An image matching device, a feature extraction module configured to extract features having a plurality of scales and a plurality of feature dimensions from an image waiting for matching, wherein the extracted features relate to an object in the image waiting for matching; a feature fusion module configured to align all or some of the extracted features in terms of feature dimension, and then align the features aligned in terms of feature dimension two by two in the order of "from small to large" and "from large to small" in terms of scale dimension to generate a fusion feature, wherein the fusion feature includes the aligned features and the features before alignment; and a matching module configured to stitch all or some of the fusion features, calculate the distance between the stitched features and all the images in the image library, and sort all the images in the image library based on the calculated result, wherein network parameters of each of the feature extraction module and the feature fusion module are set by training the apparatus according to any one of claims 1 to 3.

9. A program for causing a computer to execute a method for extracting features, wherein the method extracts features having a plurality of scales and a plurality of feature dimensions from an image waiting for matching respectively; Arrange all or some of the extracted features in feature dimensions, and then arrange the features arranged in feature dimensions two by two in the order of "from small to large" and "from large to small" in the scale dimension to generate a fused feature, where the fused feature includes the arranged features and the features before arrangement; Train a plurality of classifiers, where each of the plurality of classifiers is assigned to a corresponding fused feature and minimizes the difference between the prediction result of each fused feature arranged in the order of "from small to large" in the scale dimension and the prediction results of all fused features arranged in the order of "from large to small", or is trained to minimize the difference between the prediction result of each fused feature arranged in the order of "from large to small" and the prediction results of all fused features arranged in the order of "from small to large"; and Based on a predetermined rule, divide the fused feature into a foreground feature and a background feature, and during the training, for other classifiers, increase one class for the background feature compared to the foreground feature, where the foreground feature is a feature associated with the category of the object in the image waiting for matching, including this, a program.

10. A program for causing a computer to execute an image matching method, where the image matching method includes extracting features having a plurality of scales and a plurality of feature dimensions from an image waiting for matching, and the extracted features are related to the object in the image waiting for matching; Arrange all or some of the extracted features in feature dimensions, and then arrange the features arranged in feature dimensions two by two in the order of "from small to large" and "from large to small" in the scale dimension to generate a fused feature, where the fused feature includes the arranged features and the features before arrangement; and stitch all or some of the fused features, calculate the distance between the stitched features and all images in the image library, and rearrange all images in the image library based on the calculated result, The extraction and fusion of the features are executed by the feature extraction module and the feature fusion module in the image matching device according to Claim 8, a program.

Citation Information

Cited By

  • Target attribute identification method and related device

    CN121962778A