Feature extraction device, image matching equipment and computer readable storage medium

Through the combination of feature extraction, fusion and classifier modules, the problem of high difficulty in product recognition in large supermarkets is solved, and fast and accurate product category recognition is achieved.

CN120259678APending Publication Date: 2025-07-04FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410012267.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In large supermarkets, it is difficult to identify products, especially due to the similarity and angle differences between product images and complex backgrounds, making it difficult for the prior art to quickly and accurately identify product categories.

Method used

Feature extraction module is used to extract features of multi-scale and multi-feature dimensions from the image, and the feature fusion module is used to pull the feature dimensions and pull it from small to large and from large to small in the scale dimension. It is combined with classifier module training to minimize the difference in prediction results, and separate the foreground and background features through the foreground enhancement module.

Benefits of technology

Improves the recognition accuracy of object categories in the image and can quickly and accurately find similar objects in the image library.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259678A_ABST
    Figure CN120259678A_ABST
Patent Text Reader

Abstract

The invention discloses a feature extraction device, an image matching apparatus and a medium. The feature extraction device comprises a feature extraction module which extracts features with a plurality of scales and a plurality of feature dimensions from an image to be matched; the feature fusion module is used for aligning all or part of the extracted features on the feature dimension, and then aligning the features aligned on the feature dimension in pairs on the scale dimension from small to large and from large to small so as to generate fused features; a classifier module comprising a plurality of classifiers, each classifier being assigned to a corresponding fusion feature and being trained such that a difference between a prediction result of each fusion feature drawn from small to large in a scale dimension and prediction results of all fusion features drawn from large to small is minimum, and vice versa; and the foreground enhancement module divides the fusion features into foreground features and background features according to a preset rule, and adds a class to the background features for another classifier during training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of neural networks, and more particularly to techniques for feature extraction for fine-grained verification. Background Art

[0002] In recent years, techniques for understanding shoppers' interests in product selection based on images or image frames in images or videos have received extensive attention. This generally requires matching a query product image with corresponding images in an image library.

[0003] In a large supermarket, there are often hundreds or thousands of products. Some products of different categories look very similar. For products of the same category, their appearances vary greatly when viewed from different angles. Moreover, the backgrounds of product images are usually complex, which may include customers' hands, shopping carts, shelves, etc., all of which increase the difficulty of obtaining discriminative features of the products. Summary of the Invention

[0004] A brief summary of the present disclosure is given below in order to provide a basic understanding of certain aspects of the present disclosure. It should be understood that this summary is not an exhaustive summary of the present disclosure. It is not intended to identify the key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.

[0005] According to one aspect of the present disclosure, there is provided an apparatus for feature extraction, including: a feature extraction module configured to extract features having multiple scales and multiple feature dimensions from an image to be matched; a feature fusion module configured to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise in the scale dimension from small to large and from large to small, thereby generating fused features, the fused features including the aligned features and the features before alignment; a classifier module including a plurality of classifiers, each of the plurality of classifiers being assigned to a corresponding fused feature and being trained such that the difference between the prediction result of each fused feature aligned from small to large in the scale dimension and the prediction results of all fused features aligned from large to small is minimized, or such that the difference between the prediction result of each fused feature aligned from large to small and the prediction results of all fused features aligned from small to large is minimized; and a foreground enhancement module configured to divide the fused features into foreground features and background features according to a preset rule, and add one class to the background features compared to the foreground features for another classifier during training, the foreground features being features associated with the category recognition of the object in the image to be matched.

[0006] Preferably, the feature extraction module includes N backbone blocks, the multiple scales include N scales, and each backbone block is configured to extract features with one of the N scales, where N is an integer greater than or equal to 2.

[0007] Preferably, the multiple classifiers include N×2 classifiers, the fused features are divided into N×2 feature sets based on the scale, and the N×2 classifiers are respectively assigned to the N×2 feature sets.

[0008] Preferably, the multiple classifiers can be implemented with fully connected layers.

[0009] Preferably, the prediction result is the probability distribution of the object belonging to each of two or more classes.

[0010] Preferably, the preset rule is to divide the fused features into foreground features and background features based on the ranking on the true class in the probability distributions output by all classifiers.

[0011] Preferably, the multiple classifiers are trained based on a first loss function and a second loss function, and wherein, the additional classifier in the foreground enhancement module is trained based on the second loss function.

[0012] Preferably, the first loss function is the Kullback-Leibler divergence loss function, and wherein, the second loss function is the cross-entropy loss function.

[0013] Preferably, the total loss function is the weighted sum of the first loss function and the second loss function, and wherein, the training includes minimizing the total loss function.

[0014] Preferably, the features aligned in the feature dimension are pairwise aligned from small to large and from large to small in the scale dimension, and are executed in parallel or sequentially.

[0015] Preferably, the feature dimension is the length of the extracted features, and the scale is obtained based on the resolution of the image to be matched.

[0016] According to another aspect of the present disclosure, there is provided an image matching device, including: a feature extraction module configured to extract features having multiple scales and multiple feature dimensions from an image to be matched, wherein the extracted features are related to an object in the image to be matched; a feature fusion module configured to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise in the scale dimension from small to large and from large to small, so as to generate fused features, the fused features including the aligned features and the features before alignment; and a matching module configured to splice all or a part of the fused features together, calculate the distances between the spliced features and all the images in an image library, and sort all the images in the image library according to the calculation results, wherein the network parameters of the feature extraction module and the feature fusion module are set by training the above-described device for extracting features.

[0017] According to still another aspect of the present disclosure, there is provided a computer-readable storage medium having a program stored thereon, the program causing a computer to execute a method for extracting features when being executed by a processor, the method including: extracting features having multiple scales and multiple feature dimensions from an image to be matched respectively; aligning all or part of the extracted features in the feature dimension, and then aligning the features aligned in the feature dimension pairwise in the scale dimension from small to large and from large to small, so as to generate fused features, the fused features including the aligned features and the features before alignment; training a plurality of classifiers, each of the plurality of classifiers being assigned to a corresponding fused feature and being trained to minimize the difference between the prediction result of each fused feature aligned from small to large in the scale dimension and the prediction results of all the fused features aligned from large to small, or to minimize the difference between the prediction result of each fused feature aligned from large to small in the scale dimension and the prediction results of all the fused features aligned from small to large; and dividing the fused features into foreground features and background features according to a preset rule, and adding one class for the background features compared with the foreground features when training another classifier, the foreground features being features associated with the category recognition of an object in the image.

[0018] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium storing a program which, when executed by a processor, causes a computer to execute an image matching method. The image matching method includes: extracting features having multiple scales and multiple feature dimensions from an image to be matched, wherein the extracted features are related to an object in the image to be matched; aligning all or part of the extracted features in the feature dimension, and then aligning the features aligned in the feature dimension pairwise from small to large and from large to small in the scale dimension, thereby generating fused features, the fused features including the aligned features and the features before alignment; and splicing all or part of the fused features together, calculating the distances between the spliced features and all the images in an image library, and sorting all the images in the image library according to the calculation results, wherein the extraction and fusion of the features are performed by the feature extraction module and the feature fusion module in the above image matching device.

[0019] According to other aspects of the present disclosure, corresponding computer program codes and computer program products are also provided.

[0020] Through the device for feature extraction and the image matching device of the present disclosure, the accuracy of identifying object categories in images is improved.

[0021] These and other advantages of the present disclosure will become more apparent from the following detailed description of the preferred embodiments of the present disclosure in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To further illustrate the above and other advantages and features of the present disclosure, the specific embodiments of the present disclosure will be described in further detail below in conjunction with the accompanying drawings. The accompanying drawings, together with the following detailed description, are included in this specification and form a part of this specification. Elements having the same function and structure are denoted by the same reference numerals. It should be understood that these drawings only depict typical examples of the present disclosure and should not be regarded as limiting the scope of the present disclosure. In the drawings:

[0023] Figure 1 Schematically shows a device for feature extraction according to an embodiment of the present disclosure;

[0024] Figure 2 Schematically shows the process of feature fusion;

[0025] Figure 3 Schematically shows an image matching device according to an embodiment of the present disclosure;

[0026] Figure 4 Shows a schematic illustration of the distance calculation with respect to images in an image library;

[0027] Figure 5 shows a flowchart of a method for feature extraction according to an embodiment of the present disclosure;

[0028] Figure 6 shows a flowchart of an image matching method according to an embodiment of the present disclosure;

[0029] Figure 7 is a block diagram of an exemplary structure of a general - purpose personal computer in which the method and / or device according to an embodiment of the present disclosure can be implemented. Detailed Embodiments

[0030] Hereinafter, exemplary embodiments of the present disclosure will be described in conjunction with the accompanying drawings. For clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many embodiment - specific decisions must be made in the process of developing any such actual embodiment in order to achieve the specific goals of the developer, for example, to comply with those system - and business - related constraints, and these constraints may vary with different embodiments. In addition, it should also be understood that although the development work may be very complex and time - consuming, for those skilled in the art who benefit from the present disclosure, such development work is merely a routine task.

[0031] Here, it should also be noted that, in order to avoid obscuring the present disclosure with unnecessary details, only the device structures and / or processing steps closely related to the solution according to the present disclosure are shown in the drawings, while other details less related to the present disclosure are omitted.

[0032] To solve the problems of the prior art, the present disclosure proposes a feature extraction method for fine - grained verification. This method focuses on extracting features highly relevant to the object in the image to be matched, so as to quickly and accurately find the image containing the same type of object in the image library.

[0033] Figure 1 Schematically shows an apparatus 100 for feature extraction according to an embodiment of the present disclosure.

[0034] As Figure 1 shown, the apparatus 100 for feature extraction includes a feature extraction module 10, a feature fusion module 20, a classifier module 30, and a foreground enhancement module 40. The feature extraction module 10 can be implemented, for example, by using backbone blocks in a backbone network, and can include, for example, four backbone blocks 1, 2, 3, 4. The classifier module 30 includes eight classifiers 11 to 18. The apparatus 100 also includes a classifier 19 for the foreground enhancement module 40.

[0035] It should be understood that although Figure 1Four backbone blocks are shown, but the present disclosure is not limited thereto, and two or more backbone blocks can be used. Each backbone block extracts features in a specific scale from the image to be matched.

[0036] Preferably, the number of classifiers included in the classifier module 30 is twice the number of backbone blocks. The number of the fused feature sets in the feature fusion module 20 is also twice the number of backbone blocks, and each feature set is assigned to a corresponding classifier in the classifier module 30.

[0037] It should be noted that the classifier module 30 can be implemented, for example, using a fully connected layer in a neural network.

[0038] The following will be described in detail Figure 1 the operations of each module in the device 100

[0039] First, the images to be matched are respectively input into the four backbone blocks in the feature extraction module 10 to extract the feature sets f1 to f4 with four scales respectively.

[0040] Next, the feature fusion module 20 fuses the feature sets f1 to f4 with four scales received from the feature extraction module 10 in both the feature dimension and the scale dimension. First, the feature sets f1 to f4 with four scales are leveled to the same level at one time in the feature dimension, and then the features leveled in the feature dimension are leveled pairwise from small to large and from large to small in the scale dimension to generate the fused features F1 to F8.

[0041] For example, in the case of leveling from small to large in the scale dimension, after leveling the feature set f4 to the feature set f3, the leveled feature set f4’ is generated. The leveled feature set f4’ is added to the feature set f3 to generate the fused feature set F3, and the initial feature set f4 remains unchanged and serves as the fused feature set F4. The same operation is performed for the feature sets f3 to f1. It should be understood that in the case of leveling from small to large in the scale dimension, the feature set f1 with the largest scale dimension is no longer leveled.

[0042] For example, in the case of leveling from large to small in the scale dimension, after leveling the feature set f1 to the feature set f2, the leveled feature set f1’ is generated. The leveled feature set f1’ is added to the feature set f2 to generate the fused feature set F6, and the initial feature set f1 remains unchanged and serves as the fused feature set F5. The same operation is performed for the feature sets f2 to f4. It should be understood that in the case of leveling from large to small in the scale dimension, the feature set f4 with the smallest scale dimension is no longer leveled.

[0043] It should be noted that the feature sets f1 to f4 with four scales can be aligned to a specific level on the feature dimension C as needed. This level can be the feature dimension of any one of these feature sets or another feature dimension.

[0044] It should also be noted that when adding the aligned feature set f1' to the feature set f2, an appropriate weight can be added to f1' as needed. The same applies to the feature sets f2 to f4.

[0045] Figure 2 The process of aligning the extracted features is schematically shown. For the sake of explanation, Figure 2 only the alignment between two feature sets is shown. In Figure 2 , C represents the feature dimension (e.g., feature length), and N represents the scale dimension (i.e., the number of elements in the feature map, e.g., resolution). The larger the C value, the more semantic information the feature has, but the lower the resolution; the larger the N, the higher the resolution of the feature but the less semantic information.

[0046] Figure 2 As shown, first, on the feature dimension C, the feature set f1 is aligned to the level of the feature set f2. It should be understood that both the feature sets f3 and f4 are aligned to the level of f2.

[0047] Then, on the scale dimension N, the feature set f2 is increased to the level of the feature set f1, and the feature set f1 is decreased to the level of the feature set f2, resulting in the fused feature sets F1, F2, F5, and F6.

[0048] It should be understood that the fused feature sets F3, F4, F7, and F8 can be obtained similarly.

[0049] It should also be understood that the fused feature sets F1 to F4 are the result of fusing the extracted feature sets f1 to f4 in ascending order on the scale dimension, and the fused feature sets F5 to F8 are the result of fusing the extracted feature sets f1 to f4 in descending order on the scale dimension.

[0050] Next, in combination with Figure 1 an example of the calculation method for feature fusion will be introduced.

[0051] First, in order to fuse the features output by each backbone block, the features of each backbone block are aligned to the same feature dimension level. For example, the feature alignment can be performed through a combination of non-linear transformations such as activation functions and linear transformations. However, it should be understood that the present disclosure is not limited thereto, but any other suitable method can be used for feature transformation.

[0052] Next, fuse the feature sets f1 to f4 from small to large in the scale dimension to obtain feature_td i . For example, feature_td can be calculated by the following formula i :

[0053] feature_td i = feature_ i + upsampling(feature i+1 ) (1)

[0054] Among them, feature_ i represents the feature leveled to the same feature dimension level, where i = 1, 2,... M-1, and where M represents the number of backbone blocks.

[0055] It should be noted that the upsampling transformation of feature i+1 can be implemented by, for example, one-dimensional convolution or linear transformation.

[0056] Next, fuse the feature sets f1 to f4 from large to small in the scale dimension to obtain feature_bu j . For example, feature_bu can be calculated by the following formula j :

[0057] feature_bu j = feature_ j + downsampling(feature_ j-1 ) (2)

[0058] Among them, feature_ j represents the feature leveled to the same feature dimension level, where j = 2, 3,... M, and where M represents the number of backbone blocks.

[0059] For example, the downsampling of feature_ j-1 can be implemented by one-dimensional convolution or linear transformation.

[0060] In this article, for the sake of convenience of explanation, subscripts i and j are used to represent the feature leveling from small to large and from large to small in dimension N respectively.

[0061] The above introduces a calculation example of feature fusion. However, the present disclosure is not limited thereto, but any existing appropriate method can be used for feature fusion.

[0062] It should be noted that after aligning all the feature sets at once in the feature dimension C, pairwise alignment can be performed step by step in the scale dimension N, or pairwise alignment can be performed across levels. Step-by-step alignment means aligning between two adjacent feature sets, while cross-level alignment means aligning between two non-adjacent feature sets. Taking four feature sets f1 to f4 aligned from smallest to largest in the scale dimension as an example, in the case of step-by-step alignment, f1 can be aligned to f2, f2 can be aligned to f3, and f3 can be aligned to f4, or only f2 can be aligned to f3 without f1 being aligned to f2 and f3 being aligned to f4; in the case of cross-level alignment, f1 can be directly aligned to f3 and f2 can be directly aligned to f4, or only f1 can be aligned to f3 without f2 being aligned to f4. In short, the feature sets to be aligned can be customized according to actual needs.

[0063] It should also be noted that according to actual needs, all the extracted features can be fused, or some of the extracted features can be fused.

[0064] It should be understood that the generated fused features include both the features aligned in the feature dimension and the scale dimension, and the features extracted before alignment.

[0065] By aligning the extracted features in the feature dimension and the scale dimension, the image information becomes more abundant, thus making the prediction results of the classifier module 30 more accurate.

[0066] Return Figure 1 , after obtaining eight fused feature sets F1 to F8 through the feature fusion module 20, the classifiers 11 to 18 of the classifier module 30 respectively predict the eight fused feature sets F1 to F8 after fusion. For example, classifier 11 predicts the feature set F1, classifier 12 predicts the feature set F2, classifier 13 predicts the feature set F3, and so on.

[0067] The classifiers 11 to 18 are trained to minimize the difference between the prediction results of each fused feature aligned from smallest to largest in the scale dimension and the prediction results of all the fused features aligned from largest to smallest, or to minimize the difference between the prediction results of each fused feature aligned from smallest to largest and the prediction results of all the fused features aligned from largest to smallest. For example, the prediction result of classifier 11 for the feature set F1 is compared one by one with all the prediction results of classifiers 15 to 18 for the feature sets F5 to F8, and the difference between the prediction result of the feature set F1 and each of the prediction results of the feature sets F5 to F8 is minimized. For example, the prediction result of classifier 15 for the feature set F5 is compared one by one with all the prediction results of classifiers 11 to 14 for the feature sets F1 to F4, and the difference between the prediction result of the feature set F5 and each of the prediction results of the feature sets F1 to F4 is minimized.

[0068] It should be noted that the results output by classifiers 11 to 18 are the probability distributions of the objects (e.g., products) in the image to be matched belonging to each (product) category, where the number of categories can be predefined or determined according to the number of categories of the images in the image library. The minimum difference between the prediction results means that the probability distributions of the two feature sets are made as consistent as possible.

[0069] For example, classifiers 11 to 18 can be trained using a loss function based on the KL (Kullback-Leibler divergence). Let Y_td i represent the classification result of feature_td i , and Y_bu j represent the classification result of feature_bu j . The dimension of Y_td i is [N i , C gt and the dimension of Y_bu j is [N j , C gt , where C gt is the number of categories used for training.

[0070] The closeness of Y_td i to Y_bu j can be constrained by calculating the difference between Y_td i and Y_bu j . To keep the dimensions the same, Y_td i and Y_bu j represent the average value of this feature in dimension N in the following formula (3). The first loss function for classifiers 11 to 18 can be calculated by the following formula:

[0071]

[0072] where diff(Y_td i , Y_bu j ) can be calculated, for example, by the KL divergence (Kullback-Leibler divergence), α ij is a preset hyperparameter, and i = 1, 2,... M and j = 1, 2,... M, where M represents the number of backbone blocks.

[0073] It should be understood that the above loss function L1 is only an example, and the present disclosure is not limited thereto, but any other suitable loss function can be used as long as the same function can be achieved. For example, the loss function L1 can also be constructed based on the L-2 norm.

[0074] The feature fusion module 20 also provides the fused features to the foreground enhancement module 40. The foreground enhancement module 40 is configured to divide the fused features into foreground features and background features according to a preset rule, and add one class for the background features compared to the foreground features for another classifier during training, where the foreground features are features associated with the category recognition of the object in the image to be matched.

[0075] Specifically, in this embodiment, consider the fused feature feature_bu j , and assume that some elements in feature_bu j belong to the foreground in the image to be matched, while other elements belong to the background. As described above, Y_bu j represents the predicted probability of each element in feature_bu j belonging to each category, and this predicted probability should be high for the true category of the foreground elements and low for the true category of the background elements.

[0076] First, sort the predicted probabilities of each element in feature_bu j on the true value category from large to small. Assume that the elements ranked at the back (1 - γ1) (for example, γ1 = 1 / 4) are likely to be background elements, and the elements ranked in the top γ2 (for example, 1 / 16) are very likely to be foreground elements, where γ2 < γ1.

[0077] Then, it is constrained that the element with the highest predicted probability in the true value category must be judged as the true value. And it is constrained that the average value of the predicted probabilities of the top γ2 elements must be judged as the true value, while the last (1 - γ1) elements should not be judged as any category in the training data (for example, the images in the image library).

[0078] After identifying all the foreground features and background features, add a pseudo-label to all the background features (add one class compared to the foreground features) and input them into the classifier 19 for prediction.

[0079] It should be understood that in this disclosure, the foreground features refer to the features highly related to identifying the category of the object (product) in the image to be matched.

[0080] The second loss function for training the classifier 19 is as follows, which is based on, for example, the cross-entropy loss function:

[0081]

[0082] where y is the true distribution, is the network output distribution, and the total number of categories is n, where n is an integer greater than 1.

[0083] Therefore, the total loss function can be a weighted sum of the first loss function and the second loss function:

[0084] L total = w1L1 + w2L2 (5)

[0085] where w1 and w2 are the weights of the loss functions L1 and L2 respectively, and can be set according to actual needs.

[0086] By minimizing the total loss function L total the classifiers 11 to 19 are trained so that the network parameters of the feature extraction module 10, the feature fusion module 20, the classifier module 30, and the foreground enhancement module 40 are optimized, and in particular, the feature extraction module 10 can extract features highly relevant to the category of the object in the image to be matched.

[0087] Figure 3 Schematically shown is an image matching device 200 according to an embodiment of the present disclosure. The image matching device 200 includes a feature extraction module 10, a feature fusion module 20, and an image matching module 50, wherein the network parameters of the feature extraction module 10 and the feature fusion module 20 are set by training the above-described device 100, that is, the features of the image extracted by the feature extraction module 10 are highly relevant to the object of the category to be recognized in the image.

[0088] Figure 3 The working principles of the feature extraction module 10 and the feature fusion module 20 in Figure 1 and Figure 2 are the same as those described above in combination with

[0089] and will not be elaborated herein.

[0090] It should be noted that before feature stitching, the features to be stitched need to be averaged. For example, taking the stitched feature feature_bu j as an example, what is stitched is the average value of all elements in feature_bu j .

[0091] It should also be noted that any existing distance calculation method can be used to calculate the distance between the stitched feature and the images in the image library 60, such as Euclidean distance, cosine distance, etc.

[0092] It should be understood that the image library 60 may include, for example, images of various products in a supermarket, or may include any images related to category differentiation.

[0093] It should also be understood that all the fused features can be spliced, or only a part of the fused features can be spliced according to actual needs.

[0094] As Figure 4 shown, according to the calculated distance, the image in the image library 60 with the closest distance is used as the first image in the sequence, and so on. Therefore, it can be determined that the category of the object in the image to be matched is the same as the category of the first image.

[0095] The image matching device 200 according to the present disclosure enables the accurate identification of the category of the object in the image to be matched.

[0096] Figure 5 is a flowchart of a method 500 for feature extraction according to an embodiment of the present disclosure. The method 500 is implemented by using the above-described device 100 for feature extraction.

[0097] First, in step 501, the feature extraction module 10 in the device 100 is used to extract features with multiple scales and multiple feature dimensions from the image to be matched respectively.

[0098] Next, in step 502, the feature fusion module 20 is used to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise from small to large and from large to small in the scale dimension, so as to generate fused features, where the fused features include the aligned features and the features before alignment.

[0099] Next, in step 503, multiple classifiers in the classifier module 30 are trained. Each classifier is assigned to a corresponding fused feature and is trained to minimize the difference between the prediction results of each fused feature aligned from large to small in the scale dimension and the prediction results of all fused features aligned from small to large, or to minimize the difference between the prediction results of each fused feature aligned from small to large and the prediction results of all fused features aligned from large to small.

[0100] Finally, in step 504, the foreground enhancement module 40 is used to divide the fused features into foreground features and background features, and in training, for another classifier, one more class is added to the background features compared to the foreground features. The foreground features are the features associated with the category recognition of the object in the image to be matched.

[0101] It should be understood that the specific operation methods of the steps of the method 500 have been described above in combination with Figure 1 and Figure 2The described apparatus 100 has been described in detail and will not be elaborated here.

[0102] Figure 6 FIG. 4 shows a flowchart of an image matching method 600 according to an embodiment of the present disclosure. The method 600 is implemented by using the above-described image matching device 200.

[0103] First, in step 601, the feature extraction module 10 in the image matching device 200 is used to extract features with multiple scales and multiple feature dimensions from the image to be matched respectively. The extracted features are highly correlated with the objects in the image to be matched.

[0104] Next, in step 602, the feature fusion module 20 is used to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise from small to large and from large to small in the scale dimension, thereby generating fused features, where the fused features include the aligned features and the features before alignment.

[0105] Finally, in step 603, all or a part of the fused features are spliced together, the distances between the spliced features and all the images in the image library are calculated, and all the images in the image library are sorted according to the calculation results.

[0106] It should be understood that the specific operation methods of the steps of the method 600 have been described in detail with respect to the above-described image matching device 200 and will not be elaborated here. Figure 3 The described apparatus 100 has been described in detail and will not be elaborated here.

[0107] The above-discussed methods can be fully implemented by computer-executable programs, or can be partially or fully implemented using hardware and / or firmware. When implemented using hardware and / or firmware, or when a computer-executable program is loaded into a hardware device capable of running the program, the apparatus for feature extraction and the image matching device described above are implemented.

[0108] Each component module and unit in the above-described apparatus or device can be configured in a manner of software, firmware, hardware, or a combination thereof. The specific means or manners of configuration are well known to those skilled in the art and will not be elaborated here. In the case of implementation using software or firmware, a program constituting the software is installed from a storage medium or a network into a computer with a dedicated hardware structure (such as Figure 7 the general-purpose computer 700 shown), and when various programs are installed in this computer, it can execute various functions, etc.

[0109] Figure 7 FIG. 5 is a block diagram of an exemplary structure of a general-purpose personal computer in which the methods and / or devices according to the embodiments of the present invention can be implemented. As Figure 7As shown, the central processing unit (CPU) 701 executes various processes according to a program stored in the read-only memory (ROM) 702 or a program loaded from the storage device 708 into the random access memory (RAM) 703. In the RAM 703, data required when the CPU 701 executes various processes and the like is also stored as needed. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. An input / output interface 705 is also connected to the bus 704.

[0110] The following components are connected to the input / output interface 705: an input device 706 (including a keyboard, a mouse, etc.), an output device 707 (including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.), a storage device 708 (including a hard disk, etc.), and a communication device 709 (including a network interface card such as a LAN card, a modem, etc.). The communication device 709 executes communication processing via a network such as the Internet. As needed, a drive 710 may also be connected to the input / output interface 705. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 710 as needed, so that a computer program read therefrom is installed in the storage device 708 as needed.

[0111] In the case where the above-described series of processes are implemented by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 711.

[0112] Those skilled in the art should understand that such a storage medium is not limited to Figure 7 the removable medium 711 shown in which a program is stored and distributed separately from the device to provide the program to the user. Examples of the removable medium 711 include a magnetic disk (including a floppy disk (registered trademark)), an optical disk (including a compact disc read-only memory (CD-ROM) and a digital versatile disc (DVD)), a magneto-optical disk (including a mini disc (MD) (registered trademark)), and a semiconductor memory. Alternatively, the storage medium may be the ROM 702, a hard disk included in the storage device 708, etc., in which a program is stored and distributed to the user together with the device containing them.

[0113] The present disclosure also provides corresponding computer program codes and a computer program product storing machine-readable instruction codes. When the instruction codes are read and executed by a machine, the methods 500 and 600 according to the embodiments of the present disclosure can be executed.

[0114] Accordingly, a storage medium configured to carry the above-described program product storing machine-readable instruction codes is also included in the disclosure of the present disclosure. The storage medium includes but is not limited to a floppy disk, an optical disk, a magneto-optical disk, a memory card, a memory stick, etc.

[0115] Through the above description, the embodiments of the present disclosure provide the following technical solutions, but are not limited thereto.

[0116] Supplementary Note 1. An apparatus for feature extraction, comprising:

[0117] A feature extraction module configured to extract features with multiple scales and multiple feature dimensions from an image to be matched respectively;

[0118] A feature fusion module configured to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise in the scale dimension from small to large and from large to small, thereby generating a fused feature, the fused feature including the aligned features and the features before alignment;

[0119] A classifier module including a plurality of classifiers, each of the plurality of classifiers being assigned to a corresponding fused feature and being trained such that the difference between the prediction results of each fused feature aligned from small to large in the scale dimension and the prediction results of all fused features aligned from large to small is minimized, or such that the difference between the prediction results of each fused feature aligned from large to small and the prediction results of all fused features aligned from small to large is minimized; and

[0120] A foreground enhancement module configured to divide the fused feature into a foreground feature and a background feature according to a preset rule, and add one class to the background feature compared to the foreground feature for another classifier during training, the foreground feature being a feature associated with the category recognition of an object in the image to be matched.

[0121] Supplementary Note 2. The apparatus according to Supplementary Note 1, wherein the feature extraction module includes N backbone blocks, the plurality of scales includes N scales, and each backbone block is configured to extract a feature with one of the N scales, where N is an integer greater than or equal to 2.

[0122] Supplementary Note 3. The apparatus according to Supplementary Note 2, wherein the plurality of classifiers includes N×2 classifiers, the fused feature is divided into N×2 feature sets based on the scale, and the N×2 classifiers are respectively assigned to the N×2 feature sets.

[0123] Supplementary Note 4. The apparatus according to any one of Supplementary Notes 1 to 3, wherein the prediction result is a probability distribution of the object belonging to each of two or more categories.

[0124] Supplement 5. The device according to Supplement 4, wherein the preset rule is to divide the fusion features into the foreground features and the background features based on the ranking on the true value category in the probability distributions output by all classifiers.

[0125] Supplement 6. The device according to any one of Supplements 1 to 3, wherein the multiple classifiers are trained based on a first loss function and a second loss function, and wherein the additional classifier in the foreground enhancement module is trained based on the second loss function.

[0126] Supplement 7. The device according to Supplement 6, wherein the first loss function is the Kullback-Leibler divergence loss function, and wherein the second loss function is the cross-entropy loss function.

[0127] Supplement 8. The device according to Supplement 6, wherein the total loss function is a weighted sum of the first loss function and the second loss function, and wherein the training includes minimizing the total loss function.

[0128] Supplement 9. The device according to any one of Supplements 1 to 3, wherein the features aligned in the feature dimension are pairwise aligned from small to large and from large to small in the scale dimension, and are performed in parallel or sequentially.

[0129] Supplement 10. The device according to Supplement 9, wherein the feature dimension is the length of the extracted features.

[0130] Supplement 11. The device according to any one of Supplements 1 to 3, wherein the scale is obtained based on the resolution of the image to be matched.

[0131] Supplement 12. A method for feature extraction, comprising:

[0132] Extracting features with multiple scales and multiple feature dimensions from the image to be matched respectively;

[0133] Aligning all or part of the extracted features in the feature dimension, and then pairwise aligning the features aligned in the feature dimension from small to large and from large to small in the scale dimension, thereby generating fusion features, the fusion features including the aligned features and the features before alignment;

[0134] Training multiple classifiers, each of the multiple classifiers being assigned to a corresponding fusion feature and being trained such that the difference between the prediction results of each fusion feature aligned from small to large in the scale dimension and the prediction results of all fusion features aligned from large to small is minimized, or such that the difference between the prediction results of each fusion feature aligned from large to small and the prediction results of all fusion features aligned from small to large is minimized; and

[0135] According to a preset rule, divide the fusion feature into a foreground feature and a background feature, and add one class for the background feature compared to the foreground feature for another classifier during the training, where the foreground feature is a feature associated with the class recognition of the object in the image to be matched.

[0136] Supplementary Note 13. According to the method of Supplementary Note 12, wherein the prediction result is a probability distribution of the object belonging to each of two or more classes.

[0137] Supplementary Note 14. According to the method of Supplementary Note 12, wherein the preset rule is to divide the fusion feature into the foreground feature and the background feature based on the ranking on the true value class in the probability distributions output by all classifiers.

[0138] Supplementary Note 15. According to the method of any one of Supplementary Notes 12 to 14, wherein the plurality of classifiers are trained based on a first loss function and a second loss function, and wherein the another classifier is trained based on the second loss function.

[0139] Supplementary Note 16. According to the method of Supplementary Note 15, wherein the first loss function is a Kullback-Leibler divergence loss function, and wherein the second loss function is a cross-entropy loss function,

[0140] wherein the total loss function is a weighted sum of the first loss function and the second loss function, and

[0141] wherein the training includes minimizing the total loss function.

[0142] Supplementary Note 17. An image matching device, comprising:

[0143] A feature extraction module configured to extract features with multiple scales and multiple feature dimensions from an image to be matched, wherein the extracted features are related to an object in the image to be matched;

[0144] A feature fusion module configured to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise in the scale dimension from small to large and from large to small, so as to generate a fusion feature, where the fusion feature includes the aligned features and the features before alignment; and

[0145] A matching module configured to splice all or a part of the fusion features together, calculate the distances between the spliced features and all images in an image library, and sort all images in the image library according to the calculation results,

[0146] Among them, the network parameters of the feature extraction module and the feature fusion module are set by training according to the device described in any one of Appendices 1 to 11.

[0147] Appendix 18. An image matching method, comprising:

[0148] Extracting features with multiple scales and multiple feature dimensions from the image to be matched, wherein the extracted features are related to the object in the image to be matched;

[0149] Aligning all or part of the extracted features in the feature dimension, and then aligning the features aligned in the feature dimension pairwise from small to large and from large to small in the scale dimension, so as to generate fused features, the fused features including the aligned features and the features before alignment; and

[0150] Concatenating all or a part of the fused features together, calculating the distances between the concatenated features and all the images in the image library, and sorting all the images in the image library according to the calculated results,

[0151] wherein the extraction and fusion of the features are performed by the feature extraction module and the feature fusion module in the image matching device according to Appendix 17.

[0152] Appendix 19. A computer-readable storage medium storing a program, which when executed by a processor causes a computer to execute the method according to any one of Appendices 12 to 16.

[0153] Appendix 20. A computer-readable storage medium storing a program, which when executed by a processor causes a computer to execute the method according to Appendix 18.

[0154] Finally, it should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. In addition, without more limitations, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.

[0155] Although the embodiments of the present disclosure have been described in detail above in conjunction with the accompanying drawings, it should be understood that the above-described embodiments are only configured to illustrate the present disclosure and do not constitute a limitation to the present disclosure. Those skilled in the art can make various modifications and changes to the above embodiments without departing from the essence and scope of the present disclosure. Therefore, the scope of the present disclosure is only defined by the appended claims and their equivalent meanings.

Claims

1. An apparatus for feature extraction, comprising: A feature extraction module configured to extract features having multiple scales and multiple feature dimensions from an image to be matched respectively; A feature fusion module configured to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise in the scale dimension from small to large and from large to small, thereby generating fused features, the fused features including the aligned features and the features before alignment; A classifier module including multiple classifiers, each of the multiple classifiers being assigned to a corresponding fused feature and being trained such that the difference between the prediction result of each fused feature aligned from small to large in the scale dimension and the prediction results of all fused features aligned from large to small is minimized, or such that the difference between the prediction result of each fused feature aligned from small to large and the prediction results of all fused features aligned from large to small is minimized; And A foreground enhancement module configured to divide the fused features into foreground features and background features according to a preset rule, and add one class to the background features compared to the foreground features for another classifier during the training, the foreground features being features associated with the category recognition of the object in the image to be matched.

2. The device according to claim 1, wherein The feature extraction module includes N backbone blocks, the multiple scales include N scales, and each backbone block is configured to extract features having one of the N scales, where N is an integer greater than or equal to 2.

3. The device according to claim 2, wherein, The multiple classifiers include N×2 classifiers, the fused features are divided into N×2 feature sets based on the scale, and the N×2 classifiers are respectively assigned to the N×2 feature sets.

4. The device according to any one of claims 1 to 3, wherein The prediction result is the probability distribution of the object belonging to each of two or more categories.

5. The device according to claim 4, wherein The preset rule is to divide the fused features into the foreground features and the background features based on the ranking on the true category in the probability distributions output by all classifiers.

6. The device according to any one of claims 1 to 3, wherein The multiple classifiers are trained based on a first loss function and a second loss function, and wherein the additional classifier in the foreground enhancement module is trained based on the second loss function, and wherein the total loss function is the weighted sum of the first loss function and the second loss function, and wherein the training includes minimizing the total loss function.

7. The apparatus according to any one of claims 1 to 3, wherein The feature dimension is the length of the extracted features, and wherein the scale is obtained based on the resolution of the image to be matched.

8. An image matching device, comprising: A feature extraction module configured to extract features having multiple scales and multiple feature dimensions from an image to be matched respectively, wherein the extracted features are related to the object in the image to be matched; A feature fusion module configured to align all or part of the extracted features in the feature dimension, and then align the features aligned in the feature dimension pairwise in the scale dimension from small to large and from large to small, thereby generating fused features, the fused features including the aligned features and the features before alignment; and A matching module configured to splice together all or part of the fusion features, calculate the distances between the spliced features and all the images in the image library, and sort all the images in the image library according to the calculation results. Wherein, the network parameters of the feature extraction module and the feature fusion module are set by training according to the device described in any one of claims 1 to 7.

9. A computer-readable storage medium storing a program, which when executed by a processor causes the computer to execute a method for feature extraction, the method comprising: Extracting features with multiple scales and multiple feature dimensions from the image to be matched respectively; Aligning all or part of the extracted features in the feature dimension, and then aligning the features aligned in the feature dimension pairwise from small to large and from large to small in the scale dimension, so as to generate fusion features, where the fusion features include the aligned features and the features before alignment; Training a plurality of classifiers, each of the plurality of classifiers being assigned to a corresponding fusion feature and being trained such that the difference between the prediction results of each fusion feature aligned from small to large in the scale dimension and the prediction results of all fusion features aligned from large to small is minimized, or such that the difference between the prediction results of each fusion feature aligned from large to small and the prediction results of all fusion features aligned from small to large is minimized; and According to a preset rule, dividing the fusion features into foreground features and background features, and adding one more class to the background features than the foreground features for another classifier during the training, where the foreground features are features associated with the category recognition of the object in the image to be matched.

10. A computer-readable storage medium storing a program, which when executed by a processor causes the computer to execute an image matching method, the image matching method comprising: Extracting features with multiple scales and multiple feature dimensions from the image to be matched respectively, wherein the extracted features are related to the object in the image to be matched; Aligning all or part of the extracted features in the feature dimension, and then aligning the features aligned in the feature dimension pairwise from small to large and from large to small in the scale dimension, so as to generate fusion features, where the fusion features include the aligned features and the features before alignment; and Splicing together all or part of the fusion features, calculating the distances between the spliced features and all the images in the image library, and sorting all the images in the image library according to the calculation results. Wherein, the extraction and fusion of the features are performed by the feature extraction module and the feature fusion module in the image matching device according to claim 8.