A sampling-independent full-metric small-sample target detection method

By introducing a full-metric sample detection model and spatial attention mechanism in small sample object detection, the problem of existing methods ignoring spatial position information and sampling instability is solved, and more robust and consistent detection results are achieved.

CN115240008BActive Publication Date: 2025-05-09CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210931306.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-05-09
Estimated Expiration
2042-08-04

AI Technical Summary

Technical Problem

The existing small sample object detection method based on metric learning ignores spatial position information, and the sampling results of small samples are unstable, making it easy to sample noise data, resulting in feature prototype deviations, which in turn affects the consistency of the model's results on different sampling distributions.

Method used

A sampling-independent full-metric small sample object detection method is proposed. By constructing a full-metric sample detection model, including feature offset correction module, full-metric module, multi-scale matching module, regional proposal network module and residual block, the spatial position relationship is encoded using the spatial attention mechanism, and constrain the network through multi-scale feature matching and self-supervised learning strategies, a more robust prototype vector is generated.

Benefits of technology

By introducing spatial location information and self-supervised learning strategies, the matching suboptimal results caused by scale differences are reduced, the robustness of the model to noise data and the consistency of the results are improved, and the application scenarios of the model are expanded.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240008B_ABST
    Figure CN115240008B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of artificial intelligence technology, and specifically relates to a sampling-independent full-metric small sample target detection method, including: constructing a full-metric sample detection model, and fine-tuning the full-metric sample detection model; sampling small sample data, organizing and dividing the small sample data set, obtaining a class support set and a query set, and pre-processing the class support samples in the class support set; inputting the small sample data into the fine-tuned full-metric sample detection model, performing target detection and obtaining detection results. The present invention reduces suboptimal matching results caused by scale differences by using cross-scale semantic matching; by constructing a set of normal and damaged image pairs, a self-supervised learning strategy is used to constrain the network so that the encoder can use the context to build a more robust prototype; spatial position information is added to the prototype vector to guide the model to capture the target more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a sampling-independent full-metric small-sample target detection method. Background Art

[0002] Thanks to the continuous development of computing power and big data, large-sample deep learning solutions have made great progress and have reached human-level performance in some areas (such as classification). Due to issues such as privacy, data long-tail distribution, and labor costs, it is not realistic to collect a large sample data set in many cases, and small samples can cause conventional deep learning solutions to easily fall into overfitting.

[0003] In recent years, small sample target detection has received widespread attention. Its purpose is to efficiently learn a model that approximates large sample learning from only a small number of labeled samples. Inspired by the fact that humans can learn quickly with the help of historical experience, researchers at home and abroad will use an additional dataset that does not contain small sample category annotations to pre-train the model, providing the model with general knowledge / experience so that it can learn more efficiently in subsequent small sample learning tasks. In order to obtain better knowledge, the current conventional solutions include metric learning algorithms, transfer learning algorithms, and data augmentation algorithms. Compared with other solutions, the metric learning solution first constructs the additional dataset into a set of small sample tasks to learn general knowledge, namely the connection between category prototypes and feature maps, and then fine-tunes the model on the target small sample task.

[0004] Due to design reasons, the metric-based scheme relies heavily on category prototypes to classify targets. If the prototype quality is poor, noise will be introduced, resulting in poor recognition results. Existing schemes mainly require feature prototypes to have the characteristics of "high cohesion and low coupling" from the perspective of inductive spatial constraints. Such constraints will implicitly lose spatial information, which is very important for detection tasks.

[0005] In addition, according to statistical laws, the sampling results of small samples are quite unstable. It is very easy to sample noise data and generate biased feature prototypes, which makes the results of the current model very different in different sampling distributions, greatly limiting its application scenarios. To solve this problem, some methods introduce technologies such as feature pyramids and affinity matrices to perform feature matching at a fine-grained level, but ignore the potential scale mismatch problem.

[0006] In summary, the existing technical problems are: the current metric-based scheme ignores spatial location information and the sampling results of small samples are unstable. It is very easy to sample noise data and generate biased feature prototypes, which makes the current model have large differences in results under different sampling distributions, greatly limiting its application scenarios. Summary of the invention

[0007] In order to solve the above technical problems, the present invention proposes a sampling-independent full-metric small sample target detection method, comprising the following steps:

[0008] S1: Build a full-metric sample detection model and fine-tune the full-metric sample detection model;

[0009] The full metric sample detection model includes: a feature offset correction module, a full metric module, a multi-scale matching module, a region proposal network module, and a residual block;

[0010] S2: Sampling small sample data, organizing and dividing the small sample data to obtain the class support set S i and the query set Q i , and the class support set S is corrected by the feature shift correction module i The classes in support samples for preprocessing;

[0011] S3: Preprocessed class support samples and query set Q i Extract features from the query samples in the class and get the class prototype vector and query image feature map

[0012] S4: Class prototype vector Input the full metric module, encode the spatial position relationship through the spatial attention mechanism, and obtain the enhanced prototype vector

[0013] S5: Enhanced prototype vector and query image feature map Input into the multi-scale matching module, perform multi-scale feature matching, and obtain the semantic feature X that matches accurately in scale ext ;

[0014] S6: Semantic feature X ext Input region proposal network module to generate candidate boxes;

[0015] S7: Using candidate boxes from query image feature maps Extract candidate target features;

[0016] S8: Use the residual block to transform the class prototype vector Mapped to the target feature space, we get the class prototype vector pair

[0017] S9: Each pair of class prototype vectors Match with the candidate target features to obtain the final detection result.

[0018] Preferably, the full metric sample detection model is fine-tuned by a loss function, and the loss function expression is:

[0019]

[0020] in, Represents the loss of classification and positioning using the Faster RCNN framework; Represents class prototype sample I and reconstruction sample I pred The mean square error of .

[0021] Preferably, the small sample data is organized and divided, specifically including: according to the acquired small sample data, the pictures containing the labeled information in the small sample data are divided into class support sets S i , divide the pictures without annotation information into query set Q i .

[0022] Preferably, the class support set S is corrected by a feature shift correction module. i Preprocessing is performed, and the feature offset correction module includes a data augmenter, a feature encoder, and a decoder: the prototype sample I is input into the data augmenter, a part of the regional features of the prototype sample is randomly masked to form a noise sample I', the noise sample I' is feature encoded by the feature encoder, and the decoder is used to decode and reconstruct the feature-encoded noise sample to obtain a reconstructed sample I pred , calculate the prototype sample I and the reconstruction sample I pred The sample with the smallest mean square error is selected as the class support sample after feature offset correction, and the context area of ​​the target image in the class support sample is adaptively confirmed to reduce the interference of noisy background on the generated class prototype vector.

[0023] Furthermore, the prototype sample I and the reconstructed sample I are calculated. pred The mean square error of , the expression includes:

[0024]

[0025] in, Represents class prototype sample I and reconstruction sample I pred The mean square error of I represents the prototype sample; I pred Represents the reconstructed sample.

[0026] Furthermore, the context area of ​​the target image in the class support sample is adaptively confirmed to reduce the interference of the noisy background on the generated class prototype vector. The expression includes:

[0027]

[0028] Among them, w, h and γ represent the width, height and scaling ratio of the target image respectively, and δ represents a hyperparameter that limits the maximum proportion of the target context area.

[0029] Preferably, the S4 specifically includes: the full metric module includes: a convolution layer, a BN layer, a nonlinear mapping layer, and a sigmoid activation layer:

[0030] S41: Class prototype vector Perform average pooling operations on the horizontal X and vertical Y axes to obtain directional features and

[0031] S42: Directional features and Fuse together and input into a convolutional layer, a BN layer, and a nonlinear mapping layer to capture the relationship between different positions;

[0032] S43: Directional features are transformed according to the position relationship and Input into a convolution layer and a sigmoid activation layer respectively to obtain position features in different directions;

[0033] S44: Based on the coordinate attention mechanism, matrix multiplication is performed on the position features in different directions to obtain the coordinate attention weights, and the attention weights are combined with the class prototype vectors. Do the dot product and the class prototype vector Perform residual operation to obtain the enhanced prototype vector

[0034] Preferably, the S5 specifically includes:

[0035] S51: Enhanced prototype vector and query image feature map Input into a 1*1 convolutional layer and project into the same feature space to obtain k and v features respectively;

[0036] S52: Perform matrix multiplication operation on the support set k feature and the query set v feature to obtain the mutual relationship between the two features;

[0037] S53: Use the softmax function to normalize the weights according to the mutual relationship between the two features, and add them to the enhanced prototype vector Do a dot multiplication operation and then query the image feature map After the dot multiplication, the residual operation is performed on the support set features to obtain the semantic feature X that matches the scale exactly. ext .

[0038] Beneficial effects of the present invention:

[0039] (1) By using cross-scale semantic matching to reduce and Differences in scale lead to suboptimal matching results;

[0040] (2) By constructing a set of normal and damaged image pairs to correct feature offsets, a self-supervised learning strategy is used to constrain the network so that the encoder can use the context to build a more robust prototype;

[0041] (3) Spatial position information is added to the prototype vector to guide the model to capture the target more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a framework diagram of the full-metric sample detection model of the present invention;

[0043] Figure 2 This is a structural diagram of the characteristic offset correction module of the present invention;

[0044] Figure 3 This is a structural diagram of the coordinate attention module of the present invention;

[0045] Figure 4 This is a structural diagram of the multi-scale matching module of the present invention. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] A sampling-independent full-metric small-sample target detection method comprises the following steps:

[0048] S1: Build a full-metric sample detection model and fine-tune the full-metric sample detection model;

[0049] The full metric sample detection model, such as Figure 1 As shown, it includes: feature offset correction module, full metric module, multi-scale matching module, region proposal network module, residual block;

[0050] S2: Sampling small sample data, organizing and dividing the small sample data to obtain the class support set S i and the query set Q i , and the class support set S is corrected by the feature shift correction module i The classes in support samples for preprocessing;

[0051] S3: Preprocessed class support samples and query set Q i Extract features from the query samples in the class and get the class prototype vector and query image feature map

[0052] S4: Class prototype vector Input the full metric module, encode the spatial position relationship through the spatial attention mechanism, and obtain the enhanced prototype vector

[0053] S5: Enhanced prototype vector and query image feature map Input into the multi-scale matching module, perform multi-scale feature matching, and obtain the semantic feature X that matches accurately in scale ext ;

[0054] S6: Semantic feature X ext Input region proposal network module to generate candidate boxes;

[0055] S7: Using candidate boxes from query image feature maps Extract candidate target features;

[0056] S8: Use the residual block to transform the class prototype vector Mapped to the target feature space, we get the class prototype vector pair

[0057] S9: Each pair of class prototype vectors Match the features of the candidate targets to obtain the final detection result, that is, the category of the small sample target and the position of the target in the entire image.

[0058] Preferably, the full metric sample detection model is fine-tuned through the loss function to guide the model to learn better parameters. The loss function expression is:

[0059]

[0060] in, Represents the loss of classification and positioning using the Faster RCNN framework; Represents class prototype sample I and reconstruction sample I pred The mean square error of .

[0061] Preferably, obtaining small sample data includes: randomly selecting K={10,30} from 20 classes of data overlapping with PASCAL VOC in the COCO 2017 dataset as small sample data, performing N-way K-shot task division on the small sample data, randomly selecting one image from each category as a support image, and randomly selecting one from the remaining images as a query image, wherein the annotation information on the support image is known, while the annotation information on the query image is unknown; all support images constitute a class support set S i , all query images constitute the query set Q i .

[0062] Furthermore, the small sample data is organized and divided, specifically including: according to the acquired small sample data, the pictures containing the labeled information in the small sample data are divided into the class support set S i , divide the pictures without annotation information into query set Q i The class support set has annotation information, and we can extract the target features based on this annotation information. The query set is an unlabeled image (i.e., test image). We extract the target features from the class support set to find all class instances in the query set.

[0063] Preferably, Figure 2 This is a diagram of the semantic shift correction module. First, a Random Object Transformer (essentially a data augmenter, random mask or some other operation) is used to generate a damaged image, which is then fed into an encoder-decoder structure to restore the original image (with the original size). Figure 4 We expect the restored image to be similar to the original image. Figure 1 That is, the MSE loss should be as small as possible (MSE-mean square error).

[0064] Furthermore, the class support set S is corrected by the feature shift correction module. i Preprocessing is performed, and the feature offset correction module includes a data augmenter, a feature encoder, and a decoder: the prototype sample I is input into the data augmenter, a part of the regional features of the prototype sample is randomly masked to form a noise sample I', the noise sample I' is feature encoded by the feature encoder, and the decoder is used to decode and reconstruct the feature-encoded noise sample to obtain a reconstructed sample I pred , calculate the prototype sample I and the reconstruction sample I pred The mean square error of the target image is calculated, and the sample with the smallest mean square error is selected as the class support sample after feature offset correction. The smaller the mean square error, the better the feature offset correction effect. The context area of ​​the target image in the class support sample is adaptively confirmed to reduce the interference of noisy background on the generated class prototype vector.

[0065] Furthermore, the prototype sample I and the reconstructed sample I are calculated. pred The mean square error of , the expression includes:

[0066]

[0067] in, Represents class prototype sample I and reconstruction sample I pred The mean square error of I represents the prototype sample; I pred Represents the reconstructed sample.

[0068] Furthermore, the context area of ​​the target image in the class support sample is adaptively confirmed to reduce the interference of the noisy background on the generated class prototype vector. The expression includes:

[0069]

[0070] Among them, w, h and γ are the width, height and scaling ratio of the target, and δ is a hyperparameter used to limit the maximum proportion of the target context area.

[0071] Preferably, the full metric module includes a residual-based coordinate attention module: the coordinate attention mechanism is used to obtain the class prototype vector Add a spatial position relationship, the coordinate attention module, such as Figure 3 As shown, the full metric module includes: a convolutional layer, a BN layer, a nonlinear mapping layer, and a sigmoid activation layer.

[0072] Furthermore, for the class prototype vector Adding a spatial position relationship specifically includes:

[0073] S41: Class prototype vector Perform average pooling operations on the horizontal X and vertical Y axes to obtain directional features and

[0074] S42: Directional features and Fuse together and input into a convolutional layer, a BN layer, and a nonlinear mapping layer to capture the relationship between different positions;

[0075] S43: Directional features are transformed according to the position relationship and Input into a convolution layer and a sigmoid activation layer respectively to obtain position features in different directions;

[0076] S44: Based on the coordinate attention mechanism, matrix multiplication is performed on the position features in different directions to obtain the coordinate attention weights, and the attention weights are combined with the class prototype vectors. Do the dot product and the class prototype vector Perform residual operation to obtain the enhanced prototype vector

[0077] Furthermore, the coordinate attention value mechanism embeds position information into channel attention, thereby enabling the mobile network to obtain information of a larger area without introducing large overhead. In order to avoid the loss of position information introduced by 2D global pooling, the coordinate attention mechanism proposes to decompose channel attention into two parallel 1D feature encodings to efficiently integrate spatial coordinate information into the generated attention maps. Specifically, two 1D global pooling operations are used to aggregate the input features along the vertical and horizontal directions into two separate direction-aware feature maps. The two feature maps with embedded specific direction information are then encoded into two attention maps respectively. Each attention map captures the long-range dependencies of the input feature map along a spatial direction. The position information can therefore be saved in the generated attention map. Both attention maps are then applied to the input feature maps by multiplication to emphasize the representation of the attention area.

[0078] Preferably, the multi-scale matching module, such as Figure 4 As shown: The enhanced prototype vector and query image feature map Projected into the same semantic space through Conv convolution, the prototype features are obtained and query image features Calculate the similarity relationship of feature matching at different scales through matrix multiplication + softmax Use similarity relationship to match the most appropriate semantic feature X in scale and space ext ; Where R represents the real number space, w, h, d, w', h'd' will change according to the actual input image size, W i , H i Represents the width and height of the query set feature graph i; Indicates image depth; W j ′、H j ′ represents the width and height of the support set feature map j; n represents the number of feature images in the support set; m represents the number of feature images in the query set.

[0079] Furthermore, multi-scale feature matching specifically includes:

[0080] S51: Enhanced prototype vector and query image feature map Input into a 1*1 convolutional layer and project into the same feature space to obtain k and v features respectively;

[0081] S52: Perform matrix multiplication operation on the support set k feature and the query set v feature to obtain the mutual relationship between the two features;

[0082] S53: Use the softmax function to normalize the weights according to the mutual relationship between the two features, and add them to the enhanced prototype vector Do a dot multiplication operation and then query the image feature map After the dot multiplication, the residual operation is performed on the support set features to obtain semantic features that match the scale precisely.

[0083] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A sampling-independent full-metric small-sample target detection method, characterized in that: The following steps are involved: S1: Build a full-metric sample detection model and fine-tune the full-metric sample detection model; The full metric sample detection model includes: a feature offset correction module, a full metric module, a multi-scale matching module, a region proposal network module, and a residual block; S2: Sampling small sample data, organizing and dividing the small sample data to obtain the class support set S i and the query set Q i , and the class support set S is corrected by the feature shift correction module i The classes in support samples for preprocessing; The class support set S is corrected by the feature shift correction module i Preprocessing is performed, and the feature offset correction module includes a data augmenter, a feature encoder, and a decoder: the prototype sample I is input into the data augmenter, and a part of the regional features of the prototype sample is randomly masked to form a noise sample I ′ , through the feature encoder to the noise sample I ′ Perform feature encoding, and use a decoder to decode and reconstruct the noise sample after feature encoding to obtain the reconstructed sample I pred , calculate the prototype sample I and the reconstruction sample I pred The mean square error of the image is calculated, and the sample with the smallest mean square error is selected as the class support sample after feature offset correction. The context area of ​​the target image in the class support sample is adaptively confirmed to reduce the interference of noisy background on the generated class prototype vector. S3: Preprocessed class support samples and query set Q i Extract features from the query samples in the class and get the class prototype vector and query image feature map S4: Class prototype vector Input the full metric module, encode the spatial position relationship through the spatial attention mechanism, and obtain the enhanced prototype vector The full metric module includes: convolution layer, BN layer, nonlinear mapping layer, sigmoid activation layer: S41: class prototype vector F Si Perform average pooling operations on the horizontal X and vertical Y axes to obtain directional features and S42: Directional features and Fuse together and input into a convolutional layer, a BN layer, and a nonlinear mapping layer to capture the relationship between different positions; S43: Directional features are transformed according to the position relationship and Input into a convolution layer and a sigmoid activation layer respectively to obtain position features in different directions; S44: Based on the coordinate attention mechanism, matrix multiplication is performed on the position features in different directions to obtain the coordinate attention weights, and the attention weights are combined with the class prototype vectors. Do the dot product and the class prototype vector Perform residual operation to obtain the enhanced prototype vector S5: Enhanced prototype vector and query image feature map Input into the multi-scale matching module, perform multi-scale feature matching, and obtain the semantic feature X that matches accurately in scale ext ; S6: Semantic feature X ext Input region proposal network module to generate candidate boxes; S7: Using candidate boxes from query image feature maps Extract candidate target features; S8: Use the residual block to transform the class prototype vector Mapped to the target feature space, we get the class prototype vector pair S9: Each pair of class prototype vectors Match with the candidate target features to obtain the final detection result.

2. The sampling-independent full-metric small-sample target detection method according to claim 1, characterized in that: The full metric sample detection model is fine-tuned by the loss function, and the loss function expression is: in, Represents the loss of classification and positioning using the Faster RCNN framework; Represents class prototype sample I and reconstruction sample I pred The mean square error of .

3. The sampling-independent full-metric small-sample target detection method according to claim 1, characterized in that: Organize and divide the small sample data, specifically including: according to the obtained small sample data, divide the pictures containing annotation information in the small sample data into class support sets S i , divide the pictures without annotation information into query set Q i .

4. The sampling-independent full-metric small sample target detection method according to claim 1, characterized in that: Computational prototype sample I and reconstruction sample I pred The mean square error of , the expression includes: in, Represents class prototype sample I and reconstruction sample I pred The mean square error of I represents the prototype sample; I pred Represents the reconstructed sample.

5. The sampling-independent full-metric small-sample target detection method according to claim 1, characterized in that: And adaptively confirm the context area of ​​the target image in the class support sample to reduce the interference of noisy background on the generated class prototype vector. The expressions include: Among them, w, h and γ represent the width, height and scaling ratio of the target image respectively, and δ represents a hyperparameter that limits the maximum proportion of the target context area.

6. The sampling-independent full-metric small sample target detection method according to claim 1, characterized in that: The S5 specifically includes: S51: Enhanced prototype vector and query image feature map Input into a 1*1 convolutional layer and project into the same feature space to obtain k and v features respectively; S52: Perform matrix multiplication operation on the support set k feature and the query set v feature to obtain the mutual relationship between the two features; S53: Use the softmax function to normalize the weights according to the mutual relationship between the two features, and add them to the enhanced prototype vector Do a dot multiplication operation and then query the image feature map After the dot multiplication, the residual operation is performed on the support set features to obtain the semantic feature X that matches the scale exactly. ext .