A few-shot object image detection method based on meta-learning framework
By applying sample normalization and Z-Score normalization methods in the meta-learning framework, the problem of insufficient generalization and robustness in the detection of few-sample target images is solved, and the performance stability and accuracy of the detection are improved.
Patent Information
- Application Number
- CN202410009600.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-01-03
AI Technical Summary
The prior art has problems of insufficient generalization and robustness in the detection of few-sample target images, especially inconsistent performance on different data and prone to forgetting problems.
A small sample target image detection method based on a meta-learning framework is adopted to reduce the data volume gap between the training data and the test data by applying a sample normalization method in the support image branch, and use Z-Score normalization before final detection to avoid the impact of overcooling and overheating spots.
The performance stability and generalization ability of the detection of small-sample target image is improved, the gap between random small-sample training data in new image data is reduced, and the impact of supercooling and overheating points is reduced in the high-dimensional feature space, thereby improving the accuracy of detection.
Smart Images

Figure CN117830763B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection, and in particular to a few-sample target image detection method based on a meta-learning framework. Background Art
[0002] Object image detection is a fundamental task in computer vision and has made remarkable progress with the help of large-scale annotated datasets and well-designed powerful detectors. However, in practical applications, the frequent occurrence of few-shot object image scenes limits the capabilities of detectors, making few-shot object image detection a key method to bring object image detection from theoretical research to practical applications and deployments, which aims to solve this task with a small number of training samples.
[0003] In recent years, researchers have proposed many few-shot object image detection methods to improve the generalization ability of neural networks, which can be mainly divided into transfer learning-based methods and meta-learning-based methods. Transfer learning-based methods aim to train suitable network parameters for invariant representation, share features across domains, and focus on how to freeze fewer detector components without reducing performance. Consistent with the general object detection framework, transfer learning-based methods provide a simplified training process without complex training procedures. However, these methods are limited by the few-shot setting and require a mapping from the source domain to the target domain, which makes them lack generalization and robustness and face serious overfitting and forgetting problems.
[0004] In contrast, meta-learning usually consists of two branches, called the support set and the query set, and pays more attention to aggregating the information of these two branches to obtain the ability to "learn to learn". Compared with the transfer learning-based methods, it has better robustness and generalization ability. However, due to the few-shot setting, the meta-learning-based methods usually cannot achieve consistent performance on different data and often encounter the forgetting problem. At the same time, the detection of meta-learning-based methods is unstable in the case of few samples, and the detection performance varies greatly. In addition, in the few-shot target image detection task, the image is usually regarded as a three-dimensional matrix, and feature extraction and mapping are performed through a series of networks, and the relevant features are aggregated in the high-dimensional feature space to complete the final detection task. In this process, the high-dimensional feature space makes the vector susceptible to the influence of super-cold points and super-hot points, that is, in the high-dimensional feature space, the K nearest neighbors of many points will gather at some specific points, which will reduce the accuracy of target image detection. Summary of the invention
[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a few-sample target image detection method based on a meta-learning framework, in which a sample normalization method is applied in the support image branch to obtain a relatively consistent input, so as to reduce the data volume gap between the training data and the test data and the data volume gap between the base class images and the new class images in the meta-fine-tuning process, and Z-Score normalization is applied before the final detection to avoid the influence of overcold points and overhot points.
[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:
[0007] A few-shot target image detection method based on a meta-learning framework comprises the following steps:
[0008] S1, obtaining base class image data and new class image data;
[0009] S2, constructing a few-shot target image detection network model based on a meta-learning framework, meta-training the few-shot target image detection network model using the base class image data in step S1, and meta-fine-tuning the meta-trained few-shot target image detection network model using the base class image data and the new class image data in step S1;
[0010] S3. Obtain a few sample image data to be detected, and determine the image detection result according to the few sample image data to be detected and the few sample target image detection network model after the meta-adjustment in step S2.
[0011] Further, in step S2, the few-shot target image detection network model constructed based on the meta-learning framework includes a support image branch and a query image branch;
[0012] One end of the support image branch is connected to the query image branch, the support image branch is used to receive the support image, perform feature extraction, position encoding, task encoding and normalization on the support image to obtain a normalized vector of the feature vector of the support image, and transmit the normalized vector of the feature vector of the support image to the query image branch;
[0013] One end of the query image branch is connected to the support image branch, the query image branch is used to receive the query image, perform feature extraction and position encoding on the query image to obtain a feature vector of the query image, receive the normalized vector of the feature vector of the support image transmitted by the support image branch, aggregate the feature vector of the query image and the normalized vector of the feature vector of the support image, and use the inter-class correlation to obtain the query features of the support image enhancement, perform Z-Score normalization and linear projection on the query features of the base class support image enhancement to obtain the detection head input vector of the base class query image, and detect the detection head input vector of the base class query image to obtain the detection result of the base class query image.
[0014] Furthermore, the supported image branch includes a second feature extractor, a second Transformer encoder, a first classification head and a sample normalization module; the second feature extractor is connected to the second Transformer encoder, the second Transformer encoder is respectively connected to the first classification head and the sample normalization module, and the sample normalization module is connected to the query image branch.
[0015] Furthermore, the sample normalization module is used to normalize the feature vector of the support image. The calculation formula of the sample normalization module is expressed as:
[0016]
[0017] Where: s′ k is the normalized vector of the feature vector of the support image, k is the element number of the feature vector of the support image, s k is the eigenvector of the support image, μ is the mean of the eigenvector of the support image, and σ is the variance of the eigenvector of the support image.
[0018] Furthermore, the query image branch includes a first feature extractor, a first Transformer encoder, a first Transformer decoder, a Z-Score normalization module and a detection head connected in sequence, and the first Transformer encoder is connected to the support image branch.
[0019] Furthermore, the Z-Score normalization module is used to perform Z-Score normalization and linear projection on the query features supporting image enhancement. The calculation formula of the Z-Score normalization module is expressed as:
[0020]
[0021] Where: Z is the feature vector after Z-Score normalization, W is the linear layer parameter, D is the enhanced query feature, μ z To query the mean of feature D in all dimensions, σ z To query the variance of feature D in all dimensions, B is the linear layer parameter, R Η It is an H-dimensional feature space, and H is the dimension of the query feature D.
[0022] Furthermore, in step S2, the base class image data in step S1 is used to meta-train the few-sample target image detection network model, and the specific process is: randomly sampling the base class image data in step S1 to obtain a base class support image, and selecting a base class query image from the base class image data in step S1; the support image branch obtains a normalized vector of the feature vector of the base class support image and a classification result of the base class support image according to the base class support image; the classification result of the base class support image is used to calculate a first loss function to optimize the internal parameters of the support image branch; the query image branch obtains the detection result of the base class query image according to the base class query image and the normalized vector of the feature vector of the base class support image, and the detection result of the base class query image is used to calculate a second loss function to optimize the internal parameters of the query image branch.
[0023] Furthermore, the second loss function is calculated using the detection results of the base class query image, which is expressed as:
[0024]
[0025] in: For test results The loss function between y and the true result y′, where y′ is the true result of the query image. is the detection result of the query image, i is the input support image number, N is the number of input support images, is the true category label c′ i and the detected category labels sigmoidfocalloss,c′ i To query the true category label of the image, is the category label detected by the query image, is the actual positioning result b′ i The positioning results of the detection L 1 The linear combination of loss and GIoUloss, b′ i To query the real positioning result of the image, The detection and positioning results of the query image.
[0026] Furthermore, in step S2, the base class image data and the new class image data in step S1 are used to meta-fine-tune the meta-trained few-sample target image detection network model. Specifically, images are selected from the base class image data and combined with images in the new class image data to form a new data set, and the new data set is used to meta-fine-tune the meta-trained few-sample target image detection network model.
[0027] The present invention has the following beneficial effects:
[0028] (1) The present invention considers two gaps caused by the few-sample setting in the meta-learning framework, including the gap between the training samples and the test data in the new class image data, and the gap between the base class image data and the new class image data, which have an impact on the performance stability of the new class image data and the result reduction of the base class image data category. The present invention reduces the gap caused by the difference in data volume between meta-training and meta-fine-tuning by setting a sample normalization module in the few-sample target image detection network model. The present invention reduces the gap between representations in different training stages, enabling the model to better transfer meta-knowledge and maintain the performance of detecting base class image data to a certain extent. And the present invention reduces the gap between random few-sample training data in the new class image data. When different pictures are selected as training support images for the same category, the network weights will shift in different or even completely different directions, resulting in a lack of stability. The present invention adopts a sample normalization strategy in fine-tuning, which can reduce the gap within the class and, to a certain extent, retain the relative characteristics between different samples instead of relying solely on the central features of the class.
[0029] (2) The present invention takes into account that in few-sample target detection, due to the small number of samples, the input features of the detection head are similar to each other in the feature space, making these features susceptible to the influence of supercooling points and overhot spots. Therefore, the influence of supercooling points and overhot spots may have an important impact on few-sample target detection. In this study, the query features that support image enhancement under the meta-learning framework are more similar, resulting in the influence of supercooling points and overhot spots that are unfavorable to target sample detection. Therefore, the present invention handles this problem by setting a Z-Score normalization module in the few-sample target image detection network model, and using Z-Score normalization and linear projection before the final detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a flowchart of a few-shot target image detection method based on a meta-learning framework;
[0031] Figure 2 This is a diagram of the working principle of the few-shot target image detection network model. DETAILED DESCRIPTION
[0032] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0033] like Figure 1As shown, a few-sample target image detection method based on a meta-learning framework includes steps S1-S3, which are as follows:
[0034] S1. Obtain base class image data and new class image data.
[0035] In an optional embodiment of the present invention, the present invention obtains base class image data and new class image data from Pascal VOC (a general data set), and the intersection of the base class image data and the new class image data is an empty set.
[0036] S2. Construct a few-shot target image detection network model based on a meta-learning framework, meta-train the few-shot target image detection network model using the base class image data in step S1, and meta-fine-tune the meta-trained few-shot target image detection network model using the base class image data and the new class image data in step S1.
[0037] In an optional embodiment of the present invention, the present invention constructs a few-sample target image detection network model based on a meta-learning framework, meta-trains the few-sample target image detection network model using base class image data, and meta-fine-tunes the meta-trained few-sample target image detection network model using base class image data and new class image data.
[0038] like Figure 2 As shown, the few-sample target image detection network model constructed based on the meta-learning framework of the present invention includes a support image branch and a query image branch.
[0039] One end of the support image branch is connected to the query image branch, and the support image branch is used to receive the support image, perform feature extraction, position encoding, task encoding and normalization on the support image to obtain a normalized vector of the feature vector of the support image, and transmit the normalized vector of the feature vector of the support image to the query image branch.
[0040] The image support branch includes a second feature extractor, a second Transformer encoder, a first classification head and a sample normalization module; the second feature extractor is connected to the second Transformer encoder, the second Transformer encoder is respectively connected to the first classification head and the sample normalization module, and the sample normalization module is connected to the query image branch.
[0041] The sample normalization module is used to normalize the feature vector of the support image. The calculation formula of the sample normalization module is expressed as:
[0042]
[0043] Where: s′ kis the normalized vector of the feature vector of the support image, k is the element number of the feature vector of the support image, s k is the eigenvector of the support image, μ is the mean of the eigenvector of the support image, and σ is the variance of the eigenvector of the support image.
[0044] The present invention calculates the mean of the feature vector of the support image, which is expressed as:
[0045]
[0046] Where: μ is the mean of the feature vector of the support image, and N is the number of support images.
[0047] The present invention calculates the variance of the eigenvector of the support image, which is expressed as:
[0048]
[0049] Where: σ is the variance of the eigenvector of the support image.
[0050] One end of the query image branch is connected to the support image branch, the query image branch is used to receive the query image, perform feature extraction and position encoding on the query image to obtain a feature vector of the query image, receive the normalized vector of the feature vector of the support image transmitted by the support image branch, aggregate the feature vector of the query image and the normalized vector of the feature vector of the support image, and use the inter-class correlation to obtain the query features of the support image enhancement, perform Z-Score normalization and linear projection on the query features of the base class support image enhancement to obtain the detection head input vector of the base class query image, and detect the detection head input vector of the base class query image to obtain the detection result of the base class query image.
[0051] The query image branch includes a first feature extractor, a first Transformer encoder, a first Transformer decoder, a Z-Score normalization module and a detection head connected in sequence, and the first Transformer encoder is connected to the support image branch. The first feature extractor shares weights with the second feature extractor.
[0052] The Z-Score normalization module is used to perform Z-Score normalization and linear projection on the query features that support image enhancement. The calculation formula of the Z-Score normalization module is expressed as:
[0053]
[0054] Where: Z is the feature vector after Z-Score normalization, W is the linear layer parameter, D is the enhanced query feature, μ z is the mean of the features of query feature D in all dimensions, σ zTo query the variance of feature D in all dimensions, B is the linear layer parameter, R Η It is an H-dimensional feature space, and H is the dimension of the query feature D.
[0055] The present invention calculates the mean of the query feature D in all dimensions, which is expressed as:
[0056]
[0057] Where: μ z is the mean of feature D in all feature dimensions, c is the traversal number of feature dimensions, D c is the component of the eigenvector D in the cth dimension.
[0058] The present invention calculates the variance of the query feature D in all dimensions, which is expressed as:
[0059]
[0060] Where: z It is the variance of the query feature D in all dimensions.
[0061] The present invention uses the base class image data in step S1 to perform meta-training on a few-sample target image detection network model, and the specific process is: randomly sampling the base class image data in step S1 to obtain a base class support image, and selecting a base class query image from the base class image data in step S1; the support image branch obtains a normalized vector of the feature vector of the base class support image and a classification result of the base class support image according to the base class support image; the classification result of the base class support image is used to calculate a first loss function to optimize the internal parameters of the support image branch; the query image branch obtains the detection result of the base class query image according to the base class query image and the normalized vector of the feature vector of the base class support image, and the detection result of the base class query image is used to calculate a second loss function to optimize the internal parameters of the query image branch.
[0062] The first loss function in the present invention adopts the cosine similarity cross entropy loss function.
[0063] The present invention uses the detection result of the base class query image to calculate the second loss function, which is expressed as:
[0064]
[0065] in: For test results The loss function between y and the true result y′, where y′ is the true result of the query image. is the detection result of the query image, i is the input support image number, N is the number of input support images, is the true category label c′i and the detected category labels sigmoidfocalloss,c′ i To query the true category label of the image, is the category label detected by the query image, is the actual positioning result b′ i The positioning results of the detection L 1 The linear combination of loss and GIoUloss, b′ i To query the real positioning result of the image, The detection and positioning results of the query image.
[0066] The present invention uses the base class image data and the new class image data in step S1 to meta-fine-tune the meta-trained few-sample target image detection network model. Specifically, images are selected from the base class image data and combined with images in the new class image data to form a new data set, and the new data set is used to meta-fine-tune the meta-trained few-sample target image detection network model.
[0067] The present invention utilizes a new data set to construct training data and test data in a meta-fine-tuning process. The training data is used to train a few-sample target image detection network model, and the test data is used to test the few-sample target image detection network model, so as to perform meta-fine-tuning on the few-sample target image detection network model.
[0068] S3. Obtain a few sample image data to be detected, and determine the image detection result according to the few sample image data to be detected and the few sample target image detection network model after the meta-adjustment in step S2.
[0069] In an optional embodiment of the present invention, the present invention obtains a few-sample image data to be detected, and inputs the few-sample image data to be detected as a query image into the query image branch in the few-sample target image detection network model after meta-fine-tuning. The present invention extracts 200 base class images and few-sample images in advance, obtains corresponding features through the trained second feature extractor, and obtains a feature for each category by calculating the feature mean, which is used as the input of the support image branch to determine the image detection result.
[0070] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0071] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0072] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0073] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
[0074] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A few-sample target image detection method based on a meta-learning framework, characterized in that: The following steps are involved: S1, obtaining base class image data and new class image data; S2, constructing a few-shot target image detection network model based on a meta-learning framework, meta-training the few-shot target image detection network model using the base class image data in step S1, and meta-fine-tuning the meta-trained few-shot target image detection network model using the base class image data and the new class image data in step S1; The few-shot target image detection network model built based on the meta-learning framework includes a support image branch and a query image branch; One end of the support image branch is connected to the query image branch, the support image branch is used to receive the support image, perform feature extraction, position encoding, task encoding and normalization on the support image to obtain a normalized vector of the feature vector of the support image, and transmit the normalized vector of the feature vector of the support image to the query image branch; One end of the query image branch is connected to the support image branch, the query image branch is used to receive the query image, perform feature extraction and position encoding on the query image to obtain a feature vector of the query image, receive a normalized vector of the feature vector of the support image transmitted by the support image branch, aggregate the feature vector of the query image and the normalized vector of the feature vector of the support image, and use inter-class correlation to obtain query features for support image enhancement, perform Z-Score normalization and linear projection on the query features for base class support image enhancement to obtain a detection head input vector of the base class query image, and perform detection on the detection head input vector of the base class query image to obtain a detection result of the base class query image; The query image branch includes a first feature extractor, a first Transformer encoder, a first Transformer decoder, a Z-Score normalization module, and a detection head connected in sequence, and the first Transformer encoder is connected to the support image branch; S3. Obtain a few sample image data to be detected, and determine the image detection result according to the few sample image data to be detected and the few sample target image detection network model after the meta-adjustment in step S2.
2. The method for detecting target images with a small number of samples based on a meta-learning framework according to claim 1, characterized in that: The image support branch includes a second feature extractor, a second Transformer encoder, a first classification head and a sample normalization module; the second feature extractor is connected to the second Transformer encoder, the second Transformer encoder is respectively connected to the first classification head and the sample normalization module, and the sample normalization module is connected to the query image branch.
3. The method for detecting target images with a small number of samples based on a meta-learning framework according to claim 2, characterized in that: The sample normalization module is used to normalize the feature vector of the support image. The calculation formula of the sample normalization module is expressed as: Where: s′ k is the normalized vector of the feature vector of the support image, k is the element number of the feature vector of the support image, s k is the eigenvector of the support image, μ is the mean of the eigenvector of the support image, and σ is the variance of the eigenvector of the support image.
4. The method for detecting target images with a small number of samples based on a meta-learning framework according to claim 1, characterized in that: The Z-Score normalization module is used to perform Z-Score normalization and linear projection on the query features that support image enhancement. The calculation formula of the Z-Score normalization module is expressed as: Where: Z is the feature vector after Z-Score normalization, W is the linear layer parameter, D is the enhanced query feature, μ z is the mean of the features of query feature D in all dimensions, σ z To query the variance of feature D in all dimensions, B is the linear layer parameter, R Η It is an H-dimensional feature space, and H is the dimension of the query feature D.
5. The method for detecting target images with a small number of samples based on a meta-learning framework according to claim 1, characterized in that: In step S2, the base class image data in step S1 is used to perform meta-training on the few-sample target image detection network model. The specific process is: randomly sampling the base class image data in step S1 to obtain a base class support image, and selecting a base class query image from the base class image data in step S1; The support image branch obtains the normalized vector of the feature vector of the base class support image and the classification result of the base class support image according to the base class support image; the classification result of the base class support image is used to calculate the first loss function to optimize the internal parameters of the support image branch; the query image branch obtains the detection result of the base class query image according to the base class query image and the normalized vector of the feature vector of the base class support image, and the detection result of the base class query image is used to calculate the second loss function to optimize the internal parameters of the query image branch.
6. The method for detecting target images with a small number of samples based on a meta-learning framework according to claim 5, characterized in that: The second loss function is calculated using the detection results of the base class query image, expressed as: in: For test results The loss function between y and the true result y′, where y′ is the true result of the query image. is the detection result of the query image, i is the input support image number, N is the number of input support images, is the true category label c′ i and the detected category labels The sigmoid focal loss, c′ i To query the true category label of the image, is the category label detected by the query image, is the actual positioning result b′ i The positioning results of the detection The linear combination of L1 loss and GIoU loss, b′ i To query the real positioning result of the image, The detection and positioning results of the query image.
7. The method for detecting target images with a small number of samples based on a meta-learning framework according to claim 1, characterized in that: In step S2, the base class image data and the new class image data in step S1 are used to meta-fine-tune the meta-trained few-sample target image detection network model. Specifically, images are selected from the base class image data and combined with images in the new class image data to form a new data set, and the new data set is used to meta-fine-tune the meta-trained few-sample target image detection network model.
Citation Information
Patent Citations
Small sample target detection method and system based on category semantic feature reweighting
CN113420642A