Self-supervised fine-tuning-free few-sample classification method based on neural network
By introducing lateral connections and vector quantization into the codec network, combined with data enhancement technology, the problem of fine-tuning steps and data annotation dependence in the classification of few samples is solved, and efficient and sample classification without fine-tuning is achieved, which improves the generalization ability and accuracy of the model.
Patent Information
- Application Number
- CN202510748180.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-25
AI Technical Summary
In computer vision tasks, small sample classification requires additional fine-tuning steps, and existing self-supervised learning methods are prone to overfitting when training data is scarce, resulting in insufficient generalization ability.
Using the data enhancement method in the lateral connection design of codec network, the vector quantization of the coding network feature map and the test stage, a self-supervised fine-tuning and few-sample classification model is built, and the stability and generalization ability of the model are improved through lateral connection and vector quantization, and the classification accuracy is improved using data enhancement technology.
It realizes the classification of few samples without fine-tuning, simplifies the process, alleviates the dependence on data annotation, and improves the accuracy and generalization capabilities of the model during the testing stage.
Smart Images

Figure CN120375094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of few-shot image data classification in deep learning, and particularly to a self-supervised (unsupervised), non-fine-tuning few-shot classification method with denoising and reconstruction as training objectives. Background Art
[0002] In computer vision tasks based on deep learning, supervised learning methods are often limited to the corresponding labeled data range when applying the obtained model because they use labeled data during model training. In addition, when the training data of the neural network is scarce, it often suffers from overfitting. These two data-related embarrassing situations are exactly one of the difficult problems faced by few-shot visual classification tasks. Currently, self-supervised (unsupervised) learning methods represented by contrastive learning and masked learning have achieved remarkable results. Moreover, compared with supervised learning methods, the generalization ability of self-supervised (unsupervised) learning methods has been significantly improved. Taking the visual classification task as an example, the model trained by the self-supervised (unsupervised) method has a very high classification accuracy for categories that did not appear during the training stage. Based on this, the self-supervised (unsupervised) learning method is a way to solve the problems faced by the above few-shot visual classification tasks. Summary of the Invention
[0003] Technical Problems to be Solved
[0004] The purpose of the present invention is to provide a self-supervised, non-fine-tuning few-shot classification method based on a neural network to solve the problem that few-shot classification in computer vision tasks requires an additional fine-tuning step.
[0005] Technical Solution
[0006] The present invention provides a self-supervised, non-fine-tuning few-shot classification method based on a neural network. It includes three aspects in total: First, the design of lateral connections of the encoder-decoder network. Different from the common structure of sequentially stacking the encoder and decoder, this design constructs several horizontal connections in the horizontal direction based on the covariance matrix of the feature maps of the encoder and decoder to connect the encoding modules and decoding modules at different levels in the autoencoder; Second, the design of the feature map vector quantization method of the encoding network. It is used to discretize the feature maps in the encoder network; Third, the design of the classification discrimination basis. In the test stage, multiple data augmentation methods are used to improve the model test accuracy.
[0007] In some embodiments of the present invention, the design of the lateral connections of the encoder-decoder network includes:
[0008] First, obtain the feature maps at certain specified positions of the encoder. Take the depth dimension of the feature maps as the sample dimension, and count the number of samples according to the height and width dimensions of the feature maps. Estimate the sample covariance matrix of the feature maps based on the counted number of samples; in the same way, count the samples of the feature maps at the specified stage of the decoder and centralize them (subtract the mean of all samples from each sample). The multiplication of the covariance matrix of an encoder feature map and the decoder centralized sample matrix is a lateral connection from the specified position of the encoder to the specified position of the decoder.
[0009] In some embodiments of the present invention, the design of the feature map vector quantization method includes:
[0010] Obtain a feature map from each of the specified positions of the encoder and the decoder in sequence, as well as a predefined learnable weight matrix. First, add the feature maps of the encoder and decoder and calculate an output through a multi-layer perceptron, denoted as the Q matrix; at the same time, perform a z-score normalization operation on the predefined learnable weight matrix. Subsequently, calculate two outputs through two different fully connected layer structures for the normalized learnable weight matrix, denoted as the K and V matrices in sequence. Then, perform layer normalization operations on the K and V matrices respectively. Next, multiply the Q matrix and the layer-normalized K matrix to obtain a logit matrix, denoted as L. Then, use the logit matrix L as the input of the gumbel-softmax algorithm to calculate a class sampling matrix, denoted as M. Finally, use the result of the matrix multiplication of M and V as the final quantization result of the encoder feature map.
[0011] In some embodiments of the present invention, the design of the classification discrimination basis includes:
[0012] In the model testing stage, for an input data (i.e., an image), use the BICUBIC interpolation method to expand its height and width by several times (referred to as resolution enhancement). For the image after resolution enhancement, use the 11 image enhancement strategies in Table 1 in sequence to obtain 11 enhanced images (referred to as data noise addition). In summary, an input image can obtain 11 enhanced images after resolution enhancement and data noise addition, plus the original image, a total of 12 images. Input these 12 images into the trained encoder to calculate 12 corresponding feature maps. Subsequently, perform a splicing operation on the obtained 12 feature maps in the width (or height) dimension, and calculate its covariance matrix according to the sample statistical method of feature map vector quantization. The distance between two different samples is the Euclidean distance between the corresponding covariance matrices. Use this Euclidean distance as the discrimination basis, and adopt a typical Prototypical comparison strategy for few-shot classification of images.
[0013] Beneficial effects
[0014] The self-supervised few-shot classification method based on deep neural network of the present invention has at least the following two advantages compared with the prior art: on the one hand, the present invention can complete the few-shot visual classification task in the visual scene without fine-tuning the model according to the design of the classification discrimination basis in the above test stage, streamlining the process of the few-shot classification task in the visual scene; on the other hand, the present invention relies on the denoising and self-reconstruction strategies and does not require data annotation in the model pre-training stage, alleviating the dependence on data annotation during the neural network training. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 Schematic diagram of the overall process of the present invention
[0016] Figure 2 Schematic diagram of the lateral connection structure of the present invention (including its relationship with the vector quantization module)
[0017] Figure 3 Schematic diagram of the vector quantization structure of the present invention
[0018] Figure 4 Specific configuration diagram of a group of encoding and decoding modules of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The present invention provides a self-supervised and few-shot classification method without fine-tuning based on a deep neural network. On the basis of training a neural network with an encoding and decoding structure aiming at denoising and reconstruction, by designing the lateral connection between the encoder and the decoder, the encoder feature map vector quantization method, and the classification discrimination basis for the few-shot visual classification scenario, it is realized that the model training can be completed without annotating data at all. At the same time, the pre-trained model can complete the few-shot classification task in the visual scene without fine-tuning operation.
[0020] In order to make the technology, purpose and advantages of the present invention clearer and more understandable, the following further details the present invention in combination with specific embodiments and with reference to the accompanying drawings.
[0021] The present invention provides a self-supervised and few-shot classification method without fine-tuning based on a deep neural network. Figure 1 Schematic diagram of the implementation process of the present invention. The overall process includes four parts: the encoder and decoder network, the lateral connection between the encoder and the decoder, the encoder network feature map vector quantization method (included in the lateral connection module), and the classification discrimination basis. In one training step, first randomly select N image data and scale their size to 128x128, and after predefined random data augmentation ( Figure 1 denoted as ) input into the encoder, and according to Figure 1 the arrow direction, after the data passes through different encoder modules, lateral connection modules, and decoder modules, the output is about its own image (Figure 1 is denoted as ) the distribution parameter of the pixel Figure 1 is denoted as π in i , μ i , ). The training of the model parameters adopts the method of maximum likelihood estimation, and the distribution of the image pixels adopts the discrete logistic distribution. The loss function of the model training is:
[0022]
[0023] The classification discrimination module does not intervene in the training process of the model. In the test stage, only the encoder module and the classification discrimination module on the left in Figure 1 are retained, and all the remaining modules do not participate in the test process. In Figure 4 kinds, a specific example of the present invention for an encoding and decoding module group is provided. Among them, the feature map of the encoding module is input into the vector quantization module (including the lateral connection module) after being processed by a multi-layer perceptron. The vector quantization module is placed behind the 1x1 convolutional layer of the decoder (for the sake of the beauty of the legend, the normalization layer and the non-linear activation layer structures behind the convolutional layer are omitted).
[0024] The overall design of the vector quantization module includes the lateral connection module, and their overall relationship is as Figure 2 shown. The two modules receive the same input, which come from the encoding module of the same group ( Figure 2 is denoted as in Figure 1 and has been processed by the multi-layer perceptron in Figure 2 is denoted as ) and the decoding module ( Calculate its covariance matrix ( Figure 2 is denoted as Σ in Figure 2 and the mean ( ). At the same time, the feature map from the decoder is subjected to group normalization and Gaussian noise processing with a variance of 0.1 (the processed decoder feature map is still denoted as
[0025] Then, the design of the vector quantization module is described. Its specific structure is as Figure 3As shown. After adding the feature maps from the same set of encoders and decoders, it serves as the input to the vector quantization module. This input is multiplied by the learnable weights processed by the fully connected layer and serves as the log odds of the Gumbel-softmax. Subsequently, the output of the Gumbel-sfotmax serves as a one-hot sampling matrix for the learnable weights to sample the learnable weights. The corresponding output is a vector quantization matrix (denoted as Q). After adding this matrix to the lateral connection module, a linear transformation is performed:
[0026]
[0027] The vector quantization module provides discretized features for the decoding module in the lower layer. At the same time, the lateral connection provides first-order (the mean in the above formula) and second-order (the covariance matrix in the above formula) correlations for the decoder. These two characteristics are both features that contribute to the generalization ability of the model and will make the encoding of the encoder more stable.
[0028] Table 1
[0029] invert autocontrast Rgb_to_gray_scale posterize Color_jitter equalize Guassian_blur sharpeness solarize Gaussian_noise Channel_shuffle
[0030] Finally, the design of the classification and discrimination basis in the test stage is described. In the test stage, the trained model only retains all encoding modules. Generally speaking, the output of the last encoding module is taken as the final output. For an input image, it is scaled to 256x256 size and passed through 11 image enhancement methods in Table 1 in turn to obtain 11 enhanced images corresponding to it. Together with the original image, there are 12 input images in total. After passing through the encoder, 12 corresponding feature maps are obtained. Calculate the covariance matrix of these 12 feature maps, which is the feature representation corresponding to the input image. The classification odds are calculated according to the following formula:
[0031]
[0032] Prob(Σ,Σ i ) represents the probability that the image Σ i and the image Σ belong to the same category; d(Σ,Σ j ) represents the Euclidean distance between Σ and Σ j . In the test stage, since the decoder and lateral connection modules of the model are discarded, and there is no need to fine-tune the model in the test stage, the overall computational amount of the model is greatly reduced.
[0033] So far, the embodiments of the present disclosure have been described in detail with reference to the accompanying drawings. It should be noted that, in the accompanying drawings or the main text of the specification, the implementation manners that are not illustrated or described are all forms known to those of ordinary skill in the art and have not been described in detail. In addition, the above definitions of various elements and methods are not limited to the specific structures, shapes, or manners mentioned in the embodiments, and those of ordinary skill in the art can make simple changes or replacements to them. "Comprising" does not exclude the existence of elements or steps not listed in the claims. "A" or "an" before an element does not exclude the existence of multiple such elements. In addition, unless specifically described or steps that must occur in sequence, the order of the above steps is not limited to those listed above and can be changed or rearranged according to the required design. And the above embodiments can be used in combination with each other or combined with other embodiments based on considerations of design and reliability, that is, the technical features in different embodiments can be freely combined to form more embodiments. The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. The present invention provides a self-supervised and fine-tuning-free vision few-shot classification method based on a neural network. In a network model with an encoder-decoder structure, an encoder-decoder lateral connection structure including second-order statistics is integrated between a set of paired encoder and decoder modules. On the basis of considering the additive connection of the first-order statistics of the encoder module and the decoder module, the lateral connection incorporates the high-order statistical features of the feature map of the encoding module into the decoding process of the decoding module, adding a second-order multiplicative association to the decoding process of the decoding module.
2. In the process of incorporating the high-order statistics into the decoding module as described in claim 1, the incorporation process is to perform a multiplication operation on the sum of the feature maps of the encoder-decoder module described in claim 1 and a predefined learnable weight, and the obtained result is used as the log-odds of the Gumbel-softmax algorithm, so as to perform vector quantization sampling on the learnable weight. Finally, the second-order multiplicative and first-order additive associations described in claim 1 are applied to the output result of the vector quantization.
3. After training the neural network model with the encoder-decoder structure using the two strategies described in claims 1 and 2, only part or all of its encoders are retained as the input encoding structure in the test phase. At the same time, resolution doubling and image enhancement are used to obtain a multiplicatively increased number of encoded feature maps. Finally, the second-order statistics of the multiplied feature maps are used as the classification discriminant basis for the corresponding input.