An image classification method and device based on a self-supervised model

CN116310584BActive Publication Date: 2026-09-29CHONGQING TESLINK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310332609.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2026-09-29
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

[0005]本申请实施例提供一种基于自监督模型的图像分类方法及装置,用以解决当前针对图像分类的自监督模型存在的特征提取能力较弱、模型性能较差的问题

Benefits of technology

[0026]本申请实施例提供一种基于自监督模型的图像分类方法及装置,通过对模型网络各阶段的输入图像采用不同大小的补丁进行掩码处理,并分别计算模型网络各阶段的损失,设计了一种多阶段的目标,能够预测输入图像全分辨率下的所有像素值,大大提高了图像局部上下文信息的捕捉能力,以进一步提高模型的预测性能。并且,本申请引入了特征提取能力强的自监督模型编码器,以及轻量级设计的解码器,能够很好的提升模型的性能和精准度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310584B_ABST
    Figure CN116310584B_ABST
Patent Text Reader

Abstract

The application discloses an image classification method and device based on a self-supervised model, to solve the problem of weak feature extraction capability and poor model performance of the current self-supervised model for image classification. The method trains an initial self-supervised model according to an unlabeled training set; combines labeled data to jointly train the initial self-supervised model to obtain an image classification model; in the process of joint training, the input image is masked according to the image block size of each stage of the model network, and the loss of each stage output prediction is calculated; and the image classification model is used for image classification. The method designs a multi-stage target, which can predict all pixel values of the input image at full resolution, greatly improves the capture ability of the local context information of the image, and further improves the prediction performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of self-supervised learning technology, and in particular to an image classification method and apparatus based on a self-supervised model. Background Technology

[0002] Machine learning is a crucial component in the development of artificial intelligence. Machine learning is divided into supervised learning, unsupervised learning, and reinforcement learning. Among unsupervised learning methods, self-supervised learning is one such approach. It can learn a general feature representation without supervision, which can then be used for downstream tasks such as image classification, object detection, and semantic segmentation.

[0003] In Natural Language Processing (NPL), generative self-supervised models are widely used as the target for pre-training language models. They typically extract semantic information from large unlabeled corpora in a self-supervised manner. In the setting of a generative self-supervised model, a certain percentage of the labels in the sentences input to the model are masked, and the goal is to predict the original information corresponding to the masked labels based solely on their context.

[0004] However, in the field of vision, images have higher dimensionality, noise, and redundant formats compared to text. Current self-supervised models for image classification still suffer from weak feature extraction capabilities and poor model performance. Summary of the Invention

[0005] This application provides an image classification method and apparatus based on a self-supervised model to address the problems of weak feature extraction capability and poor model performance in current self-supervised models for image classification.

[0006] This application provides an image classification method based on a self-supervised model, comprising:

[0007] Train an initial self-supervised model based on the unlabeled training set;

[0008] By combining labeled data, the initial self-supervised model is jointly trained to obtain an image classification model.

[0009] During the joint training process, the input image is masked according to the image patch size of each stage of the model network, and the loss of the output prediction at each stage is calculated.

[0010] The image classification model described above is used to classify images.

[0011] In one example, the input image is masked according to the image patch size of each stage of the model network. Specifically, this includes: determining the image patch resolution of each stage of the model network; and masking the input image with patches of the same size as the image patch resolution of the corresponding stage.

[0012] In one example, calculating the loss for the output prediction at each stage specifically includes: calculating the loss for each stage based on the encoder's output prediction for each stage and the input image after masking processing at the corresponding stage; and performing gradient backpropagation training based on the loss.

[0013] In one example, the encoder is at least one of the following: a swin encoder, a PVT encoder.

[0014] In one example, the decoder used in the joint training process is a lightweight image decoder, TJpgDec.

[0015] In one example, the input image is masked, which specifically includes: applying a random mask to the input image with a set probability based on the determined patch size; and then randomly shifting the masked image.

[0016] In one example, before training the initial self-supervised model, the method further includes: acquiring source images and establishing an initial training set; performing data augmentation on the initial training set to obtain the unlabeled training set; the data augmentation includes at least one of image normalization, multi-scale cropping, and rotation.

[0017] In one example, the formula used for normalization is:

[0018]

[0019] Where x represents the input data, x * The output is normalized so that all data are between [0,1]. Normalization improves convergence speed and model accuracy. max(x) represents the maximum value, min(x) represents the minimum value, and the difference between the maximum and minimum values ​​is used to normalize the source image. The multi-scale cropping process includes scaling the source image at different scales based on a set scaling factor. The rotation process includes rotating the source image based on a set rotation angle.

[0020] In one example, the model network is a Swing converter network; according to the image patch size of each stage of the model network, the input image is masked, specifically including: according to the image patch size of 4×4, 8×8, 16×16, and 32×32 of different stages of the Swing converter network, the input image is masked with patches of resolution 4×4, 8×8, 16×16, and 32×32 respectively.

[0021] This application provides an image classification device based on a self-supervised model, comprising:

[0022] The first training module trains an initial self-supervised model based on the unlabeled training set;

[0023] The second training module combines labeled data to jointly train the initial self-supervised model to obtain an image classification model.

[0024] The processing module performs masking processing on the input image according to the image patch size of each stage of the model network during the joint training process, and calculates the loss of the output prediction at each stage.

[0025] The classification module uses the image classification model to classify images.

[0026] This application provides an image classification method and apparatus based on a self-supervised model. By applying patches of different sizes to the input image at each stage of the model network for masking and calculating the loss at each stage, a multi-stage target is designed. This approach can predict all pixel values ​​at full resolution of the input image, significantly improving the ability to capture local contextual information and further enhancing the model's predictive performance. Furthermore, this application introduces a self-supervised model encoder with strong feature extraction capabilities and a lightweight decoder, which can effectively improve the model's performance and accuracy.

[0027] Compared to supervised image classification algorithms, self-supervised image classification algorithms save significant manual annotation costs and training set preparation time. They can learn image features from large-scale unlabeled data without any manual annotation, achieving or even surpassing the accuracy of supervised learning methods. Furthermore, data augmentation expands the dataset size and increases data diversity, which improves the robustness and versatility of the model during subsequent training. Additionally, joint training with given labeled data allows the model to achieve or even surpass the accuracy of supervised learning methods with minimal manual annotation, effectively improving the network's classification of target objects and promoting better task and architecture alignment in transfer learning. Attached Figure Description

[0028] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0029] Figure 1 A flowchart of an image classification method based on a self-supervised model provided in this application embodiment;

[0030] Figure 2 This is a schematic diagram of the structure of an image classification device based on a self-supervised model provided in an embodiment of this application. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0032] Figure 1 The flowchart of the image classification method based on a self-supervised model provided in this application embodiment specifically includes the following steps:

[0033] S101: Train the initial self-supervised model based on the unlabeled training set.

[0034] In this embodiment of the application, an initial self-supervised model is trained using an unlabeled training set based on a self-supervised learning method, and the performance of subsequent models can be verified using a validation set to determine whether the model training is complete.

[0035] The amount of data in the unlabeled training set and the validation set can be determined according to a preset ratio, and the amount of data in the unlabeled training set is usually larger. For example, the training set contains 1.28M images and the validation set contains 5000 images, both of which contain the same 1000 categories.

[0036] Compared with supervised image classification algorithms, self-supervised image classification algorithms can save a lot of manual annotation costs when building training sets and reduce the preparation time of training sets. They can learn image features from large-scale unlabeled data without using any manually labeled data, and can achieve or even surpass the accuracy achieved by supervised learning methods.

[0037] In one embodiment, before training the initial self-supervised model, source images need to be collected to establish an initial training set. The initial training set may have issues with data homogeneity, such as a limited number of body parts or behaviors. In such cases, the model is prone to overfitting during training. Therefore, data augmentation is needed to obtain an unlabeled training set for model training. Data augmentation can expand the size of the dataset and increase its diversity, which is beneficial for improving the robustness and diversity of the model during subsequent training.

[0038] Specifically, data augmentation includes at least one of the following: image normalization, multi-scale cropping, and rotation.

[0039] When normalizing, the following formula (1) is used:

[0040]

[0041] Where x represents the input data, x * This represents the normalized output data, ensuring all data points fall within the range [0,1]. `max(x)` represents the maximum value, and `min(x)` represents the minimum value. Normalization can improve the convergence speed and accuracy of the model.

[0042] Multi-scale cropping is a process that scales the source image at different scales based on a set scaling factor. Different source images in the initial training set may differ significantly in terms of angle, brightness, contrast, and target size, which differs from common action recognition datasets used for model training. Therefore, a multi-scale training strategy can be employed to scale the initial training set to different scales for input, thereby improving the network's adaptability to recognizing actions from targets of varying sizes.

[0043] For example, wrapAffine can be used to scale the source image to six scales for input, as shown in the following formula (2):

[0044]

[0045] Among them, f x and f y represents the focal length (i.e. scaling factor) of the x-axis and y-axis respectively, x and y represent the width and height of the input data before scaling, and x′ and y′ represent the width and height after scaling.

[0046] Rotation processing rotates the source image based on a set rotation angle. For example, rotation can be performed using wrapAffine, as shown in formula (3) below:

[0047]

[0048] Where θ represents the rotation angle, x and y represent the width and height of the input before scaling, and x′ and y′ represent the width and height after scaling.

[0049] S102: Combine labeled data to jointly train the initial self-supervised model to obtain an image classification model.

[0050] In the embodiments of this application, after the initial self-supervised model training is completed, it can be jointly trained based on the given labeled data for subsequent image classification tasks. This enables the image classification model to achieve or even surpass the accuracy achieved by supervised learning methods with only a small amount of manually labeled data, thereby effectively improving the network's classification of target objects and promoting better task alignment and architecture alignment in transfer learning.

[0051] S1021: During joint training, the input image is masked according to the image patch size of each stage of the model network, and the loss of the output prediction at each stage is calculated.

[0052] During joint training, the input image needs to be masked, and the original information corresponding to the masked region needs to be predicted. Currently, the image patch size of the mask can be an integer multiple of the image patch size of the model. This processing method uses the same size mask image patch at each stage of the network.

[0053] Specifically, when the model network is a Swing converter network, for example, if the image patch size in the first stage is 4×4 and the image patch size in the last stage is 32×32, then the mask image patch designed at this time only needs to be an integer multiple of the image patch size in the first stage.

[0054] In one embodiment, when designing the image patch size for the mask, the different image patch resolutions at each stage of the model network can be determined, and patches of the same size as the image patch resolution at the corresponding stage can be used to mask the input image. This processing method uses patches of different sizes for masking at each stage of the network.

[0055] Specifically, when the model network is a Swing converter network, the input image is masked by patches of different sizes with resolutions of 4×4, 8×8, 16×16, and 32×32, depending on the image patch size of 4×4, 8×8, 16×16, and 32×32 at different stages of the Swing converter network.

[0056] In one embodiment, based on the encoder's output predictions for each stage of the network and the input image after masking at the corresponding stage, the loss for each stage can be calculated separately, and gradient backpropagation training can be performed based on the loss. Compared to most current generative self-supervised models that only calculate the mask loss for the last layer of features output by the encoder, capturing only the global semantic information of the image without capturing the local contextual information, this training method designs a multi-stage objective by calculating the loss for each stage separately. This enables the prediction of all pixel values ​​at full resolution of the input image, greatly improving the ability to capture local contextual information and further enhancing the model's predictive performance.

[0057] Furthermore, most current generative self-supervised methods use relatively simple encoders, employing models such as ResNet or VIT, which struggle to accurately represent features, thus impacting the model's final performance. This application introduces a self-supervised model encoder with strong feature extraction capabilities, employing at least the Swin encoder and the Pyramid Vision Transformer (PVT) encoder, to improve the encoder's extraction performance.

[0058] Furthermore, most generative self-supervised models currently employ complex prediction targets and complex decoder designs, resulting in poor performance. In contrast, this application designs a simple image masking process to predict the target and uses a lightweight decoder, such as [example code], to achieve better performance.

[0059] In one embodiment, when masking the input image, the input image needs to be randomly masked with a set probability according to the determined patch size, and then the masked image is randomly shifted to enhance the mask generalization.

[0060] S103: Use an image classification model to classify images.

[0061] In this embodiment, a multi-stage objective is designed by masking the input images at each stage of the model network with patches of different sizes and calculating the loss at each stage. This objective can predict all pixel values ​​at full resolution of the input image, greatly improving the ability to capture local contextual information and further enhancing the model's predictive performance. Furthermore, this application introduces a self-supervised model encoder with strong feature extraction capabilities and a lightweight decoder, such as the lightweight image decoder TJpgDec, which significantly improves the model's performance and accuracy.

[0062] The above describes the image classification method based on a self-supervised model provided in this application. Based on the same inventive concept, this application also provides a corresponding image classification device based on a self-supervised model, such as... Figure 2As shown.

[0063] Figure 2 The schematic diagram of the image classification device based on the self-supervised model provided in the embodiments of this application specifically includes:

[0064] The first training module 201 trains an initial self-supervised model based on the unlabeled training set;

[0065] The second training module 202, in conjunction with labeled data, performs joint training on the initial self-supervised model to obtain an image classification model;

[0066] The processing module 2021 performs masking processing on the input image according to the image patch size of each stage of the model network during the joint training process, and calculates the loss of the output prediction at each stage.

[0067] The classification module 203 uses the image classification model to perform image classification.

[0068] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0069] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0070] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An image classification method based on a self-supervised model, characterized in that, include: Train an initial self-supervised model based on the unlabeled training set; By combining labeled data, the initial self-supervised model is jointly trained to obtain an image classification model. During the joint training process, the input image is masked according to the image patch size at each stage of the model network, and the loss of the output prediction at each stage is calculated. Gradient backpropagation training is then performed based on the loss. The model network is a Swing converter network; according to the image block size of each stage of the model network, the input image is masked, specifically including: according to the image block size of 4×4, 8×8, 16×16, 32×32 of different stages of the Swing converter network, the input image is masked with patches of resolution 4×4, 8×8, 16×16, 32×32 respectively. The image classification model described above is used to classify images.

2. The method according to claim 1, characterized in that, Calculate the loss for the output prediction at each stage, specifically including: The loss for each stage is calculated based on the encoder's output predictions for each stage and the input image that has been masked for the corresponding stage.

3. The method according to claim 2, characterized in that, The encoder used is a Swin encoder.

4. The method according to claim 1, characterized in that, The decoder used in the joint training process is a lightweight image decoder, TJpgDec.

5. The method according to claim 1, characterized in that, Masking the input image specifically includes: Based on the determined patch size, apply a random mask to the input image with a set probability. The image after random masking is then randomly shifted and masked again.

6. The method according to claim 1, characterized in that, Before training the initial self-supervised model, the method further includes: Collect source images and establish an initial training set; The initial training set is augmented to obtain the unlabeled training set; the data augmentation includes at least one of image normalization, multi-scale cropping, and rotation.

7. The method according to claim 6, characterized in that, The formula used for normalization is: in Indicates input data, This represents the normalized output, ensuring all data points are between [0,1]. This indicates taking the maximum value. This indicates taking the minimum value; The multi-scale cropping process includes: Based on the set scaling factor, the source image is scaled at different scales. The rotation process includes: The source image is rotated based on a set rotation angle.

8. An image classification device based on a self-supervised model, characterized in that, include: The first training module trains an initial self-supervised model based on the unlabeled training set; The second training module combines labeled data to jointly train the initial self-supervised model to obtain an image classification model. The processing module, during the joint training process, performs masking processing on the input image according to the image patch size of each stage of the model network, calculates the loss of the output prediction at each stage, and performs gradient backpropagation training based on the loss; wherein, The model network is a Swing converter network; according to the image block size of each stage of the model network, the input image is masked, specifically including: according to the image block size of 4×4, 8×8, 16×16, 32×32 of different stages of the Swing converter network, the input image is masked with patches of resolution 4×4, 8×8, 16×16, 32×32 respectively. The classification module uses the image classification model to classify images.

Citation Information

Patent Citations

  • Brain tumor self-supervision pre-training method and device based on attention symmetry self-coding

    CN115035093A