Method, system, medium and equipment for generating positive image sample to realize image classification

By guiding the denoising diffusion model to generate semantically consistent positive samples, combined with comparison learning and linear classifier training methods, the problems of image samples diversity and semantic consistency in the field of autonomous driving are solved, and the robustness and generalization ability of the model are improved.

CN120032163AActive Publication Date: 2025-05-23ZHEJIANG UNIV

Patent Information

Application Number
CN202411989922.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In the field of autonomous driving, it is difficult for the prior art to increase the diversity of image samples while maintaining semantic consistency, resulting in insufficient generalization of learned representations.

Method used

By guiding the denoising diffusion model to generate positive samples with consistent semantics, the image encoder is trained unsupervised, and cascades it with the linear classifier to form an image classification model, train in a supervised manner, fix the parameters of the image encoder, and optimize the linear classifier parameters to achieve image category judgment.

Benefits of technology

It significantly reduces the domain shift between samples, improves the robustness and generalization ability of learning representations, and increases the diversity of samples while maintaining semantic consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032163A_ABST
    Figure CN120032163A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a system, a medium and equipment for generating an image positive sample to realize image classification, and belongs to the field of computer vision. According to the method, firstly, text semantics are extracted and rewritten through a multi-modal model, visual prompt is enhanced, then a denoising diffusion model is guided to generate a positive sample consistent with original image semantics, and an image encoder is subjected to unsupervised training through comparative learning, so that semantic information is accurately obtained; and training a linear classifier by using the labeled feature vectors, thereby constructing an image classification model for an image classification task, and finally completing image category judgment under the condition that a large amount of labeled data is not needed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a method, system, medium and device for generating positive image samples to realize image classification. Background Art

[0002] Generally speaking, image classification algorithms in certain professional fields such as autonomous driving require a large number of fine annotations of images in the training data set. However, in actual scenarios, large-scale professional annotation of data requires professionals to devote a lot of time and energy to data annotation, which not only increases the difficulty and cost of obtaining training data, but also greatly limits the possibility of applying image classification algorithms to actual scenarios of autonomous driving. Considering this problem, self-supervised contrastive learning methods have shown exciting prospects in the research of image classification.

[0003] In the field of autonomous driving, contrastive learning, as an efficient self-supervised learning method, has been widely used in the training process of perception systems by relying on its ability to improve the robustness and generalization ability of the model. Typical contrastive learning methods such as SimCLR, MoCo, and VICReg generate positive samples from original images through data augmentation technology to maximize the similarity between positive samples and minimize the difference between negative samples. This strategy relies heavily on the generation mechanism of positive and negative samples, which has a direct and significant impact on the quality of learned representations.

[0004] Although contrastive learning has shown significant results, its limitations in sample generation are equally obvious. First, the currently widely used data augmentation methods, such as cropping, rotation, and flipping, although they can increase the number of samples, often fail to effectively expand the diversity of samples. These traditional methods have difficulty in ensuring the semantic consistency between samples in the process of generating samples, and are prone to domain shifts between samples, resulting in insufficient generalization of the learned representations. Secondly, although the use of generative models such as generative adversarial networks (GANs) can create new visually diverse samples, these methods still have deficiencies in controlling the semantic consistency of generated samples, which may cause the generated positive sample pairs to be semantically biased from the original samples.

[0005] In addition, although existing research has begun to focus on using multimodal large models (MLLM) to improve the quality of positive samples, this method shows new possibilities in controlling the consistency and diversity between positive samples. However, research in this field is still insufficient. How to effectively integrate data of different modalities in practical applications of multimodal models and how to increase sample diversity while maintaining semantic consistency are technical challenges that need to be solved urgently. Solving these problems requires in-depth technical innovation and comprehensive model optimization in order to achieve a wider application of self-supervised learning methods in autonomous driving systems. Summary of the Invention

[0006] The object of the present invention is to solve the problem in the prior art of how to increase the sample diversity while maintaining semantic consistency in the field of autonomous driving, and to provide a method, system, medium and device for generating positive image samples to achieve image classification. The method of the present invention can be applied to scenarios in the field of autonomous driving where the amount of training data is insufficient or the labels are scarce. Only by generating positive samples with consistent semantics through the framework can the data diversity be enhanced and the representation ability of the model be improved.

[0007] In order to achieve the above object of the invention, the present invention specifically adopts the following technical solutions:

[0008] In the first aspect, the present invention provides a method for generating positive image samples to achieve image classification, which includes the following steps:

[0009] S1. Based on the obtained original image, guide the denoising diffusion model to generate positive samples, and perform unsupervised training on the image encoder in a contrastive learning manner based on the generated positive samples;

[0010] S2. Cascade the image encoder obtained by unsupervised training with a linear classifier to be trained to form an image classification model, and train the image classification model in a supervised manner. Fix the parameters of the image encoder in the image classification model and optimize the parameters of the linear classifier so that it can achieve image category judgment;

[0011] S3. Input the image to be classified into the trained image classification model, and output the classification result of the image to be classified.

[0012] On the basis of the above solution, each step can be implemented in the following preferred specific ways.

[0013] As a preference of the above first aspect, in step S1, the specific process of generating positive samples is as follows:

[0014] S11. Input the original image and a sentence for prompting semantic extraction into a pre-trained multimodal large model to generate a first feature description of the original image in the text space;

[0015] S12. Send the first feature description and a sentence for prompting semantic rewriting into a large language model to perform semantic rewriting on the first feature description and generate a second feature description of the original image in the text space;

[0016] S13. Adding Gaussian noise to the original image to generate a random noise image, obtaining a mask image of the original image and a target image, weightedly mixing the mask image of the original image and the target image according to pixel values ​​to generate a mask mixed image, using one of the original image, the random noise image, the mask image of the original image, and the mask mixed image as a first visual cue, and performing data enhancement on the first visual cue to obtain a second visual cue;

[0017] S14. Input the second feature description and the second visual cue into the denoising diffusion model to generate a semantically-aware synthetic image, and use the semantically-aware synthetic image as a positive sample that is semantically consistent with the original image.

[0018] As a preferred embodiment of the first aspect, in step S1, the specific process of unsupervised training the image encoder is:

[0019] In each iteration round, the original image is data enhanced to generate a first processed image, the semantically-aware synthetic image is data enhanced to generate a second processed image, the first processed image is input into the image encoder in the MoCo contrastive learning framework for feature extraction to obtain a first feature map, the second processed image is input into the momentum image encoder in the MoCo contrastive learning framework for feature extraction to obtain a second feature map, the contrast loss is calculated based on the first feature map and the second feature map, and the training is continuously iterated. After reaching the preset iteration rounds, the training of the image encoder is completed.

[0020] As a preferred embodiment of the above-mentioned first aspect, in step S2, in the process of training the linear classifier, the original image is input into the trained image encoder, the features output by the image encoder are passed through the linear classifier, the output dimension of the linear classifier is set to the number of categories, the linear classifier outputs the image category prediction result, the true label corresponding to the original image is obtained, the first cross entropy loss between the true label and the image category prediction result is calculated, and the parameters of the linear classifier are updated based on minimizing the first cross entropy loss. The training is continuously iterated, and after reaching a preset number of iterations, the linear classifier converges to obtain a trained linear classifier.

[0021] As a preferred embodiment of the first aspect, the linear classifier adopts a fully connected layer.

[0022] As a preferred embodiment of the first aspect above, the output dimension of the linear classifier is set to 100.

[0023] In a second aspect, the present invention provides a system for generating positive image samples to implement image classification, comprising:

[0024] A first model training module is used to generate positive samples by guiding the denoising diffusion model based on the acquired original image, and to perform unsupervised training on the image encoder in a contrastive learning manner based on the generated positive samples;

[0025] The second model training module is used to cascade the image encoder obtained by unsupervised training and a linear classifier to be trained to form an image classification model, train the image classification model in a supervised manner, fix the parameters of the image encoder in the image classification model, and optimize the parameters of the linear classifier so that it can realize image category judgment;

[0026] The result acquisition module is used to input the image to be classified into the trained image classification model and output the classification result of the image to be classified.

[0027] In a third aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the method of generating image positive samples to achieve image classification as described in any of the solutions in the first aspect above.

[0028] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method of generating positive image samples to realize image classification as described in any of the schemes of the first aspect above is implemented.

[0029] In a fifth aspect, the present invention provides a computer electronic device comprising a memory and a processor;

[0030] The memory is used to store computer programs;

[0031] The processor is used to implement the method of generating image positive samples to achieve image classification as described in any solution of the first aspect when executing the computer program.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] The present invention improves the way of generating positive samples in contrastive learning by introducing a semantically-aware positive sample generation mechanism, and effectively solves the deficiencies of traditional data enhancement methods in terms of sample diversity and semantic consistency. Compared with existing data enhancement techniques such as cropping, rotation, and flipping, the present invention significantly reduces the domain shift between samples and improves the robustness and generalization ability of learning representations by generating positive sample pairs with higher semantic consistency. In addition, the present invention utilizes a multimodal large model (MLLM) to increase the diversity of samples while maintaining semantic consistency between positive samples, overcomes the limitations of existing generation methods in semantic control, optimizes the smoothness of the sample space, and facilitates the seamless integration of generated samples with various mainstream contrastive learning frameworks and traditional manually designed data enhancement methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 Schematic diagram of the steps of the method of the present invention;

[0035] Figure 2 A schematic diagram of a sentence used to prompt semantic rewriting in the method of the present invention;

[0036] Figure 3 A schematic diagram of training an image encoder and a linear classifier using the method of the present invention;

[0037] Figure 4 is a system block diagram of the present invention;

[0038] Figure 5 A schematic diagram of the components of computer electronic equipment. DETAILED DESCRIPTION

[0039] In order to make the above-mentioned purpose, features and advantages of the present invention more obvious and easy to understand, the specific implementation mode of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. The technical features in each embodiment of the present invention can be combined accordingly without conflicting with each other.

[0040] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for the purpose of distinguishing descriptions, and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features.

[0041] like Figure 1 As shown, in a preferred implementation of the present invention, the method for generating positive image samples to achieve image classification includes the following steps S1 to S3. The specific implementation process is described in detail below.

[0042] S1. Based on the acquired original image, the denoising diffusion model is guided to generate positive samples, and the image encoder is unsupervisedly trained in a contrastive learning manner based on the generated positive samples.

[0043] It should be noted that, in step S1 of the present invention, an image data set consisting of original images is first obtained. The image data set used in this embodiment is Car196, and the main application of this data set in the field of autonomous driving is to support vehicle identification and classification tasks.

[0044] It should be noted that in step S1 of the present invention, after obtaining the image data set, the denoising diffusion model is guided to generate positive samples. Each time the positive sample is generated, two steps of text semantic extraction and visual cue enhancement are included. Taking the generation of a positive sample as an example, the specific process is as follows:

[0045] 1) Text semantic extraction

[0046] S11. In each iteration round of generating positive samples, the original image and a sentence for prompting semantic extraction are input into a pre-trained multimodal large model (MLLM) to generate a first feature description of the original image in the text space.

[0047] It should be noted that, in step S11 of the embodiment of the present invention, Figure 3 As shown, a sentence for prompting semantic extraction (i.e. Figure 3 An example of a semantic extraction prompt in the text is: "Please describe this image in detail." The above multimodal large model specifically adopts the BLIP-2 model, which ultimately generates the first feature description of the original image in the text space.

[0048] The implementation of the above-mentioned BLIP-2 model belongs to the prior art. In this embodiment, the model is briefly described so that those skilled in the art can understand. The BLIP-2 model is a multimodal model that can combine image and text data to understand visual scenes and generate natural language descriptions. The input of the model includes images and text prompts (optional), and the output is a text description or an extracted feature vector. In this embodiment, images and text prompts are input into the BLIP-2 model, and the output of the model is a text description.

[0049] 2) Text semantic rewriting

[0050] S12. Send the first feature description and a sentence for prompting semantic rewriting into a large language model (LLM), perform semantic rewriting on the first feature description, and generate a second feature description of the original image in the text space.

[0051] It should be noted that, in step S12 of the embodiment of the present invention, several examples for guiding semantic rewriting are added to the above sentence to enhance sample diversity. Figure 3 An example of a semantic rewrite hint in Figure 2 In this embodiment, the large language model used is LLaMA, a product publicly released by Metaverse Platform Company (Meta), which rewrites the first feature description and outputs the second feature description.

[0052] In this embodiment, 100 samples are prepared, and 5 samples are randomly selected each time to construct sentences for prompting semantic rewriting. This is to ensure that the diversity of samples enhances the consistency of semantics before and after, thereby enhancing the richness of the samples.

[0053] 3) Enhanced visual cues

[0054] S13. Add Gaussian noise to the original image to generate a random noise image, obtain a mask image of the original image and a target image, weightedly mix the mask image of the original image and the target image according to pixel values ​​to generate a mask mixed image, use the original image, the random noise image, the mask image of the original image, and one of the mask mixed images as a first visual cue, perform data enhancement on the first visual cue, and obtain a second visual cue.

[0055] It should be noted that in step S13 of the embodiment of the present invention, in order to more flexibly control the generation type of positive samples, four different visual cues are designed, namely, the original image, the random noise image, the mask image of the original image, and the mask mixing image (Mask Mixing Image). Then, one of the images is selected as the first visual cue, and Gaussian noise is additionally added to the first visual cue and randomly cropped to generate the second visual cue. Among them, the mask image of the original image is generated by YOSO (You Only Segment Once). YOSO is a model for image segmentation, mainly used to generate mask images. A mask image (Mask Image) is an image used to mark or highlight a specific area in computer vision. Generally speaking, a mask image is composed of a binary image, which is used to distinguish between a target area and a background area.

[0056] 4) Positive sample generation

[0057] S14. Input the second feature description and the second visual cue into the denoising diffusion model to generate a semantically-aware synthetic image, and use the semantically-aware synthetic image as a positive sample that is semantically consistent with the original image.

[0058] It should be noted that in step S14 of the present invention, the basic principle of the denoising diffusion probabilistic model (DDPM) is mainly used to generate image samples. The denoising diffusion model mentioned here specifically refers to a type of denoising diffusion model. The common point of this type of model is to achieve tasks such as image generation through a step-by-step denoising and denoising process. The input of the DDPM model is a random noise image and conditional information (optional), and the output is a high-quality image after multi-step denoising. In this embodiment, the input of the DDPM model is an image and text semantic description, and the output is a high-quality image after multi-step denoising. For the structure of the denoising diffusion model, it can include different deformation structures, such as StableDiffusion and ControlNet deformation structures, so the present invention does not make too many restrictions on its specific structure. The denoising diffusion model used in this embodiment is ControlNet-sd15-seg, and its implementation method belongs to the prior art and will not be repeated.

[0059] It should be noted that in step S1 of the present invention, after the positive samples are generated, the image encoder is trained unsupervisedly in a contrastive learning manner. The training process of the image encoder is as follows:

[0060] In each iteration round, the original image is data enhanced to generate a first processed image, the semantically-aware synthetic image is data enhanced to generate a second processed image, the first processed image is input into the image encoder in the MoCo contrastive learning framework for feature extraction to obtain a first feature map, the second processed image is input into the momentum image encoder in the MoCo contrastive learning framework for feature extraction to obtain a second feature map, the contrast loss is calculated based on the first feature map and the second feature map, and the training is continuously iterated. After reaching the preset iteration rounds, the training of the image encoder is completed.

[0061] In the process of training the image encoder of the present invention, the existing contrastive learning framework MoCo is used. Other contrastive learning methods can also be selected here, and corresponding contrastive losses can be designed, so there is no limitation in the present invention. In this embodiment, the MoCo-v2 contrastive learning framework is used.

[0062] S2. The image encoder obtained by unsupervised training is cascaded with a linear classifier to be trained to form an image classification model, and the image classification model is trained in a supervised manner. The parameters of the image encoder in the image classification model are fixed, and the parameters of the linear classifier are optimized so that it can realize image category judgment.

[0063] It should be noted that in step S2 of the present invention, the image encoder obtained by contrastive learning training is used as a feature extractor, which is cascaded with a linear classifier to form an image classification model. At the same time, the parameters of the image encoder are fixed, and a linear classifier is trained through supervised fine-tuning.

[0064] In the process of training the linear classifier, the original image is input into the trained image encoder, the features output by the image encoder are passed through the linear classifier, the output dimension of the linear classifier is set to the number of categories, the linear classifier outputs the image category prediction result, the true label corresponding to the original image is obtained, the first cross-entropy loss (Cross-Entropy Loss) between the true label and the image category prediction result is calculated, and the parameters of the linear classifier are updated based on minimizing the first cross-entropy loss. The training is continuously iterated. After reaching the preset iteration rounds, the linear classifier converges and a trained linear classifier is obtained.

[0065] It should be noted that the present invention uses a part of the original image when training the image encoder, and does not need the real label corresponding to this part of the original image. When the image encoder is trained and the linear classifier needs to be trained, another part of the original image is used and the real label of this part of the original image is used at the same time. Training a linear classifier through supervised fine-tuning aims to improve the relevance and interpretability of the prediction.

[0066] It should be noted that the linear classifier of the present invention can adopt a variety of network structures, can use existing models, or can be designed by those skilled in the art according to actual needs. In this embodiment, a fully connected layer is used as a linear classifier, and its output dimension is set to 100.

[0067] S3. Input the image to be classified into the trained image classification model and output the classification result of the image to be classified.

[0068] It should be noted that in step S3 of the present invention, after the linear classifier is trained, the image classification model can be used to perform image classification, and the specific process will not be repeated here.

[0069] The present invention will now use a specific example to demonstrate the application effect of the method for generating positive image samples to achieve image classification described in S1 to S3 in the above embodiments on a specific data set, so as to facilitate understanding of the essence of the present invention.

[0070] Example

[0071] The specific implementation process of the method for generating positive image samples to realize image classification adopted in this embodiment is as described above and will not be repeated here.

[0072] In order to demonstrate the technical effect of the present embodiment, the present invention performs an ablation experiment based on MoCo-v2 on the ImageNet-100 dataset. The ImageNet-100 dataset is a subset of ImageNet, which is a large image dataset consisting of 100 categories. The ImageNet-100 dataset is divided into a "train" training set and a "val" validation set to evaluate the classification performance of the model on different categories, and the linear classification accuracy is used as an evaluation indicator. In the method of the present invention, the contrastive learning training of the image encoder is modified to be trained only with the original image, and the rest of the process remains unchanged. The image classification model finally obtained by this method is used as the baseline model, that is, the baseline model does not use semantically perceived synthetic images, nor does it contain semantic extraction cues and visual cues.

[0073] This embodiment designs three different semantic extraction prompt sentences: (a) "Question: What does this image describe? Answer:" (b) "Please describe this image in detail." (c) "What objects are included in this image?", and these three semantic extraction prompt sentences are respectively applied in the method of the present invention, and compared with the baseline model without prompts. The results are shown in Table 1.

[0074] Table 1. Ablation experiment of semantic extraction hint component of the present invention

[0075]

[0076] It can be observed from Table 1 that: 1) compared with the baseline model, the accuracy of the present invention is significantly improved no matter which prompt is selected; 2) providing a more detailed image description is beneficial to the accuracy.

[0077] In addition to the above-mentioned semantic extraction cues, this embodiment designs four visual cues (i.e., original image, random noise image, mask image of original image, or mask mixed image), and applies these three visual cues in the method of the present invention respectively, and compares them with the baseline model without cues. The results are shown in Table 2.

[0078] Table 2. Ablation experiment of the visual cue component of the present invention

[0079]

[0080]

[0081] As shown in Table 2, when the masked mixed image is used as a visual cue, the highest accuracy can be achieved. On the contrary, when the random noise image is used as a visual cue, the basic information of the original image is lost and unreasonable positive samples are introduced, resulting in a decrease in accuracy. However, when any of the above visual cues is used, the accuracy of the present invention is higher than that of the baseline model without using visual cues.

[0082] This embodiment also provides an ablation experiment on the impact of the model when multiple semantic-aware synthetic images are generated, and the samples are enhanced by semantic extraction cues and visual cues respectively. The results are shown in Table 3. In Table 3, "×1" and "×2" respectively represent twice and three times the number of semantic-aware synthetic images. The experiment before "-" only generates one semantic-aware synthetic image for each original image. As shown in Table 3, during the training process, one semantic-aware synthetic image is randomly selected from multiple semantic-aware synthetic images belonging to the same image to ensure the constancy of the computational load. As the number of semantic-aware synthetic images increases, the accuracy of the model is further improved, indicating the great potential of the present invention. The fact that semantic enhancement or visual cue enhancement alone can improve accuracy proves the rationality of the present invention.

[0083] Table 3. Ablation experiments of semantic enhancement and visual cue enhancement of the present invention

[0084] Semantic Enhancement Visual Enhancement Linear classification accuracy - - 78.4 ×1 - 80.1 - ×1 80.4 ×1 ×1 81.2 ×2 ×2 81.9

[0085] It should also be noted that the method for generating positive image samples to realize image classification in the above embodiment can essentially be executed by a computer program or module. Therefore, similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a system for generating positive image samples to realize image classification corresponding to the method for generating positive image samples to realize image classification provided in the above embodiment, such as Figure 4 As shown, it includes:

[0086] A first model training module is used to generate positive samples by guiding the denoising diffusion model based on the acquired original image, and to perform unsupervised training on the image encoder in a contrastive learning manner based on the generated positive samples;

[0087] The second model training module is used to cascade the image encoder obtained by unsupervised training and a linear classifier to be trained to form an image classification model, train the image classification model in a supervised manner, fix the parameters of the image encoder in the image classification model, and optimize the parameters of the linear classifier so that it can realize image category judgment;

[0088] The result acquisition module is used to input the image to be classified into the trained image classification model and output the classification result of the image to be classified.

[0089] It is understandable that the method of generating positive image samples to achieve image classification described in S1 to S3 above can be substantially implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the method of generating positive image samples to achieve image classification provided in the above embodiment, which includes a computer program / instruction. When the computer program / instruction is executed by a processor, the method of generating positive image samples to achieve image classification as described in the above embodiment can be implemented.

[0090] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the method for generating positive image samples to implement image classification provided in the above embodiment, such as Figure 5 As shown, it includes a memory and a processor;

[0091] The memory is used to store computer programs;

[0092] The processor is used to implement the method of generating positive image samples to realize image classification in the above embodiment when executing the computer program.

[0093] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.

[0094] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the method for generating image positive samples to realize image classification provided in the above embodiment, and a computer program is stored on the storage medium. When the computer program is executed by the processor, it can realize the method for generating image positive samples to realize image classification in the above embodiment.

[0095] It is understandable that the above storage medium may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. The storage medium may also be a U disk, a mobile hard disk, a magnetic disk or an optical disk, etc., which can store program codes.

[0096] It is understandable that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0097] It should also be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0098] The above-described embodiment is only a preferred solution of the present invention, but it is not intended to limit the present invention. A person skilled in the relevant technical field may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent replacement or equivalent transformation falls within the protection scope of the present invention.

Claims

1. A method for generating positive image samples to achieve image classification, characterized in that: The following steps are involved: S1. Based on the acquired original image, the denoising diffusion model is guided to generate positive samples, and the image encoder is unsupervisedly trained in a contrastive learning manner based on the generated positive samples; S2. Concatenate the image encoder obtained by unsupervised training and a linear classifier to be trained to form an image classification model, train the image classification model in a supervised manner, fix the parameters of the image encoder in the image classification model, and optimize the parameters of the linear classifier so that it can realize image category judgment; S3. Input the image to be classified into the trained image classification model and output the classification result of the image to be classified.

2. The method for generating positive image samples for image classification according to claim 1, characterized in that: In step S1, the specific process of generating positive samples is: S11. Inputting the original image and a sentence for prompting semantic extraction into the pre-trained multimodal large model to generate a first feature description of the original image in the text space; S12. sending the first feature description and a sentence for prompting semantic rewriting into a large language model, performing semantic rewriting on the first feature description, and generating a second feature description of the original image in the text space; S13. Adding Gaussian noise to the original image to generate a random noise image, obtaining a mask image of the original image and a target image, weightedly mixing the mask image of the original image and the target image according to pixel values ​​to generate a mask mixed image, using one of the original image, the random noise image, the mask image of the original image, and the mask mixed image as a first visual cue, and performing data enhancement on the first visual cue to obtain a second visual cue; S14. Input the second feature description and the second visual cue into the denoising diffusion model to generate a semantically-aware synthetic image, and use the semantically-aware synthetic image as a positive sample that is semantically consistent with the original image.

3. The method for generating positive image samples to realize image classification according to claim 2, characterized in that: In step S1, the specific process of unsupervised training of the image encoder is: In each iteration round, the original image is data enhanced to generate a first processed image, the semantically-aware synthetic image is data enhanced to generate a second processed image, the first processed image is input into the image encoder in the MoCo contrastive learning framework for feature extraction to obtain a first feature map, the second processed image is input into the momentum image encoder in the MoCo contrastive learning framework for feature extraction to obtain a second feature map, the contrast loss is calculated based on the first feature map and the second feature map, and the training is continuously iterated. After reaching the preset iteration rounds, the training of the image encoder is completed.

4. The method for generating positive image samples for image classification according to claim 1, characterized in that: In step S2, during the process of training the linear classifier, the original image is input into the trained image encoder, the features output by the image encoder are passed through the linear classifier, the output dimension of the linear classifier is set to the number of categories, the linear classifier outputs the image category prediction result, the true label corresponding to the original image is obtained, the first cross entropy loss between the true label and the image category prediction result is calculated, and the parameters of the linear classifier are updated based on minimizing the first cross entropy loss. The training is continuously iterated, and after reaching a preset number of iterations, the linear classifier converges to obtain a trained linear classifier.

5. The method for generating positive image samples for image classification according to claim 1 or 4, characterized in that: The linear classifier adopts one fully connected layer.

6. The method for generating positive image samples for image classification according to claim 4, characterized in that: The output dimension of the linear classifier is set to 100.

7. A system for generating positive image samples to achieve image classification, characterized in that: include: A first model training module is used to generate positive samples by guiding the denoising diffusion model based on the acquired original image, and to perform unsupervised training on the image encoder in a contrastive learning manner based on the generated positive samples; The second model training module is used to cascade the image encoder obtained by unsupervised training and a linear classifier to be trained to form an image classification model, train the image classification model in a supervised manner, fix the parameters of the image encoder in the image classification model, and optimize the parameters of the linear classifier so that it can realize image category judgment; The result acquisition module is used to input the image to be classified into the trained image classification model and output the classification result of the image to be classified.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for generating positive image samples and realizing image classification as described in any one of claims 1 to 6 can be implemented.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for generating positive image samples and realizing image classification as claimed in any one of claims 1 to 6 is implemented.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to implement the method for generating positive image samples and realizing image classification as claimed in any one of claims 1 to 6 when executing the computer program.

Citation Information

Patent Citations

  • Efficient medical image marking and learning system

    CN113314205A

  • Methods, systems, media, and apparatus for unsupervised action detection

    CN118194099A

  • Multi-modal large language model for realizing fine-grained visual perception of remote sensing image

    CN119027960A

  • Method and apparatus for vision-language understanding

    US20240144651A1

  • Endoscope image classification model training method and device, and endoscope image classification method

    WO2023030521A1

Cited By

  • Power grid system-oriented interpretability analysis method, system, equipment and medium

    CN122222334A