Small sample image amplification method and system
By acquiring various types of original image data sets and small sample target image data sets, preprocessing and training, combining discrete codebooks and code prediction modules, amplified images are generated, which solves the problems of difficulty in obtaining small sample image data in the prior art and low similarity in image generation, and achieves high-quality and diverse image generation.
Patent Information
- Application Number
- CN202510103060.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, when generating training data, it is difficult to obtain small sample image data, and the transformation-based method results in low similarity and limited diversity of the generated images.
By obtaining multiple types of original image data sets and small sample target image data sets, preprocessing is performed to obtain types-related editing vectors, combining discrete codebooks and code prediction modules, one-stage training and two-stage training are performed to generate amplified images.
It realizes the generation of multiple types of similar images under weak supervision, overcomes the problems of disappearance and collapse of generated image details, and improves the diversity and quality of generated images.
Smart Images

Figure CN119992255A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and more specifically, to a small sample image amplification method and system. Background Art
[0002] With the rapid development of computer vision, target detection technology has made significant progress in the past few years, especially in the field of deep learning. These methods have the ability to accurately and efficiently identify endangered animals, but these target detection models can only provide relatively reliable target recognition performance when they have abundant training data. When faced with an unseen environment or rare obstacles, the training effect is difficult to guarantee due to the lack of corresponding data samples in the model training process, the uneven data quality and the complexity of environmental factors. In addition, the target detection model is affected by multiple factors such as complex background, multiple targets, day and night brightness in actual scenes, resulting in large uncertainty in application, and its detection accuracy will be significantly reduced. Humans are required to distinguish whether it is correct. Manual recognition of these data is not only inefficient, but also prone to errors. Therefore, training data directly affects the accuracy of the target detection model.
[0003] For the small sample image generation algorithms used when generating training data, the existing methods are fusion-based methods and meta-learning-based methods. However, their effectiveness depends on the substantial similarity between the input images, and multiple inputs are required, and the output images can only be limited to the features of the original images. On the other hand, the transformation-based method does not change the content of the image itself, and is the simplest way to enhance the image dataset. Excessive transformation may result in limited diversity of the dataset and generate low-value data. Because this method attempts to identify and apply intra-class changes to unobserved category instances, thereby synthesizing other images of the same category.
[0004] The prior art discloses a method for generating skin disease images based on a convolutional neural network and Prompt. The method includes: collecting a skin disease image dataset containing different parts as training data and preprocessing it; building a skin disease image generation model; using the preprocessed skin disease image dataset to train StyleGAN2 and generate skin disease images; using the generated skin disease images and the text descriptions associated with them as training data for CLIP to train CLIP; using the trained CLIP to assist StyleGAN2 in obtaining the final desired skin disease image data; and evaluating and verifying the generated skin disease image data. The similarity of the images generated by this method according to the type is low. Summary of the invention
[0005] The present invention aims to solve the defects of the prior art that the training data is difficult to obtain and the similarity of the images generated according to the categories is low, and provides a small sample image amplification method and system. The amplification method can generate multiple categories at the same time, and multiple similar images of each category.
[0006] The primary purpose of the present invention is to solve the above technical problems. The technical solution of the present invention is as follows:
[0007] A small sample image amplification method, comprising:
[0008] S1: Obtain a multi-category original image dataset, including a number of original images of multiple categories; obtain a small sample target image dataset, including a number of small sample target images and corresponding prompt word libraries;
[0009] S2: preprocessing the plurality of types of original image data sets to obtain type-related editing vectors;
[0010] S3: performing one-stage training on the discrete codebook according to the plurality of small sample target images and the type-related edit vectors to obtain a trained discrete codebook;
[0011] S4: training the established code prediction module according to the plurality of small sample target images and the corresponding prompt word library and category-related edit vectors to obtain a trained code prediction module;
[0012] S5: inputting the amplified small sample image and the prompt word into the trained code prediction module to obtain an index sequence;
[0013] S6: Matching the index sequence in a trained discrete codebook to obtain a reconstruction vector of a corresponding position of the index sequence in the trained discrete codebook;
[0014] S7: Generate an enlarged image corresponding to the small sample image using the reconstruction vector.
[0015] Furthermore, in step S2, the plurality of types of original image data sets are preprocessed, including:
[0016] S201: inputting all images in the multi-class original image dataset into an image encoder respectively to obtain image embedding vectors corresponding to the images in the multi-class original image dataset;
[0017] S202: averaging the image embedding vectors corresponding to the images in all the multiple categories of original image data sets to obtain category-related editing vectors.
[0018] Furthermore, in step S3, a one-stage training is performed on the discrete codebook, including:
[0019] S301: selecting one of a plurality of small sample target images as a first image;
[0020] S302: Input the first image into an image encoder to obtain a first image embedding vector;
[0021] S303: Calculate the first image embedding vector and the type-related edit vector to obtain a second image embedding vector;
[0022] S304: inputting the second image embedding vector into a discrete codebook to obtain a first reconstruction edit vector;
[0023] S305: inputting the first reconstruction editing vector into an image generator to obtain a first reconstructed image;
[0024] S306: Calculate a first loss function according to the first image and the first reconstructed image, and optimize a discrete codebook;
[0025] S307: Select another image from the plurality of small sample target images as a new first image; repeat steps S302 to S306 until the first loss function is minimized, and a trained discrete codebook is obtained.
[0026] Furthermore, in step S303, the second image embedding vector is obtained by calculating the first image embedding vector and the type-related edit vector, including:
[0027]
[0028] represents the i-th second image embedding vector, i represents the image sequence number, represents the category-dependent edit vector, represents the i-th first image embedding vector.
[0029] Furthermore, the first loss function is as follows:
[0030]
[0031] represents the first reconstructed image, represents the first image, represents the category-dependent edit vector, Represents the embedding vector of the i-th second image.
[0032] Furthermore, in step S4, the training of the established code prediction module includes:
[0033] S401: selecting an image from a multi-category original image data set and a prompt word from a corresponding prompt library as a second image and a first prompt word;
[0034] S402: Input the second image into an image encoder to obtain a third image embedding vector;
[0035] S403: inputting the first prompt word into a prompt word module to obtain a first prompt word embedding vector;
[0036] S404: Obtain the fourth embedding vector by calculating using the category-related edit vector and the third image embedding vector;
[0037] S405: Input the fourth embedding vector and the first prompt word embedding vector into a code prediction module to obtain a first index sequence;
[0038] S406: Input the first index sequence into a trained discrete codebook to obtain a first reconstruction vector;
[0039] S407: inputting the first reconstruction vector into an image generator to obtain a second reconstructed image;
[0040] S408: Calculating a second loss function according to the second image and the second reconstructed image, and optimizing a code prediction module;
[0041] S409: Select another image in the multi-class original image data set and a prompt word in the corresponding prompt library as a new second image and a new first prompt word; repeat steps S402 to S408 until the second loss function is minimized.
[0042] Further, the code prediction module includes: a first position encoding layer, a second position encoding layer, a third position encoding layer, a first transformer layer, a second transformer layer, a third transformer layer, a linear layer, and a normalization layer;
[0043] The fourth embedding vector and the first prompt word embedding vector are input into the first position encoding layer, the second position encoding layer, and the third position encoding layer for connection. The output end of the first position encoding layer is connected to the input end of the first transformer layer, the output end of the second position encoding layer and the output end of the first transformer are connected to the input end of the second transformer layer, the output end of the third position encoding layer and the output end of the second transformer layer are connected to the input end of the third transformer layer, the output end of the third transformer layer is connected to the input end of the normalization layer, the output end of the normalization layer is connected to the input end of the linear layer, and the output end of the linear layer outputs the first index sequence.
[0044] Further, the prompt word module includes: a CLIP prompt word encoding layer, a prompt word encoding layer, a cross attention layer, a first CLIP image encoding layer, a second CLIP image encoding layer, a cosine similarity comparison layer, a first decoding layer, and a second decoding layer;
[0045] The prompt word is input to the input end of the CLIP prompt word encoding layer, the output end of the CLIP prompt word encoding layer is connected to the input end of the prompt word encoding layer, the output end of the prompt word encoding layer is connected to the input end of the cross-attention layer, the output end of the cross-attention layer is connected to the input end of the code prediction module, the output end of the code prediction module is connected to the input end of the first decoding layer, the output end of the first decoding layer is connected to the input end of the first CLIP prompt word encoding layer, the output end of the second decoding layer is connected to the input end of the second CLIP prompt word encoding layer, the output end of the first CLIP prompt word encoding layer and the output end of the second CLIP prompt word encoding layer are connected to the input end of the cosine similarity comparison layer.
[0046] Furthermore, the second loss function is as follows:
[0047]
[0048] Q i , K i 、V i Represents the parameters of the transformer layer of the i-th image, i represents the image number, sg represents the stop gradient operator, represents the first reconstruction vector of the i-th image, represents the i-th fourth embedding vector.
[0049] A small sample image amplification system, comprising:
[0050] Dataset acquisition module: obtain multiple types of original image datasets, including several original images of multiple types; obtain small sample target image datasets, including several small sample target images and corresponding prompt word libraries;
[0051] Preprocessing module: preprocessing the plurality of types of original image data sets to obtain type-related editing vectors;
[0052] One-stage training module: performing one-stage training on the discrete codebook according to the plurality of small sample target images and the type-related edit vectors to obtain a trained discrete codebook;
[0053] The second-stage training module: training the established code prediction module according to the several small sample target images and the corresponding prompt word library and the category-related editing vectors to obtain a trained code prediction module;
[0054] Index sequence acquisition module: input the amplified small sample image and prompt words into the trained code prediction module to obtain the index sequence;
[0055] Discrete codebook matching module: matches the index sequence in a trained discrete codebook to obtain a reconstruction vector of a corresponding position of the index sequence in the trained discrete codebook;
[0056] Image generation module: using the reconstruction vector to generate an enlarged image corresponding to the small sample image.
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] The present invention mines the editing vectors related to the category from a large number of easily available pictures, and performs attribute editing on the image under weak supervision, so as to achieve the purpose of generating small sample images. The present invention can avoid the complex and unstable training and transformation process. The present invention utilizes the image data generation model to overcome the shortcomings of disappearance and collapse of details in the generated image, which seriously affects the perceived image quality; at the same time, it can also make the generated image have better diversity. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A flowchart of a small sample image amplification method provided in Example 1.
[0060] Figure 2 This is a flowchart for preprocessing multiple types of original image data sets provided in Example 1.
[0061] Figure 3 This is a flow chart of performing one-stage training on a discrete codebook provided in Example 1.
[0062] Figure 4 A schematic diagram of the principle of performing one-stage training on a discrete codebook provided in Example 1.
[0063] Figure 5 A schematic diagram of the principle of performing one-stage training on a discrete codebook provided in Example 1.
[0064] Figure 6 A flowchart of training the established code prediction module provided in Example 1.
[0065] Figure 7 A schematic diagram of the principles of training the established code prediction module provided in Example 1.
[0066] Figure 8 A schematic diagram of the principles of training the established code prediction module provided in Example 1.
[0067] Fig. 9 This is a structural diagram of the code prediction module provided in Example 1.
[0068] Fig.10 This is a structural diagram of the prompt word module provided in Example 1.
[0069] Fig.11 This is a schematic diagram of the principle of the prompt word module provided in Example 1. DETAILED DESCRIPTION
[0070] The drawings are for illustrative purposes only and should not be construed as limiting the present patent;
[0071] In order to better illustrate the present embodiment, some parts in the drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product;
[0072] It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0073] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0074] Example 1
[0075] like Figure 1 As shown, a small sample image amplification method includes:
[0076] S1: Obtain a multi-category original image dataset, including a number of original images of multiple categories; obtain a small sample target image dataset, including a number of small sample target images and corresponding prompt word libraries;
[0077] S2: preprocessing the plurality of types of original image data sets to obtain type-related editing vectors;
[0078] S3: performing one-stage training on the discrete codebook according to the plurality of small sample target images and the type-related edit vectors to obtain a trained discrete codebook;
[0079] S4: training the established code prediction module according to the plurality of small sample target images and the corresponding prompt word library and category-related edit vectors to obtain a trained code prediction module;
[0080] S5: inputting the amplified small sample image and the prompt word into the trained code prediction module to obtain an index sequence;
[0081] S6: Matching the index sequence in a trained discrete codebook to obtain a reconstruction vector of a corresponding position of the index sequence in the trained discrete codebook;
[0082] S7: Generate an enlarged image corresponding to the small sample image using the reconstruction vector.
[0083] It should be noted that the generator is used in conjunction with the encoder. In one embodiment, the encoder uses P2P and the generator uses styleGAN.
[0084] Furthermore, if Figure 2 As shown, in step S2, the plurality of types of original image data sets are preprocessed, including:
[0085] S201: inputting all images in the multi-class original image dataset into an image encoder respectively to obtain image embedding vectors corresponding to the images in the multi-class original image dataset;
[0086] S202: averaging the image embedding vectors corresponding to the images in all the multiple categories of original image data sets to obtain category-related editing vectors.
[0087] It should be noted that the following formula is used to take the average:
[0088]
[0089] N m represents the total number of image embedding vectors, Represents the image embedding vector.
[0090] Furthermore, if Figure 3 , Figure 4 , Figure 5 As shown, in step S3, the discrete codebook is trained in one stage, including:
[0091] S301: selecting one of a plurality of small sample target images as a first image;
[0092] S302: Input the first image into an image encoder to obtain a first image embedding vector;
[0093] S303: Calculate the first image embedding vector and the type-related edit vector to obtain a second image embedding vector;
[0094] S304: inputting the second image embedding vector into a discrete codebook to obtain a first reconstruction edit vector;
[0095] S305: inputting the first reconstruction editing vector into an image generator to obtain a first reconstructed image;
[0096] S306: Calculate a first loss function according to the first image and the first reconstructed image, and optimize a discrete codebook;
[0097] S307: Select another image from the plurality of small sample target images as a new first image; repeat steps S302 to S306 until the first loss function is minimized, and a trained discrete codebook is obtained.
[0098] It should be noted that the use of discrete codebooks to reduce the dimension and store semantic information reduces the disappearance and collapse of animal organs compared to images generated by editing-based methods. The advantage of discrete codebooks is that smaller finite discrete spaces have better robustness and reconstruction quality than continuous infinite spaces. When inputting images of unseen categories, the encoder may generate ambiguous latent codes, which seriously affects subsequent editing and generation. The discrete structure forces the codebook to retain high-quality details, and the low-quality latent codes have a higher probability of matching with the accurate codes in the codebook.
[0099] It should be noted that in the first stage of training, the discrete codebook uses nearest neighbor matching for matching training, and the formula is as follows:
[0100]
[0101] c k represents the element of the kth discrete codebook in the discrete codebook.
[0102] Furthermore, in step S303, the second image embedding vector is obtained by calculating the first image embedding vector and the type-related edit vector, including:
[0103]
[0104] represents the i-th second image embedding vector, i represents the image sequence number, represents the category-dependent edit vector, represents the i-th first image embedding vector.
[0105] Furthermore, the first loss function is as follows:
[0106]
[0107] represents the first reconstructed image, represents the first image, represents the category-dependent edit vector, Represents the embedding vector of the i-th second image.
[0108] Furthermore, if Figure 6 , Figure 7 , Figure 8 As shown, in step S4, the training of the established code prediction module includes:
[0109] S401: selecting an image from a multi-category original image data set and a prompt word from a corresponding prompt library as a second image and a first prompt word;
[0110] S402: Input the second image into an image encoder to obtain a third image embedding vector;
[0111] S403: inputting the first prompt word into a prompt word module to obtain a first prompt word embedding vector;
[0112] S404: Obtain the fourth embedding vector by calculating using the category-related edit vector and the third image embedding vector;
[0113] S405: Input the fourth embedding vector and the first prompt word embedding vector into a code prediction module to obtain a first index sequence;
[0114] S406: Input the first index sequence into a trained discrete codebook to obtain a first reconstruction vector;
[0115] S407: inputting the first reconstruction vector into an image generator to obtain a second reconstructed image;
[0116] S408: Calculating a second loss function according to the second image and the second reconstructed image, and optimizing a code prediction module;
[0117] S409: Select another image in the multi-class original image data set and a prompt word in the corresponding prompt library as a new second image and a new first prompt word; repeat steps S402 to S408 until the second loss function is minimized.
[0118] The formula for obtaining the fourth embedding vector by using the category-related edit vector and the third image embedding vector is as follows:
[0119]
[0120] represents the i-th fourth embedding vector, i represents the image number, represents the category-dependent edit vector, represents the i-th third image embedding vector.
[0121] Furthermore, if Fig. 9 As shown, the code prediction module includes: a first position encoding layer, a second position encoding layer, a third position encoding layer, a first transformer layer, a second transformer layer, a third transformer layer, a linear layer, and a normalization layer;
[0122] The fourth embedding vector and the first prompt word embedding vector are input into the first position encoding layer, the second position encoding layer, and the third position encoding layer for connection. The output end of the first position encoding layer is connected to the input end of the first transformer layer, the output end of the second position encoding layer and the output end of the first transformer are connected to the input end of the second transformer layer, the output end of the third position encoding layer and the output end of the second transformer layer are connected to the input end of the third transformer layer, the output end of the third transformer layer is connected to the input end of the normalization layer, the output end of the normalization layer is connected to the input end of the linear layer, and the output end of the linear layer outputs the first index sequence.
[0123] The fourth embedding vector and the first prompt word embedding vector are first summed and then used as inputs of the first position encoding layer, the second position encoding layer, and the third position encoding layer.
[0124] It should be noted that the introduction of the code prediction module solves the problem of lack of diversity in generated images caused by using the codebook alone, and optimizes the generalization and robustness of the model.
[0125] Furthermore, if Fig.10 , Fig.11 As shown, the prompt word module includes: a CLIP prompt word encoding layer, a prompt word encoding layer, a cross attention layer, a first CLIP image encoding layer, a second CLIP image encoding layer, a cosine similarity comparison layer, a first decoding layer, and a second decoding layer;
[0126] The prompt word is input to the input end of the CLIP prompt word encoding layer, the output end of the CLIP prompt word encoding layer is connected to the input end of the prompt word encoding layer, the output end of the prompt word encoding layer is connected to the input end of the cross-attention layer, the output end of the cross-attention layer is connected to the input end of the code prediction module, the output end of the code prediction module is connected to the input end of the first decoding layer, the output end of the first decoding layer is connected to the input end of the first CLIP prompt word encoding layer, the output end of the second decoding layer is connected to the input end of the second CLIP prompt word encoding layer, the output end of the first CLIP prompt word encoding layer and the output end of the second CLIP prompt word encoding layer are connected to the input end of the cosine similarity comparison layer.
[0127] In a specific embodiment, the loss function of the prompt word module is:
[0128] L sub =1-cos{E clip (G ′ (w ′ ),E clip (G ′ (w)))}
[0129] G′ represents the processing result of the decoding layer, w ′ represents the output of the first decoding layer, w represents the output of the second decoding layer, cos{} represents the cosine similarity comparison, E clip Represents the processing results of the CLIP prompt word encoding layer.
[0130] It should be noted that the use of the prompt word learning module introduces the prior knowledge of the pre-trained model, enabling the model to achieve stable attribute editing and improve the quality of the generated image. Without using the prompt word learning module, the purpose of generating small sample images can also be achieved by relying only on the discrete codebook and code prediction module.
[0131] This method can perform targeted editing on the input original image according to the preset prompt words. The model will randomly select prompt words and input them into CLIP to encode them into text vectors. The text vectors and the decomposed category-independent vectors are input into the code prediction module to generate the required additional data set.
[0132] When the generalization of the model needs to be improved, the prompt word module is removed or blank prompt words are added during training, and only the code prediction module is used for generation. In this case, the images generated by the model are more diverse, but the difference between the generated images and the original images will be greater.
[0133] The prompt words include category information, color characteristics, shape characteristics, and environmental or background elements.
[0134] Furthermore, the second loss function is as follows:
[0135]
[0136] Q i , K i 、V i Represents the parameters of the transformer layer of the i-th image, i represents the image number, sg represents the stop gradient operator, represents the first reconstruction vector of the i-th image, represents the i-th fourth embedding vector.
[0137] A small sample image amplification system, comprising:
[0138] Dataset acquisition module: obtain multiple types of original image datasets, including several original images of multiple types; obtain small sample target image datasets, including several small sample target images and corresponding prompt word libraries;
[0139] Preprocessing module: preprocessing the plurality of types of original image data sets to obtain type-related editing vectors;
[0140] One-stage training module: performing one-stage training on the discrete codebook according to the plurality of small sample target images and the type-related edit vectors to obtain a trained discrete codebook;
[0141] The second-stage training module: training the established code prediction module according to the several small sample target images and the corresponding prompt word library and the category-related editing vectors to obtain a trained code prediction module;
[0142] Index sequence acquisition module: input the amplified small sample image and prompt words into the trained code prediction module to obtain the index sequence;
[0143] Discrete codebook matching module: matches the index sequence in a trained discrete codebook to obtain a reconstruction vector of a corresponding position of the index sequence in the trained discrete codebook;
[0144] Image generation module: using the reconstruction vector to generate an enlarged image corresponding to the small sample image.
[0145] Example 2
[0146] Based on the small sample image amplification method described in Example 1, that is, this embodiment adopts the same small sample image amplification method as Example 1. It can be applied to the field of endangered animal image generation.
[0147] The monitoring and protection of endangered animals is crucial to maintaining ecological balance and biodiversity. Traditional monitoring methods such as manual patrols and camera traps are inefficient and may cause interference to animals, as these animals often live in remote or inaccessible areas, and it is difficult to cover large natural areas. Therefore, drones equipped with high-resolution cameras, infrared cameras and GPS positioning systems are commonly used to capture images and location data of endangered animals, and image recognition technology is used to process the images sent back by drones to obtain the species and number of endangered animals.
[0148] Traditional deep learning target detection algorithms for wildlife directly predict the category and location of objects without generating candidate regions, and have extremely fast detection speeds. Although this method has the ability to accurately and efficiently identify endangered animals, it can only provide relatively reliable target recognition performance when there is abundant training data. However, since animals often live in remote or inaccessible areas, training data is difficult to obtain, resulting in low detection performance of target detection algorithms.
[0149] The present invention combines the discrete codebook space characteristics with the code prediction module to overcome the disappearance and collapse of animal organs while maintaining good diversity. This method is particularly suitable for solving the problem of scarce endangered animal samples, helping to achieve accurate target detection of endangered animals, and providing a key tool for wildlife protection and ecological balance research.
[0150] The present invention exhibits strong cross-scene migration capabilities, can effectively cope with the interference of various complex background information in real field scenes, has good robustness, and has strong adaptability to field environments. Its adaptability to data-scarce environments reduces reliance on manual labeling.
[0151] The same or similar reference numerals correspond to the same or similar components;
[0152] The terms used in the drawings to describe positional relationships are only used for illustrative purposes and should not be construed as limiting this patent;
[0153] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A small sample image amplification method, characterized in that: include: S1: Obtain a multi-category original image dataset, including a number of original images of multiple categories; obtain a small sample target image dataset, including a number of small sample target images and corresponding prompt word libraries; S2: preprocessing the plurality of types of original image data sets to obtain type-related editing vectors; S3: performing one-stage training on the discrete codebook according to the plurality of small sample target images and the type-related edit vectors to obtain a trained discrete codebook; S4: training the established code prediction module according to the plurality of small sample target images and the corresponding prompt word library and category-related edit vectors to obtain a trained code prediction module; S5: inputting the amplified small sample image and the prompt word into the trained code prediction module to obtain an index sequence; S6: Matching the index sequence in a trained discrete codebook to obtain a reconstruction vector of a corresponding position of the index sequence in the trained discrete codebook; S7: Generate an enlarged image corresponding to the small sample image using the reconstruction vector.
2. A small sample image amplification method according to claim 1, characterized in that: In step S2, the plurality of types of original image data sets are preprocessed, including: S201: inputting all images in the multi-class original image dataset into an image encoder respectively to obtain image embedding vectors corresponding to the images in the multi-class original image dataset; S202: averaging the image embedding vectors corresponding to the images in all the multiple categories of original image data sets to obtain category-related editing vectors.
3. A small sample image amplification method according to claim 1, characterized in that: In step S3, a one-stage training is performed on the discrete codebook, including: S301: selecting one of a plurality of small sample target images as a first image; S302: Input the first image into an image encoder to obtain a first image embedding vector; S303: Calculate the first image embedding vector and the type-related edit vector to obtain a second image embedding vector; S304: inputting the second image embedding vector into a discrete codebook to obtain a first reconstruction edit vector; S305: inputting the first reconstruction editing vector into an image generator to obtain a first reconstructed image; S306: Calculate a first loss function according to the first image and the first reconstructed image, and optimize a discrete codebook; S307: Select another image from the plurality of small sample target images as a new first image; repeat steps S302 to S306 until the first loss function is minimized, and a trained discrete codebook is obtained.
4. A small sample image amplification method according to claim 3, characterized in that: In step S303, the second image embedding vector is obtained by calculating the first image embedding vector and the type-related edit vector, including: represents the i-th second image embedding vector, i represents the image sequence number, represents the category-dependent edit vector, represents the i-th first image embedding vector.
5. A small sample image amplification method according to claim 3, characterized in that: The first loss function is as follows: represents the first reconstructed image, represents the first image, represents the category-dependent edit vector, Represents the embedding vector of the i-th second image.
6. A small sample image amplification method according to claim 1, characterized in that: In step S4, the training of the established code prediction module includes: S401: selecting an image from a multi-category original image data set and a prompt word from a corresponding prompt library as a second image and a first prompt word; S402: Input the second image into an image encoder to obtain a third image embedding vector; S403: inputting the first prompt word into a prompt word module to obtain a first prompt word embedding vector; S404: Obtain the fourth embedding vector by calculating using the category-related edit vector and the third image embedding vector; S405: Input the fourth embedding vector and the first prompt word embedding vector into a code prediction module to obtain a first index sequence; S406: Input the first index sequence into a trained discrete codebook to obtain a first reconstruction vector; S407: inputting the first reconstruction vector into an image generator to obtain a second reconstructed image; S408: Calculating a second loss function according to the second image and the second reconstructed image, and optimizing a code prediction module; S409: Select another image in the multi-class original image data set and a prompt word in the corresponding prompt library as a new second image and a new first prompt word; repeat steps S402 to S408 until the second loss function is minimized.
7. A small sample image amplification method according to claim 6, characterized in that: The code prediction module includes: a first position encoding layer, a second position encoding layer, a third position encoding layer, a first transformer layer, a second transformer layer, a third transformer layer, a linear layer, and a normalization layer; The fourth embedding vector and the first prompt word embedding vector are input into the first position encoding layer, the second position encoding layer, and the third position encoding layer for connection. The output end of the first position encoding layer is connected to the input end of the first transformer layer, the output end of the second position encoding layer and the output end of the first transformer are connected to the input end of the second transformer layer, the output end of the third position encoding layer and the output end of the second transformer layer are connected to the input end of the third transformer layer, the output end of the third transformer layer is connected to the input end of the normalization layer, the output end of the normalization layer is connected to the input end of the linear layer, and the output end of the linear layer outputs the first index sequence.
8. A small sample image amplification method according to claim 6, characterized in that: The prompt word module includes: a CLIP prompt word encoding layer, a prompt word encoding layer, a cross attention layer, a first CLIP image encoding layer, a second CLIP image encoding layer, a cosine similarity comparison layer, a first decoding layer, and a second decoding layer; The prompt word is input to the input end of the CLIP prompt word encoding layer, the output end of the CLIP prompt word encoding layer is connected to the input end of the prompt word encoding layer, the output end of the prompt word encoding layer is connected to the input end of the cross-attention layer, the output end of the cross-attention layer is connected to the input end of the code prediction module, the output end of the code prediction module is connected to the input end of the first decoding layer, the output end of the first decoding layer is connected to the input end of the first CLIP prompt word encoding layer, the output end of the second decoding layer is connected to the input end of the second CLIP prompt word encoding layer, the output end of the first CLIP prompt word encoding layer and the output end of the second CLIP prompt word encoding layer are connected to the input end of the cosine similarity comparison layer.
9. A small sample image amplification method according to claim 6, characterized in that: The second loss function is as follows: Q i , K i 、V i Represents the parameters of the transformer layer of the i-th image, i represents the image number, sg represents the stop gradient operator, represents the first reconstruction vector of the i-th image, represents the i-th fourth embedding vector.
10. A small sample image amplification system, applied to the amplification method according to any one of claims 1 to 9, characterized in that: include: Dataset acquisition module: obtain multiple types of original image datasets, including several original images of multiple types; obtain small sample target image datasets, including several small sample target images and corresponding prompt word libraries; Preprocessing module: preprocessing the plurality of types of original image data sets to obtain type-related editing vectors; One-stage training module: performing one-stage training on the discrete codebook according to the plurality of small sample target images and the type-related edit vectors to obtain a trained discrete codebook; The second-stage training module: training the established code prediction module according to the several small sample target images and the corresponding prompt word library and the category-related editing vectors to obtain a trained code prediction module; Index sequence acquisition module: input the amplified small sample image and prompt words into the trained code prediction module to obtain the index sequence; Discrete codebook matching module: matches the index sequence in a trained discrete codebook to obtain a reconstruction vector of a corresponding position of the index sequence in the trained discrete codebook; Image generation module: using the reconstruction vector to generate an enlarged image corresponding to the small sample image.
Citation Information
Patent Citations
Natural language driven small sample image generation system and method and storage medium
CN118154715A