Litchi disease and insect pest image generation method based on Stable Diffusion and Grouping

By introducing Stable Diffusion and Grounding technologies into image generation technology, the problems of accuracy of image generation of litchi pests and diseases and lack of data sets are solved, high-quality and diversified image generation is achieved, and agricultural pest detection and prevention are supported.

CN119991849AInactive Publication Date: 2025-05-13SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082020.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing image generation technology is difficult to generate accurate and diverse litchi pest images, and the lack of litchi pest data sets limits the development of pest detection technology.

Method used

Using the lychee pest image generation method based on Stable Diffusion and Grounding, pest targets are segmented through Open-Vocabulary SAM, image features are extracted using CLIP, and text prompts containing target categories, quantity and scene environment are designed, and Grounding information is introduced to guide the image generation process.

Benefits of technology

High-quality and diverse litchi pest images were generated, which improved the authenticity and credibility of the images, accurately expressed the location and categories of pests, effectively expanded the litchi pest data set, and improved agricultural production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991849A_ABST
    Figure CN119991849A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating a litchi disease and insect pest image based on Stable Diffusion and Grouping. According to the method, the Open-Vocabulary SAM is utilized to segment the pest and disease damage target, and the CLIP is utilized to extract the image features, so that the high-quality image containing the detail features of the pest and disease damage can be generated, the authenticity and the credibility of the image can be improved, a large number of high-quality litchi pest and disease damage images can be generated, the scale of the litchi pest and disease damage data set can be effectively expanded, and the image quality can be improved. The problems of difficulty in data acquisition and data shortage of a traditional method are solved, and data support is provided for training and evaluation of a litchi disease and insect pest detection model. The generated high-quality litchi pest and disease damage image can help agricultural technicians to quickly and accurately identify the types and the number of pest and disease damage and take effective prevention and control measures, so that the agricultural production efficiency is improved. By improving the litchi disease and insect pest detection and prevention efficiency, the method can effectively reduce loss caused by litchi disease and insect pests, the yield and fruit quality of litchis are guaranteed, and healthy development of the litchi industry is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of litchi pest and disease identification, and specifically relates to a litchi pest and disease image generation method based on Stable Diffusion and Grounding. Background Art

[0002] Litchi is one of the most important economic crops in tropical and subtropical regions. Its fruit is not only delicious but also rich in various nutrients. However, it may be affected by pests and diseases throughout its growth cycle. Therefore, effective detection of litchi pests and diseases is crucial to ensure yield and fruit quality. In recent years, automatic pest and disease identification methods based on deep learning technology have gradually become the mainstream research direction in this field due to their high efficiency and accuracy. However, the key to achieving this goal lies in a large number of high-quality data sets. However, pests and diseases are strongly affected by environmental conditions such as weather, temperature and vector insects, especially the small targets and random movement of pests, which makes it difficult to obtain a large number of pest and disease image data. Although the data set can be expanded to a certain extent through color transformation (HSV color transformation), Mosaic, Mixup, Fliplr and Flipud data enhancement methods, the image features are less distinguishable from the original data set, and the data is still not rich enough. In recent years, researchers have tried to use image generation technology to expand the scale of the data set. The likelihood-based diffusion model can generate more diverse and high-quality images and has been applied in many fields. To improve the accuracy of weed recognition, researchers used the diffusion model Stable Diffusion to generate weed images. Experiments show that compared with BigGAN, StyleGAN2, and StyleGAN3, the diffusion model achieves the best trade-off between sample fidelity and diversity and obtains the highest FID value. StableDiffusion is a text-to-image generation model based on the diffusion model that can generate high-quality and diverse images.

[0003] However, in complex backgrounds, the delicate objects generated by Stable Diffusion may lack details and perform poorly, and cannot use additional positioning input to guide the generation process and specify the specific location and category of objects in the image. Summary of the invention

[0004] The purpose of the present invention is to provide a litchi pest and disease image generation method based on Stable Diffusion and Grounding in order to solve the above-mentioned problems.

[0005] The technical solution adopted by the present invention is as follows: a litchi pest and disease image generation method based on Stable Diffusion and Grounding, the method comprising the following steps:

[0006] S1: Image acquisition: Collect images containing six types of litchi pests and diseases in the litchi orchard, annotate them, record the categories and bounding box coordinates, and divide the collected image dataset into training set and test set for model training and evaluation;

[0007] S2: Use CLIP Image Encoder to extract image features from litchi pest and disease images;

[0008] S3: Use the Open-Vocabulary SAM model to segment the pest target and background, and extract image features from the segmented image;

[0009] S4: Use CLIP Text Encoder to extract text features from category names; convert bounding box coordinates into Fourier embedding to represent the location information of the object in the image; use Grounding Tokenizer to fuse image features, text features and location information to generate Grounding Tokens;

[0010] S5: Design text prompts containing target categories, quantities, and scene environments to guide Stable Diffusion to generate images; use CLIP Text Encoder to encode text prompts into Caption Tokens;

[0011] S6: Based on the Stable Diffusion architecture, a gated self-attention mechanism is added to accept and utilize GroundingTokens and Caption Tokens to guide the image generation process;

[0012] Freeze the original Stable Diffusion model weights and only train the parameters of the newly added gated self-attention layer and Grounding Tokenizer module;

[0013] S7: Use Refiner to refine the image generated by Stable Diffusion to improve image quality and details;

[0014] S8: Use the litchi pest and disease detection dataset to train the model; use reconstruction loss to measure model performance, and use AdamW optimizer to update model parameters;

[0015] S9: Use FID and CLIP score to evaluate the quality of the generated litchi pest and disease images; the lower the FID value, the closer the distribution of the generated image is to the real image, that is, the higher the quality of the generated image; the higher the CLIP score value, the higher the correlation between the generated image and the original image.

[0016] In a preferred embodiment, in step S2, the litchi disease and pest images are derived from orchard scenes, and a single image often contains multiple diseases and pests. Compared with the litchi fruits and trees in the background, the main litchi diseases and pests are target types not included in the pre-training data set.

[0017] In a preferred embodiment, in step S3, after obtaining the litchi pest and disease image mask using Open-Vocabulary SAM, the target pest and disease is accurately extracted from the original image by Mask Cropping, and then input into CLIP Image Encoder to obtain image embedding; Open-Vocabulary SAM is based on SAM and CLIP, SAM2CLIP and CLIP2SAM are two knowledge transfer modules, the former transfers the knowledge of SAM to CLIP through distillation and a learnable Transformer adapter, and the latter transfers the knowledge of CLIP to SAM, thereby improving the recognition ability of the model; the input image is encoded into image embedding by CLIP Image Encoder, the input coordinate frame is encoded by SAM promptencoder, and then the SAM2CLIP module uses adaptive and distillation methods to bridge the gap in the feature representation of SAM and CLIP learning, the CLIP2SAM module uses the knowledge of CLIP to enhance the recognition ability of SAM Mask Decoder, and SAM Mask Decoder decodes to generate segmentation masks and label annotations.

[0018] In a preferred embodiment, in step S4, CLIP Text Encoder is used to obtain text features; the bounding box is represented by the coordinates of the upper left corner and the lower right corner, and is transformed into Fourier embedding using Fourier Embedder to represent the position information of the object in the image; then it is fused by Grounding tokenizer to generate new feature representation Grounding Tokens; Grounding tokenizer is a multi-layer perceptron that connects the input in the feature dimension; using h e represents Grounding Tokens, then the above process can be expressed as:

[0019] h e=Grounding_tokenizer(f image (x),f text (c),Fourier_embedder(l)) (1).

[0020] In a preferred embodiment, in step S5, the text prompt consists of the pest category, quantity and scene environment, and the specific format is: "A photo of {count}{class A}, {count}{class B}, and {count}{class C}, {scene environments}"; wherein, {scene environments} is used to describe the overall style and aesthetic quality of the image, and is set to: "photorealistic, high resolution, natural lighting, outdoor, litchiorchard"; at the same time, in order to remind the model of elements or features that should be avoided when generating images, a negative prompt is designed, which helps to more accurately control the generated image content; the negative prompt is set to: "cartoon, illustration, digital art, unrealistic colors, lowres, cropped, worst quality, lowquality"; the input text prompt is processed by the CLIP text encoder to obtain caption tokens.

[0021] In a preferred embodiment, in step S6, the GLIGEN method is used to extend StableDiffusion; after the extension, the positioning ability of the model is significantly enhanced, so that it can accept and use positioning inputs such as categories and bounding boxes to guide the image generation process, and generate images based on the grounding input;

[0022] The original Transformer module consists of a self-attention layer and a cross-attention layer; the image features are processed in the self-attention layer; and the text features are integrated with the image features in the cross-attention layer;

[0023] Using v to represent Visual Tokens, these two layers can be written as:

[0024] v=v+SelfAttn(v) (2)

[0025] v=v+CrossAttn(v,h c ) (3)

[0026] Freeze these two attention layers and add a new gated self-attention layer that receives GroundingTokens(g):

[0027] v=v+β·tanh(γ)·TS(SelfAttn(g)) (4)

[0028] where TS() is a token selection operation that considers only v, and γ is a learnable scalar initialized to 0; β varies at inference time for a predetermined sampling to improve quality and controllability, while it is set to 1 during training.

[0029] In a preferred embodiment, in step S7, Refiner is an independent Latent DiffusionModel, and the samples generated by stable diffusion are further subjected to noise-denoising processing to improve the visual quality of the samples; Refiner is a graph-to-graph module, which has strong portability and can be integrated into other generative models as a cascade component; in order to obtain higher image quality, the present invention integrates refiner after stable diffusion, inputs the same text prompt, iteratively refines the latent vector generated by stable diffusion, and decodes it back into the image.

[0030] In a preferred embodiment, in step S8, the model is trained using a litchi pest and disease detection dataset so that the Grounding input is injected into the stable diffusion with other original components unchanged, thereby improving the performance of the generative model in generating images of six types of litchi pests and diseases; the model is fine-tuned on the basis of Stable Diffusion v1.4, the pre-trained weights of Stable Diffusion v1.4 are frozen, and only the parameters of the newly added gated self-attention layer and Grounding Tokenizer module are trained; and the reconstruction loss is used to compare the distance between the real noise and the noise predicted by the model to measure the performance of the model, and the AdamW optimizer is used to calculate the gradient and update the model parameters;

[0031] Specifically, for an image sampled from the training dataset and the corresponding grounding input y, the image is encoded into a latent representation zt, and the model predicts the noise f{θ,θ′}(zt,t,y) at time step t, drawn from a standard normal distribution The real noise ∈ sampled in , calculate the square of the L2 distance between the real noise and the predicted noise, and finally take the expectation of all sampled samples. The specific calculation is as follows, where θ represents the original parameters and θ′ represents all new parameters:

[0032]

[0033] In a preferred embodiment, in step S9, the generated image is evaluated using Fréchet Inception Distance and CLIP score; FID and CLIP score are commonly used to evaluate the performance of the generated model; FID uses the pre-trained Inception V3 model to extract image features, and measures the quality of the generated image by comparing the distribution difference between the generated image and the real image in the feature space; the lower the FID value, the closer the distribution of the generated image is to the real image, that is, the higher the quality of the generated image; the calculation formula of FID is:

[0034] FID=||μ r -μ g || 2 +Tr(∑ r +∑ g -2(∑ r ∑ g ) 1 / 2 ) (6)

[0035] Among them, μ r , μ g are the feature means of the real image and the generated image, ∑ r ,∑ g are the covariance matrices of the real image and the generated image, respectively, and Tr represents the trace of the matrix;

[0036] Through CLIP, a multimodal learning method proposed by OpenAI, the original image and the generated image are mapped to the same embedding space, and their cosine similarity in the embedding space is calculated to obtain the CLIP score between the generated image and the original image. The higher the value, the higher the similarity. The CLIP score focuses on the correlation between the generated image and the original image, while FID focuses more on the visual quality and authenticity of the generated image.

[0037] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0038] 1. In the present invention, Open-Vocabulary SAM is used to segment pest targets, and CLIP is used to extract image features, which can generate high-quality images containing detailed features of pests and diseases, such as lesions, insect bodies, etc., to improve the authenticity and credibility of the images. By designing text prompts containing target categories, quantities, and scene environments, this method can generate a variety of litchi pest images, covering different scenes, different quantities, and combinations of different pest types to meet the needs of different research scenarios. The introduction of Grounding information, including category names, bounding box coordinates, etc., can guide the Stable Diffusion model to accurately represent the location and category of pests and diseases during the image generation process, thereby improving the accuracy and controllability of the image.

[0039] 2. In the present invention, a large number of high-quality litchi pest and disease images can be generated, which effectively expands the scale of litchi pest and disease data sets, solves the problems of data acquisition difficulties and data scarcity in traditional methods, and provides data support for the training and evaluation of litchi pest and disease detection models. The generated high-quality litchi pest and disease images can be used as an auxiliary tool for litchi pest and disease identification and diagnosis, helping agricultural technicians to quickly and accurately identify the types and quantities of pests and diseases, and take effective prevention and control measures to improve agricultural production efficiency. By improving the detection and prevention efficiency of litchi pests and diseases, this method can effectively reduce the losses caused by litchi pests and diseases, ensure litchi yield and fruit quality, and promote the healthy development of the litchi industry.

[0040] 3. In the present invention, the Grounding technology is applied to the image generation model, which effectively improves the quality and accuracy of the model in content generation, and provides new ideas and methods for the development of image generation technology. This method has been successfully applied to litchi pest and disease image generation, proving its scalability. In the future, it can be expanded to other fields, such as medical image generation, remote sensing image generation, etc., to promote the application of image generation technology in more fields. This method combines images, texts and Grounding information to achieve multimodal learning, providing new cases and experiences for the development of multimodal learning technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 A schematic diagram of the structure of a generated model architecture diagram of the present invention;

[0042] Figure 2 Schematic diagram of the structure of adding a gated self-attention layer in the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0044] Example:

[0045] Reference Figure 1-2 ,A litchi pest and disease image generation method based on Stable Diffusion and Grounding.

[0046] The technical problem to be solved by this invention is that the existing image generation technology is difficult to generate accurate and diverse litchi pest and disease images, and there is a lack of image generation methods for litchi pest and disease. The lack of litchi pest and disease datasets limits the development of pest and disease detection technology. Therefore, this method aims to generate litchi pest and disease images in real orchard scenes based on the core architecture of Stable Diffusion. The overall architecture is as follows Figure 2 As shown. This method uses the target detection dataset, and the Grounding information in the image is represented by the target bounding box coordinates. The input image, target category and coordinates are processed by CLIP (Contrastive Language-Image Pretraining), Open-Vocabulary SAM and Grounding_tokenizer to obtain Grounding Tokens. The text prompts are encoded by the Clip text encoder and converted into caption tokens, which together with the Grounding Tokens guide the model to sample in the latent space and iteratively generate the corresponding latent vector. The latent vector is refined and decoded by Refiner to obtain the generated image.

[0047] The present invention discloses a litchi pest and disease image generation method based on Stable Diffusion and Grounding, aiming to generate high-quality litchi pest and disease images with more precise control, richer expression methods and stronger multimodal understanding capabilities, and provide data support for the identification, diagnosis and prevention of litchi pest and disease. For complex scenes or delicate objects, the images generated by Stable Diffusion may lack details, or perform poorly on fine features, appear blurry or unclear, and cannot specify the specific location and category of objects in the image, and the controllability is not high. Combining grounding information can improve the quality and accuracy of Stable Diffusion in content generation, and solve the limitations of traditional generation models in processing complex scenes.

[0048] At the same time, Open-Vocabulary SAM is introduced to process images of litchi pests and diseases with complex backgrounds and multiple types, so as to focus on the detailed features of pests and diseases, and help the model accurately represent the details of pests and diseases when generating images.

[0049] The present invention combines images, text prompts and Grounding information to guide the litchi pest image generation process, segments litchi pests from images through Open-Vocabulary SAM, and extracts image features and text features through CLIP. Fourier Embedder is used to convert bounding box coordinates into Fourier embedding, CLIP Text Encoder is used to extract text features from class names, and Grounding Tokenizer is used to fuse features to obtain Grounding Tokens.

[0050] Add a gated self-attention mechanism to fuse Grounding Tokens into stable diffusion. Refiner refines the image generated by Stable Diffusion to improve the quality and details of the generated image, and decodes the refined image into the final generated image. This method can generate high-quality and diverse litchi pest and disease images with higher authenticity and detail accuracy, which can be used to assist in the identification, diagnosis and prevention of litchi pests and diseases and improve agricultural production efficiency.

[0051] S1: CLIP image embedding,:

[0052] Image embedding is extracted from litchi pest and disease images by CLIP Image Encoder. Litchi pest and disease images usually come from orchard scenes, and a single image often contains multiple pests and diseases, and the background is complex. Compared with the litchi fruits and trees in the background, the main litchi diseases and pests are unfamiliar to Stable Diffusion and are target types not included in its pre-training data set. Therefore, the present invention introduces the Open-Vocabulary SAM model to segment the pest and disease targets from the background, focusing on the detailed features of the pests and diseases (such as lesions, insect bodies, etc.), in order to retain these subtle features during the image generation process, and improve the ability of Stable Diffusion to accurately generate litchi pest and disease images. After obtaining the litchi pest and disease image mask using Open-VocabularySAM, the target pest and disease is accurately extracted from the original image by Mask Cropping, and then input into CLIP Image Encoder to obtain image embedding. Open-Vocabulary SAM is based on SAM and CLIP. SAM2CLIP and CLIP2SAM are two knowledge transfer modules. The former transfers the knowledge of SAM to CLIP through distillation and learnable Transformer adapters, while the latter transfers the knowledge of CLIP to SAM, thereby improving the recognition ability of the model. The input image is encoded as an image embedding by CLIP Image Encoder, and the input coordinate box is encoded by SAM prompt encoder. Then the SAM2CLIP module uses adaptation and distillation methods to bridge the gap in the feature representation of SAM and CLIP learning. The CLIP2SAM module uses the knowledge of CLIP to enhance the recognition ability of SAM Mask Decoder, which decodes and generates segmentation masks and label annotations.

[0053] S2: Grounding Tokens:

[0054] The main purpose of the Grounding process is to establish a connection between the text description and the specific area in the image, providing a way for the model to align language information with visual information, and helping the model learn the correspondence between the text description and the image area. The model first obtains the positioning information of the image (bounding box coordinates), which provides the spatial position and semantic content of different areas in the image, and transforms and fuses them to obtain Grounding Tokens. Specifically, for the input image x, the category name c and its corresponding bounding box l, the image feature v is obtained by CLIP Image Encoder, and the text feature is obtained by CLIP Text Encoder. The bounding box is represented by the coordinates of the upper left corner and the lower right corner, and is transformed into a Fourier embedding using FourierEmbedder to represent the location information of the object in the image. It is then fused by the Groundingtokenizer to generate a new feature representation Grounding Tokens. The Grounding tokenizer is a multi-layer perceptron that connects the input in the feature dimension. Use h e Represents GroundingTokens, then the above process can be expressed as:

[0055] h e =Grounding_tokenier(f image (x),f text (c),Fourier_embedder(l)) (1)

[0056] S3: Text prompt:

[0057] Text prompts guide Stable Diffusion to generate images. In order to generate accurate and clear litchi pest and disease images in orchard scenes containing multiple pest and disease categories, this method designs text prompts containing target categories and quantities. The text prompts consist of pest and disease categories (class), quantities (count), and scene environments (scene environments). The specific format is: "A photo of {count}{class A}, {count}{class B}, and {count}{class C}, {scene environments}". Among them, {scene environments} is used to describe the overall style and aesthetic quality of the image, which is set to: "photorealistic, high resolution, natural lighting, outdoor, litchi orchard". At the same time, in order to remind the model of elements or features that should be avoided when generating images, a negative prompt is designed to help control the generated image content more accurately. The negative prompt is set to: "cartoon, illustration, digital art, unrealistic colors, lowres, cropped, worst quality, lowquality". The input text prompts are processed by the CLIP text encoder to obtain caption tokens.

[0058] S4:Stable Diffusion:

[0059] This method is based on the Stable Diffusion architecture. It adds new modules to adapt to the task of generating images of litchi pests and diseases, freezes the weights of the original model, and adjusts the new modules to achieve the best multi-category litchi pest and disease image generation effect. Because Stable Diffusion has been pre-trained with a large amount of image and text data, it has accumulated the deep knowledge required to generate realistic images. Retaining the original knowledge in the model is particularly important when expanding new functions. Doing so can save training costs and make full use of the existing rich knowledge. In order to apply Stable Diffusion to the generation of orchard scene images containing multiple categories of pests and diseases, additional conditions such as bounding boxes and category information need to be added. Although Stable Diffusion is good at generating images based on given prompts, it cannot use additional positioning inputs to guide the generation process and specify the specific location and category of objects in the image. In order to overcome the limitation that Stable Diffusion cannot use additional positioning inputs in the image generation process, this method adopts the method of GLIGEN as follows Figure 2 As shown in the figure, StableDiffusion is extended by adding a gated self-attention mechanism. After the extension, the localization ability of the model is significantly enhanced, enabling it to accept and utilize localization inputs such as categories and bounding boxes to guide the image generation process and generate images conditioned on the grounding input.

[0060] The original Transformer module consists of a self-attention layer and a cross-attention layer. The self-attention layer processes image features, while the cross-attention layer integrates text features with image features.

[0061] Using v to represent Visual Tokens, these two layers can be written as:

[0062] v=v+SelfAttn(v) (2)

[0063] v=v+CrossAttn(v,h c ) (3)

[0064] Freeze these two attention layers and add a new gated self-attention layer that receives GroundingTokens(g):

[0065] v=v+β·tanh(γ)·TS(SelfAttb(g)) (4)

[0066] where TS() is a token selection operation that considers only v, and γ is a learnable scalar initialized to 0. β is varied at inference time for a predetermined sampling to improve quality and controllability, while it is set to 1 during training.

[0067] S5:Refiner:

[0068] Refiner is an independent Latent Diffusion Model that performs noise-denoising on the samples generated by stable diffusion to improve the visual quality of the samples. Refiner is a graph-to-graph module with strong portability and can be integrated into other generative models as a cascade component. In order to obtain higher image quality, the present invention integrates refiner after stable diffusion, inputs the same text prompt, iteratively refines the latent vector generated by stable diffusion, and decodes it back into the image.

[0069] S6: Training and loss:

[0070] The present invention will use the litchi pest and disease detection dataset to train the model so that the Grounding input can be injected into StableDiffusion with other original components unchanged, improving the performance of the generative model in generating 6 types of litchi pest and disease images. The model is fine-tuned on the basis of Stable Diffusion v1.4, freezing the pre-trained weights of Stable Diffusion v1.4 and only training the parameters of the newly added gated self-attention layer (Gated Self-Attention) and GroundingTokenizer modules. The reconstruction loss is used to compare the distance between the real noise and the noise predicted by the model to measure the performance of the model, and the AdamW optimizer is used to calculate the gradient and update the model parameters.

[0071] Specifically, for an image sampled from the training dataset and the corresponding grounding input y, the image is encoded into a latent representation zt, and the model predicts the noise f{θ,θ′}(zt,t,y) at time step t, drawn from a standard normal distribution The real noise ∈ sampled in , calculate the square of the L2 distance between the real noise and the predicted noise, and finally take the expectation of all sampled samples. The specific calculation is as follows, where θ represents the original parameters and θ′ represents all new parameters:

[0072]

[0073] Test example:

[0074] In order to analyze the results of image generation, this method uses Fréchet Inception Distance (FID) and CLIP (Contrastive Language-Image Pretraining) score to evaluate the generated images. FID and CLIP score are often used to evaluate the performance of generative models. FID uses the pre-trained Inception V3 model to extract image features and measures the quality of generated images by comparing the distribution differences between generated images and real images in the feature space. The lower the FID value, the closer the distribution of the generated image is to the real image, that is, the higher the quality of the generated image. The calculation formula of FID is:

[0075] FID=||μ r -μ g || 2 +Tr(∑ r +∑ g -2(∑ r ∑ g ) 1 / 2 ) (6)

[0076] Among them, μ r , μ g are the feature means of the real image and the generated image, ∑ r ,∑ g are the covariance matrices of the real image and the generated image respectively, and Tr represents the trace of the matrix.

[0077] Through CLIP, a multimodal learning method proposed by OpenAI, the original image and the generated image are mapped to the same embedding space, and their cosine similarity in the embedding space is calculated to obtain the CLIP score between the generated image and the original image. The higher the value, the higher the similarity. CLIP score focuses on the correlation between the generated image and the original image, while FID focuses more on the visual quality and authenticity of the generated image.

[0078] The FID and CLIP score results of the litchi pest and disease images generated by this method are shown in the following table.

[0079] Generate model evaluation results table

[0080]

[0081]

[0082] From the above, it can be known that in the present invention, Open-Vocabulary SAM is used to segment pest targets, and CLIP is used to extract image features, which can generate high-quality images containing detailed features of pests and diseases, such as lesions, insect bodies, etc., to improve the authenticity and credibility of the images. By designing text prompts containing target categories, quantities, and scene environments, this method can generate a variety of litchi pest images, covering different scenes, different quantities, and different combinations of pest types to meet the needs of different research scenarios. The introduction of Grounding information, including category names, bounding box coordinates, etc., can guide the StableDiffusion model to accurately represent the location and category of pests and diseases during the image generation process, thereby improving the accuracy and controllability of the image.

[0083] In the present invention, a large number of high-quality litchi pest and disease images can be generated, which effectively expands the scale of litchi pest and disease data sets, solves the problems of data acquisition difficulties and data scarcity in traditional methods, and provides data support for the training and evaluation of litchi pest and disease detection models. The generated high-quality litchi pest and disease images can be used as an auxiliary tool for litchi pest and disease identification and diagnosis, helping agricultural technicians to quickly and accurately identify the types and quantities of pests and diseases, and take effective prevention and control measures to improve agricultural production efficiency. By improving the efficiency of litchi pest and disease detection and prevention, the method can effectively reduce the losses caused by litchi pests and diseases, ensure litchi yield and fruit quality, and promote the healthy development of the litchi industry.

[0084] In the present invention, the Grounding technology is applied to the image generation model, which effectively improves the quality and accuracy of the model in content generation, and provides new ideas and methods for the development of image generation technology. The method has been successfully applied to litchi pest and disease image generation, proving its scalability. In the future, it can be expanded to other fields, such as medical image generation, remote sensing image generation, etc., to promote the application of image generation technology in more fields. The method combines images, texts and Grounding information to achieve multimodal learning, providing new cases and experiences for the development of multimodal learning technology.

[0085] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0086] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A litchi pest and disease image generation method based on Stable Diffusion and Grounding, characterized in that: The method comprises the following steps: S1: Image acquisition: Collect images containing six types of litchi pests and diseases in the litchi orchard, annotate them, record the categories and bounding box coordinates, and divide the collected image dataset into training set and test set for model training and evaluation; S2: Use CLIP Image Encoder to extract image features from litchi pest and disease images; S3: Use the Open-Vocabulary SAM model to segment the pest target and background, and extract image features from the segmented image; S4: Use CLIP Text Encoder to extract text features from category names; convert bounding box coordinates into Fourier embedding to represent the location information of the object in the image; use Grounding Tokenizer to fuse image features, text features and location information to generate Grounding Tokens; S5: Design text prompts containing target categories, quantities, and scene environments to guide Stable Diffusion to generate images; use CLIP Text Encoder to encode text prompts into Caption Tokens; S6: Based on the Stable Diffusion architecture, a gated self-attention mechanism is added to accept and utilize GroundingTokens and Caption Tokens to guide the image generation process; Freeze the original Stable Diffusion model weights and only train the parameters of the newly added gated self-attention layer and GroundingTokenizer module; S7: Use Refiner to refine the image generated by Stable Diffusion to improve image quality and details; S8: Use the litchi pest and disease detection dataset to train the model; use reconstruction loss to measure model performance, and use AdamW optimizer to update model parameters; S9: Use FID and CLIP score to evaluate the quality of the generated litchi pest and disease images; the lower the FID value, the closer the distribution of the generated image is to the real image, that is, the higher the quality of the generated image; the higher the CLIP score value, the higher the correlation between the generated image and the original image.

2. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In step S2, the litchi disease and pest images are from an orchard scene, and a single image often contains multiple diseases and pests. Compared with the litchi fruits and trees in the background, the main litchi diseases and pests are target types not included in the pre-training data set.

3. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In the step S3, after obtaining the litchi pest and disease image mask using Open-Vocabulary SAM, the target pest and disease is accurately extracted from the original image by Mask Cropping, and then input into CLIP Image Encoder to obtain image embedding; Open-Vocabulary SAM is based on SAM and CLIP, SAM2CLIP and CLIP2SAM are two knowledge transfer modules, the former transfers the knowledge of SAM to CLIP through distillation and a learnable Transformer adapter, and the latter transfers the knowledge of CLIP to SAM, thereby improving the recognition ability of the model; the input image is encoded into image embedding by CLIP Image Encoder, the input coordinate frame is encoded by SAM prompt encoder, and then the SAM2CLIP module uses adaptive and distillation methods to bridge the gap in the feature representation of SAM and CLIP learning, and the CLIP2SAM module uses the knowledge of CLIP to enhance the recognition ability of SAM Mask Decoder, and SAM Mask Decoder decodes to generate segmentation masks and label annotations.

4. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In step S4, CLIP Text Encoder is used to obtain text features; the bounding box is represented by the coordinates of the upper left corner and the lower right corner, and is transformed into Fourier embedding using Fourier Embedder to represent the location information of the object in the image; it is then fused by Grounding tokenizer to generate a new feature representation GroundingTokens; Grounding tokenizer is a multi-layer perceptron that connects the input in the feature dimension; h e represents Grounding Tokens, then the above process can be expressed as: h e =Grounding_tokenizer(f image (x),f text (c),Fourier_embedder(l)) (1)。 5. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In step S5, the text prompt consists of the pest category, quantity and scene environment, and the specific format is: "A photo of {count}{class A}, {count}{class B}, and {count}{class C}, {scene environments}"; wherein, {scene environments} is used to describe the overall style and aesthetic quality of the image, and is set to: "photorealistic, high resolution, natural lighting, outdoor, litchi orchard"; at the same time, in order to remind the model of elements or features that should be avoided when generating images, a negative prompt is designed to help more accurately control the generated image content; the negative prompt is set to: "cartoon, illustration, digital art, unrealistic colors, lowres, cropped, worst quality, low quality"; the input text prompt is processed by the CLIP text encoder to obtain caption tokens.

6. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In step S6, the GLIGEN method is used to expand Stable Diffusion; after the expansion, the positioning ability of the model is significantly enhanced, so that it can accept and use positioning inputs such as categories and bounding boxes to guide the image generation process, and generate images based on the grounding input; The original Transformer module consists of a self-attention layer and a cross-attention layer; the image features are processed in the self-attention layer; and the text features are integrated with the image features in the cross-attention layer; Using v to represent Visual Tokens, these two layers can be written as: v=v+SelfAttn(v) (2) v=v+CrossAttn(v,h c ) (3) Freeze these two attention layers and add a new gated self-attention layer that receives Grounding Tokens (g): v=v+β·tanh(γ)·TS(SelfAttn(g)) (4) where TS() is a token selection operation that considers only v, and γ is a learnable scalar initialized to 0; β varies at inference time for a predetermined sampling to improve quality and controllability, while it is set to 1 during training.

7. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In the step S7, Refiner is an independent Latent Diffusion Model, which performs noise-denoising processing on the samples generated by stable diffusion to improve the visual quality of the samples; Refiner is a graph-to-graph module with strong portability and can be integrated into other generative models as a cascade component; in order to obtain higher image quality, the present invention integrates refiner after stable diffusion, inputs the same text prompt, iteratively refines the latent vector generated by stable diffusion, and decodes it back into the image.

8. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In the step S8, the model is trained using the litchi pest and disease detection dataset so as to inject the Grounding input into StableDiffusion with other original components unchanged, thereby improving the performance of the generative model in generating images of six types of litchi pests and diseases; the model is fine-tuned on the basis of Stable Diffusion v1.4, the pre-trained weights of StableDiffusion v1.4 are frozen, and only the parameters of the newly added gated self-attention layer and Grounding Tokenizer module are trained; and the reconstruction loss is used to compare the distance between the real noise and the noise predicted by the model to measure the performance of the model, and the AdamW optimizer is used to calculate the gradient and update the model parameters; Specifically, for an image sampled from the training dataset and the corresponding grounding input y, the image is encoded into a latent representation zt, and the model predicts the noise f{θ,θ′}(zt,t,y) at time step t, drawn from a standard normal distribution The real noise ∈ sampled in , calculate the square of the L2 distance between the real noise and the predicted noise, and finally take the expectation of all sampled samples. The specific calculation is as follows, where θ represents the original parameters and θ′ represents all new parameters:

9. The litchi pest and disease image generation method based on Stable Diffusion and Grounding as claimed in claim 1, characterized in that: In step S9, the generated image is evaluated using Fréchet Inception Distance and CLIP score; FID and CLIP score are often used to evaluate the performance of the generated model; FID uses the pre-trained InceptionV3 model to extract image features, and measures the quality of the generated image by comparing the distribution difference between the generated image and the real image in the feature space; the lower the FID value, the closer the distribution of the generated image is to the real image, that is, the higher the quality of the generated image; the calculation formula of FID is: FID=||μ r -m g || 2 +Tr(Σ r +∑ g -2(∑ r ∑ g ) 1 / 2 ) (6) Among them, μ r , μ g are the feature means of the real image and the generated image, Σ r ,Σ g are the covariance matrices of the real image and the generated image, respectively, and Tr represents the trace of the matrix; Through CLIP, a multimodal learning method proposed by OpenAI, the original image and the generated image are mapped to the same embedding space, and their cosine similarity in the embedding space is calculated to obtain the CLIP score between the generated image and the original image. The higher the value, the higher the similarity. The CLIP score focuses on the correlation between the generated image and the original image, while FID focuses more on the visual quality and authenticity of the generated image.