Data enhancement method and system based on diffusion model
By using a data augmentation method based on a diffusion model, enhanced images that conform to the biological characteristics of shrimp are generated, solving the problem of insufficient data and improving the accuracy and efficiency of shrimp classification and identification.
Patent Information
- Application Number
- CN202511082897.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-28
AI Technical Summary
Existing shrimp classification methods suffer from insufficient data samples, lack of diversity, and limitations of traditional data augmentation methods, resulting in low classification accuracy.
A data augmentation method based on a diffusion model is adopted. By acquiring the original dataset of shrimp images, a set of cue words is generated. The forward noise addition and reverse noise reduction mechanisms of the diffusion model are used to generate diverse augmented images in combination with the cue words. New augmented images are generated through a latent variable sequence and cue word guidance mechanism. Augmented images containing random variations are generated, and augmented images that conform to the biological characteristics of shrimp are generated through a latent variable sequence and cue word guidance mechanism.
The method generates diverse shrimp images, enriches the training dataset, improves the accuracy and efficiency of shrimp classification and recognition, and avoids the problem of generated images not conforming to biological characteristics that exists in existing technologies.
Smart Images

Figure CN121032841A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, and specifically to a data augmentation method and system based on a diffusion model. Background Technology
[0002] In the aquaculture sector, shrimp farming is an important industry branch, with species such as the Chinese white shrimp, Litopenaeus vannamei, and Litopenaeus molluscini being the main farmed species. Accurate shrimp classification is crucial for various stages, including shrimp larvae sorting, farming management, disease prevention and control, and market sales. For example, during shrimp larvae procurement, mixing different species can negatively impact farming quality and economic benefits; during the farming process, accurate identification of shrimp species facilitates targeted farming management and disease control.
[0003] The shrimp classification methods in related technologies have the following drawbacks: Manual sorting is inefficient and prone to errors: Currently, shrimp larvae sorting mainly relies on manual labor, requiring workers to be proficient in the characteristics of different shrimp species. However, this method is not only inefficient and labor-intensive, but also prone to sorting errors. This is especially true for species that closely resemble Litopenaeus vannamei, Litopenaeus chinensis, and Litopenaeus moguls, making them even more difficult to distinguish during the larval stage. This significantly impacts shrimp quality and sales, resulting in economic losses for aquaculture farms.
[0004] The automatic sorting method has shortcomings: Existing automatic shrimp sorting methods usually classify shrimp based on data such as weight, size, and outline. However, weight data requires weighing the shrimp after they are out of the water, which is inconvenient to operate in actual aquaculture environments. Furthermore, relying solely on these physical characteristics cannot effectively distinguish shrimp species with similar appearances, resulting in low classification accuracy.
[0005] With the development of computer vision and artificial intelligence technologies, image recognition technology is gradually being applied to shrimp classification. However, this technology faces the serious challenge of insufficient data samples during its application.
[0006] Sample collection is difficult: In actual aquaculture scenarios, collecting a large number of high-quality and diverse shrimp image samples requires a lot of time and effort, and is limited by various factors such as environment and shrimp growth stage, making it difficult to obtain sufficient training data.
[0007] Insufficient data diversity: Existing shrimp image datasets often suffer from problems such as limited sample size, single perspective, inconsistent lighting conditions, and limited shrimp pose variation. This results in poor generalization ability of classification models trained on these datasets, making it difficult to accurately identify shrimp species under various conditions in practical applications.
[0008] Data augmentation methods and limitations in related technologies: Traditional data augmentation methods include simple image transformation operations such as rotation, flipping, scaling, cropping, and color transformation. While these methods can increase the number of data samples to some extent, they only perform simple geometric or color transformations on the original image and cannot generate samples with new morphological features and details, thus failing to effectively improve the model's ability to recognize typical characteristics of shrimp species.
[0009] Generative adversarial networks (GANs) and other generative models have also been applied to data augmentation. However, these methods have some obvious limitations when generating shrimp images. For example, the training process of GAN models is unstable and prone to pattern collapse, resulting in inconsistent sample quality. Furthermore, it is difficult to guarantee the biological validity of the generated images, which can easily produce artifacts and images that do not conform to the morphological characteristics of shrimp, thus affecting the training effect of classification models. Summary of the Invention
[0010] The purpose of this invention is to provide a data augmentation method and system based on a diffusion model, which aims to generate diverse shrimp images and improve the accuracy and efficiency of shrimp classification and recognition.
[0011] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, embodiments of the present invention provide a data augmentation method based on a diffusion model, the method comprising the following steps: S100, Obtain the original dataset containing shrimp images, and generate a set of prompt words corresponding to the shrimp images; S200, based on the forward noise addition mechanism of the diffusion model, the shrimp image is subjected to progressive random noise superposition processing to obtain the latent spatial representation under different noise intensities, forming a noisy latent variable sequence; the latent variable sequence includes the noisy image at each time step; S300 uses a sequence of latent variables and a set of cue words as dual inputs. It embeds a cue word guidance mechanism in the reverse denoising process of the diffusion model and outputs an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp. S400 integrates the generated enhanced images into the original dataset to obtain the enhanced dataset.
[0012] Optionally, in S100, obtaining the original dataset containing shrimp images and generating a set of prompt words corresponding to the shrimp images includes: S110, acquire shrimp images, identify typical features of shrimp in the shrimp images, and form a feature vector from the typical features of shrimp; the typical features include the shrimp's body outline, the color and texture of the shrimp's body, and the shape of the shrimp's antennae and eyes; S120, combining the typical features and the pre-built shrimp feature knowledge base, generate prompt words corresponding to the shrimp in the shrimp image; the shrimp feature knowledge base contains prompt words for different varieties of shrimp, and the prompt words are used to describe the typical features of the corresponding shrimp. S130, Based on the typical characteristics of shrimp and the corresponding prompt words, construct a feature-prompt word mapping model, and map the feature vector of shrimp to a vector representation composed of prompt words through the mapping model; S140, determine the feature score of each prompt word in the vector representation. When the feature score of a prompt word exceeds the feature score threshold, the prompt word is used as the prompt word corresponding to the shrimp image and added to the prompt word set. Optionally, in S130, the construction of a feature-cue word mapping model based on typical features of shrimp and corresponding cue words includes: S131, Form a feature vector from the typical characteristics of each shrimp species, and determine the prompt words corresponding to the feature vector; S132, input the feature vector and the corresponding prompt word into the Transformer network for training, and obtain the trained Transformer network as the feature-prompt word mapping model.
[0013] Optionally, in S200, the step of performing progressive random noise superposition processing on shrimp images based on the forward noise addition mechanism of the diffusion model, recording the latent spatial representation under different noise intensities, and forming a noisy latent variable sequence includes: S210, The prompt words in the prompt word set are input into the trained and optimized diffusion model. The diffusion model reconstructs the initial Gaussian noise required to generate the shrimp image based on the input prompt words through a reverse inference algorithm. S220, Based on the initial Gaussian noise field, the shrimp image is subjected to progressive random noise superposition processing to obtain noise images at each time step, and the noise images at each time step are sequentially arranged to form a latent variable sequence.
[0014] Optionally, the initial Gaussian noise required for reconstructing and generating the shrimp image using the reverse inference algorithm includes: The shrimp image is back-diffused using the following formula, by iterating the shrimp image from the maximum time step to the minimum time step, to obtain the initial Gaussian noise of the shrimp image:
[0015] in, , They are at the time step , Noisy images, For time steps, For the maximum time step, , These represent the time steps respectively. , The variance reduction factor These are the parameters of the diffusion model; Is the model at time step Noise estimation.
[0016] Optionally, in S300, the step of embedding a cue word guidance mechanism in the reverse denoising process of the diffusion model, using the latent variable sequence and cue word set as dual inputs, and outputting an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp, includes: S310, which uses text embedding technology to transform the set of prompt words into text embedding vectors; S320, starting from the initial Gaussian noise obtained by inversion, the shrimp image is gradually denoised and updated. In the iterative update at each time step, the text embedding vector is combined to guide the generation process, and the output is an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp.
[0017] Optionally, the iterative update formula for the shrimp image is:
[0018] in, , They are at the time step , Noisy images, , These represent the time steps. , The variance reduction factor These are the parameters of the diffusion model; In time step Noise estimation incorporating cue word information It is a noise vector sampled from a standard Gaussian distribution; The shrimp image is denoised using the following formula to obtain an enhanced image: ; in, It is noise from random sampling. This represents the variance reduction factor. Indicates the percentage of noise added. ,0< <1, 0< <1; Indicated based on time step Noisy images Given time step t and condition c, predict the amount to be added to the noisy image. The estimated value of noise in the data; In a certain intermediate generation step have: ; in, This represents a denoising operation based on random noise. Indicates based on guidance information The guided operation, It is in the steps A coefficient that controls the guiding intensity; The enhanced image in step s represents the image from... Gradually denoise and restore the enhanced image.
[0019] Secondly, embodiments of the present invention provide a data augmentation system based on a diffusion model, the system comprising: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any of the preceding statements.
[0020] The beneficial effects of this invention are as follows: By employing a data augmentation method based on a diffusion model, this invention effectively solves the problems of insufficient data samples and limitations of traditional data augmentation methods in shrimp classification. Specifically, this method acquires the original dataset and generates a set of cue words. Utilizing the forward noise addition and reverse denoising mechanisms of the diffusion model, combined with cue word guidance, it generates augmented images that contain both random variations and conform to the biological characteristics of shrimp, greatly enriching the diversity of shrimp images. Compared to traditional data augmentation methods that can only perform simple geometric or color transformations, the samples generated by this method possess new morphological features and details, helping to improve the model's ability to recognize typical characteristics of shrimp species. Compared to generative models such as Generative Adversarial Networks (GANs), this method does not suffer from problems such as unstable training processes, pattern collapse, and difficulty in ensuring the biological rationality of generated images. By integrating the generated augmented images into the original dataset, the resulting augmented dataset can provide richer, higher-quality training data for shrimp classification models, thereby potentially significantly improving the accuracy and efficiency of shrimp classification and recognition. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of the data augmentation method based on the diffusion model in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating data augmentation of shrimp images based on a diffusion model in an embodiment of the present invention; Figure 3 This is a schematic diagram of the data augmentation system based on the diffusion model in an embodiment of the present invention. Detailed Implementation
[0023] The following will clearly and completely describe the concept, specific structure, and technical effects of the present invention in conjunction with embodiments and accompanying drawings, so as to fully understand the purpose, solution, and effects of the present invention. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present invention can be combined with each other.
[0024] To overcome the technical bottlenecks in shrimp image data augmentation, large-scale visual model image editing techniques based on diffusion theory have become an important research direction in this field in recent years and have been widely applied in data augmentation practice. Based on differences in editing paradigms, current mainstream diffusion model image editing techniques can be mainly divided into the following two categories: Dragging editing, inspired by the DragGAN method, involves users defining multiple sets of coordinate pairs. The system uses an optimization algorithm to migrate content from the vicinity of the starting point to the target location, while ensuring high fidelity and semantic rationality in the editing result. The main technical advantages of this method lie in its intuitive operation and fine-grained editing capabilities; however, its technical limitations are significant, making it difficult to achieve coarse-grained editing operations such as background replacement and object replacement. Text-driven editing requires users to provide the original image, corresponding text descriptions, and target descriptions. The system must ensure that the editing result meets both the target semantic requirements and maintains consistency with the original image. This method can achieve precise modification of specific areas while keeping other areas unchanged. Its significant technical advantage lies in the flexibility of natural language description, which makes it easy for users to freely express their editing intentions; however, it has obvious technical shortcomings, such as the difficulty in semantic representation of fine-grained actions like raising a hand, and the need to invest more computing resources in maintaining the overall consistency of the image.
[0025] Therefore, in order to overcome the problems of insufficient data samples and low classification accuracy in existing shrimp classification technologies, there is an urgent need for an efficient and accurate data augmentation method that can generate a large number of high-quality, diverse shrimp images that conform to biological laws, so as to expand the training set and improve the performance of shrimp classification and recognition models.
[0026] See Figure 1 This invention provides a data augmentation method based on a diffusion model, the method comprising the following steps: S100, Obtain the original dataset containing shrimp images, and generate a set of prompt words corresponding to the shrimp images; S200, based on the forward noise addition mechanism of the diffusion model, the shrimp image is subjected to progressive random noise superposition processing to obtain the latent spatial representation under different noise intensities, forming a noisy latent variable sequence; the latent variable sequence includes the noisy image at each time step; S300 uses a sequence of latent variables and a set of cue words as dual inputs. It embeds a cue word guidance mechanism in the reverse denoising process of the diffusion model and outputs an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp. S400 integrates the generated enhanced images into the original dataset to obtain the enhanced dataset.
[0027] This invention addresses the limitations of diffusion models in precise inversion and the inadequacies of existing data augmentation schemes by proposing a guided random diffusion inversion method. The core mechanism of this method lies in fully utilizing the inherent characteristics of diffusion models. While maintaining randomness to ensure data diversity, it introduces guiding information to achieve precise directional control of the generation process, thereby effectively overcoming the technical bottleneck that "precise inversion can only generate deterministic images."
[0028] The specific implementation process includes three key stages:
[0029] The first stage involves the preprocessing of data and guidance information. First, a dataset of original shrimp images is constructed, and corresponding prompts (covering shrimp morphology, texture, color, and other feature descriptions) are generated through feature extraction techniques. Second, based on the forward noise addition mechanism of the diffusion model, progressive random noise superposition is applied to the original images, while recording the latent spatial representations under different noise intensities, ultimately forming a noisy sequence of latent variables.
[0030] The second stage implements a guided stochastic inversion process. Using the latent variable sequence and cue words as dual inputs, a cue word guidance mechanism is embedded in the inverse denoising process of the diffusion model. This stage preserves the characteristics of randomly sampled noise in each denoising operation, enabling the model to explore diverse generation methods in the noisy space under cue word constraints (such as forcibly preserving key morphological features of shrimp). Through iterative calculation, the final output is a newly generated image that contains both random variations and conforms to the biological characteristics of shrimp.
[0031] The third stage involves constructing a closed-loop data augmentation system. The generated noisy images are integrated into the original dataset for subsequent optimization training of the shrimp recognition model. Based on performance feedback from the recognition task, the prompt word construction strategy (refining feature description accuracy) and noise sampling scheme (optimizing the balance between randomness and controllability) are dynamically adjusted, continuously expanding the high-quality augmented dataset to systematically improve the model's generalization performance.
[0032] This invention innovatively integrates a "cue word guidance mechanism" with a "random noise exploration strategy" to achieve synergistic optimization of randomness and directionality during image generation. This technology not only retains the inherent data augmentation advantages of diffusion models but also effectively controls the generation direction through guidance information, fundamentally solving the technical limitation of traditional precise inversion methods that "can only generate deterministic images." This provides a more adaptable and practical data augmentation solution for computer vision tasks such as shrimp recognition.
[0033] This invention provides a data augmentation method based on a diffusion model. By employing image inversion data augmentation, it generates high-quality images of *Litopenaeus vannamei*, *Litopenaeus vannamei*, and *Litopenaeus molluscini*, expanding the training set. Specifically, this method should be able to: generate shrimp images with rich morphological features and details, accurately reflecting the typical characteristics of different shrimp species such as color, texture, body shape, and markings; generate diverse shrimp images, including different lighting conditions, background environments, viewing angles, and poses, to improve the generalization ability of the classification model; ensure the biological rationality of the generated images, avoiding artifacts and images that do not conform to the morphological characteristics of shrimp; and be compatible with existing shrimp classification models, allowing the generated samples to be directly used for model training, improving the accuracy and efficiency of classification and recognition.
[0034] In some embodiments, in S100, obtaining the original dataset containing shrimp images and generating a set of prompt words corresponding to the shrimp images includes: S110, acquire shrimp images, identify typical features of shrimp in the shrimp images, and form a feature vector from the typical features of shrimp; the typical features include the shrimp's body outline, the color and texture of the shrimp's body, and the shape of the shrimp's antennae and eyes; S120, combining the typical features and the pre-built shrimp feature knowledge base, generate prompt words corresponding to the shrimp in the shrimp image; the shrimp feature knowledge base contains prompt words for different varieties of shrimp, and the prompt words are used to describe the typical features of the corresponding shrimp. S130, Based on the typical characteristics of shrimp and the corresponding prompt words, construct a feature-prompt word mapping model, and map the feature vector of shrimp to a vector representation composed of prompt words through the mapping model; S140, determine the feature score of each prompt word in the vector representation. When the feature score of a prompt word exceeds the feature score threshold, the prompt word is used as the prompt word corresponding to the shrimp image and added to the prompt word set. The specific screening mechanism is as follows: setting a feature score threshold. When prompt words Corresponding feature scores If the feature score exceeds the threshold, the prompt word is included in the final set; otherwise, it is removed. Its mathematical expression can be simplified to: .
[0035] In this embodiment, by recognizing and analyzing typical features of shrimp images, valuable information can be accurately extracted from the original dataset, providing strong support for subsequent data augmentation operations. For example, different species of shrimp have significant differences in body shape and outline; the Chinese white shrimp is relatively slender, while the Litopenaeus vannamei is relatively short and stout. By recognizing these features and converting them into feature vectors, a foundation can be laid for generating accurate prompts.
[0036] In the process of generating prompt words by combining typical features with a shrimp feature knowledge base, the rich information contained in the knowledge base plays a crucial role. It not only covers the basic feature descriptions of common shrimp species but may also include feature information of shrimp in special cases, making the generated prompt words more comprehensive and accurate. For example, the color and texture of some shrimp may change at different growth stages; the knowledge base can provide this detailed information to help generate prompt words that highly match images of specific shrimp species.
[0037] A feature-cue word mapping model was constructed to further optimize the cue word generation process. This model can transform complex feature vectors into easily understood and processed cue word vector representations, facilitating subsequent cue word selection and integration. By determining a feature score threshold, the most representative cue words are selected and included in the cue word set, ensuring that the cue words in the set are closely related to shrimp images and have high value.
[0038] In practical applications, this prompt word generation method, which combines feature recognition and a knowledge base, significantly improves the quality and relevance of prompt words during data augmentation. It makes subsequent image generation operations based on the diffusion model more accurate and effective, resulting in augmented images that better reflect shrimp biological characteristics. This provides high-quality data support for training shrimp classification and recognition models, further enhancing model performance and generalization ability. Simultaneously, as shrimp research deepens and new data accumulates, the shrimp feature knowledge base can be continuously updated and improved, further optimizing the prompt word generation process and ensuring the overall data augmentation system remains consistently effective and adaptable.
[0039] In some embodiments, S130, the construction of a feature-cue word mapping model based on typical characteristics of shrimp and corresponding cue words includes: S131, Form a feature vector from the typical characteristics of each shrimp species, and determine the prompt words corresponding to the feature vector; S132, input the feature vector and the corresponding prompt word into the Transformer network for training, and obtain the trained Transformer network as the feature-prompt word mapping model.
[0040] In this embodiment, by inputting feature vectors and prompt words into a Transformer network for training, the powerful self-attention mechanism of the Transformer network can effectively capture the complex semantic relationships between feature vectors and prompt words. The Transformer network can automatically learn the interactions between different features and how they correspond to the corresponding prompt word descriptions.
[0041] During training, the network continuously adjusts its parameters to minimize the error between the predicted prompts and the actual corresponding prompts. This training method enables the Transformer network to gradually master the accurate mapping relationship from shrimp feature vectors to prompts, thereby constructing a high-performance feature-prompt mapping model.
[0042] Compared to traditional mapping methods, the Transformer network-based mapping model exhibits better generalization ability. It can not only handle feature combinations that have appeared in the training data, but also provide reasonable prompt word mappings for new and unseen feature combinations based on learned semantic rules. For example, when encountering a shrimp species that combines features from two known species, this mapping model can still generate relatively accurate prompt words by understanding the existing feature-prompt word relationships.
[0043] Furthermore, the parallel computing capabilities of the Transformer network make it highly efficient when processing large amounts of data. In practical applications, when faced with massive raw shrimp datasets, this efficiency can significantly shorten the training time of the mapping model, accelerate the data augmentation process, and enable the entire system to quickly respond to and process new image data.
[0044] The feature-cue word mapping model constructed in this way provides more accurate cue word input for the subsequent data augmentation process based on the diffusion model. This helps to generate enhanced images that are more consistent with the real characteristics of shrimp and have richer diversity, further improving the performance and accuracy of shrimp classification and recognition models, and thus playing an important role in shrimp research and related industrial applications.
[0045] In some embodiments, S200, the progressive random noise superposition processing of shrimp images based on the forward noise addition mechanism of the diffusion model, recording the latent spatial representation under different noise intensities, and forming a noisy latent variable sequence includes: S210, The prompt words in the prompt word set are input into the trained and optimized diffusion model. The diffusion model reconstructs the initial Gaussian noise required to generate the shrimp image based on the input prompt words through a reverse inference algorithm. S220, Based on the initial Gaussian noise field, the shrimp image is subjected to progressive random noise superposition processing to obtain noise images at each time step, and the noise images at each time step are sequentially arranged to form a latent variable sequence.
[0046] In this embodiment, by inputting prompt words into a trained and optimized diffusion model, an inverse inference algorithm is used to generate initial Gaussian noise, laying the foundation for subsequent noise superposition processing. This method of generating initial noise based on prompt words ensures that the addition of noise is not completely random, but closely related to the features of the shrimp image. Because the prompt words contain key feature information such as the shrimp's size, color, and texture, the diffusion model can generate more targeted initial Gaussian noise based on this information.
[0047] In the progressive random noise superposition process based on an initial Gaussian noise field, the noise addition at each time step is carefully designed. As the time step progresses, the noise intensity gradually increases, causing the image to change gradually in the process. For example, in the initial stage, the noise intensity is low, and the image may only show some slight blurring or texture changes; however, as the time step goes on, the noise intensity increases, and the details of the image are gradually masked by the noise, eventually forming a sequence of noisy images with different noise intensities.
[0048] This progressive noise superposition method is significant. On one hand, it simulates various noise interference scenarios that images may encounter in the real world, making the noisy images generated by the diffusion model more realistic and diverse. On the other hand, by recording the latent spatial representations under different noise intensities, it provides rich information for the subsequent reverse denoising process. These latent spatial representations contain information on the feature changes of the image during noise interference, allowing the model to better reconstruct enhanced images that conform to the biological characteristics of shrimp during reverse denoising.
[0049] The resulting noisy latent variable sequence, i.e., the noisy images at each time step arranged sequentially, serves as a crucial input for the subsequent guided stochastic inversion process. It provides the data foundation for embedding a cue word guidance mechanism during the diffusion model's reverse denoising process. This allows the model to gradually remove noise based on this noisy sequence, combined with cue word constraints, generating new images that contain both random variations and conform to shrimp biological characteristics. This further enhances the effectiveness and practicality of the diffusion model-based data augmentation method, providing higher-quality training data for shrimp classification and recognition models and promoting improved model performance.
[0050] Furthermore, this progressive random noise superposition method offers flexibility and adjustability. The generation process of the noisy image can be controlled by adjusting parameters such as the noise intensity growth rate and the noise distribution type. Depending on different application scenarios and needs, image sequences with varying noise characteristics can be generated to meet the enhancement requirements of shrimp classification and recognition models at different training stages or with different data demands. For example, in the early stages of model training, image sequences with relatively gradual changes in noise intensity may be needed to help the model gradually adapt to noise interference and learn the basic features of the image; while in the later stages of training, to improve the model's robustness, the amplitude of noise intensity changes can be appropriately increased to generate more challenging noisy images, prompting the model to learn more complex image features.
[0051] It's important to note that parameter optimization is a crucial step in this gradual random noise overlay process. Extensive experimentation and data analysis, combined with consideration of the characteristics of shrimp images and performance feedback from the classification model, are necessary to determine the optimal combination of noise addition parameters. For example, by comparing the classification accuracy and recall of enhanced images generated at different noise intensity growth rates after model training, the noise addition scheme that maximizes model performance can be identified. Furthermore, as shrimp image data accumulates and the understanding of diffusion models deepens, these parameters can be dynamically adjusted and optimized to further improve the data augmentation effect and model performance.
[0052] This embodiment performs progressive random noise superposition processing based on the forward noise addition mechanism of the diffusion model, forming a noisy latent variable sequence. It plays a crucial role in the data augmentation method based on the diffusion model, laying a solid foundation for the smooth progress of the entire data augmentation process and the final generation of high-quality shrimp augmented images.
[0053] In some embodiments, the initial Gaussian noise required to reconstruct the shrimp image using a reverse deduction algorithm includes: The shrimp image is back-diffused using the following formula, by iterating the shrimp image from the maximum time step to the minimum time step, to obtain the initial Gaussian noise of the shrimp image:
[0054] in, , They are at the time step , Noisy images, For time steps, For the maximum time step, , These represent the time steps. , The variance reduction factor These are the parameters of the diffusion model; Is the model at time step Noise estimation.
[0055] In this embodiment, the backdiffusion operation using the above formula enables the gradual derivation of initial Gaussian noise from a given shrimp image. This back-diffusion algorithm exhibits high accuracy and logical consistency. During the iterative process from the maximum time step to the minimum time step, each step is calculated based on the noise image at the current time step, the variance reduction factor, and the model's estimate of the noise.
[0056] From a mathematical perspective, the parameters in the formula interact with each other, resulting in a variance reduction factor. , The degree of change of the noisy image between different time steps is determined by the model's estimation of the noise. This provides direction for each iteration. Through this precise calculation, the generated initial Gaussian noise can be closely correlated with the features of the shrimp image.
[0057] In practical applications, the initial Gaussian noise generated by this reverse inference algorithm endows the subsequent progressive random noise superposition process with specific properties. Because the initial Gaussian noise is derived from the characteristics of shrimp images, it can better simulate the interference that might occur to shrimp images in the real world during noise superposition. For example, different species of shrimp images have unique morphological and color characteristics, and the initial Gaussian noise generated by this algorithm will also bear the "imprint" of these characteristics. Therefore, the resulting noise image sequence after noise superposition can more accurately reflect the changes in shrimp images under different noise interferences.
[0058] In some embodiments, S300, the step of embedding a cue word guidance mechanism in the reverse denoising process of the diffusion model, using the latent variable sequence and cue word set as dual inputs, and outputting an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp, includes: S310, which uses text embedding technology to transform the set of prompt words into text embedding vectors; S320, starting from the initial Gaussian noise obtained by inversion, the shrimp image is gradually denoised and updated. In the iterative update at each time step, the text embedding vector is combined to guide the generation process, and the output is an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp.
[0059] In this embodiment, the set of cue words is transformed into text embedding vectors using text embedding technology. This step lays the foundation for incorporating cue word information into the subsequent reverse denoising process of the diffusion model. Text embedding vectors can accurately express the semantic information contained in the cue words in a numerical way, enabling the diffusion model to better understand and utilize this information during the denoising process.
[0060] In the process of progressively denoising and updating shrimp images starting from the initial Gaussian noise derived from the inverse model, each iterative update at every time step plays a crucial role. Combining text embedding vectors to guide the generation process allows the diffusion model to shape the generated image based on the shrimp biological characteristics contained in the prompts while removing noise. For example, the description of shrimp body shape in the prompts influences the model's shape restoration during denoising through the text embedding vectors, ensuring that the generated enhanced image conforms to the true characteristics of shrimp in terms of body shape. Similarly, information about shrimp color in the prompts guides the model to correctly restore or adjust the image's color during denoising, ensuring that the generated image's color matches the biological characteristics of shrimp.
[0061] This method, which incorporates text embedding vectors to guide the generation process in each iterative update at every time step, ensures that the generated enhanced images contain a degree of random variation while strictly adhering to the biological characteristics of shrimp. The presence of random variation increases image diversity, helping to improve the shrimp classification and recognition model's ability to identify images of shrimp with different morphologies. Adherence to the biological characteristics of shrimp guarantees the effectiveness of the generated images for model training, preventing the generated images from deviating too much from the actual characteristics of shrimp and causing the model to learn incorrect information.
[0062] It's important to note that excessive random variation can cause generated images to deviate from the true characteristics of shrimp, failing to provide effective data for model training. Conversely, overemphasizing biological characteristics may result in images lacking diversity, making it difficult for the model to learn comprehensive shrimp features. Therefore, it's necessary to find the optimal balance by adjusting relevant parameters, such as the weighting of the text embedding vector's influence during denoising.
[0063] In some embodiments, the iterative update formula for the shrimp image is:
[0064] in, , They are at the time step , Noisy images, , These represent the time steps. , The variance reduction factor These are the parameters of the diffusion model; In time step Noise estimation incorporating cue word information It is a noise vector sampled from a standard Gaussian distribution.
[0065] In this embodiment, the specific update steps of shrimp images in the diffusion model reverse denoising process are described in detail using the above iterative update formula. Each term in the formula has a clear physical meaning and function; they work together to gradually restore the noisy image to an enhanced image that contains random variations while conforming to the biological characteristics of shrimp.
[0066] variance reduction factor , These two factors determine the magnitude of noise variation between different time steps. By adjusting these two factors, the rate and extent of image denoising can be controlled. For example, when... and When the difference is small, the noise changes relatively smoothly in adjacent time steps, and the image denoising process is relatively stable; however, when the difference is large, the noise changes more drastically, the image denoising speed will be faster, but it may have some impact on the restoration of image details.
[0067] The noise estimation step, which incorporates cue word information at each time step, is crucial for integrating the shrimp biological characteristics contained in the cue words into the denoising process. Information such as body shape, color, and texture in the cue words, after being converted into text embedding vectors using text embedding technology, further influences the noise estimation. This allows the model to adjust the image according to the actual characteristics of the shrimp during denoising, avoiding the generation of images that do not conform to the biological characteristics of shrimp.
[0068] The noise vector, sampled from a standard Gaussian distribution, introduces a degree of randomness into image generation. This randomness contributes to the diversity of the generated enhanced images, preventing all enhanced images from exhibiting similar shapes. The degree of random variation in the generated images can be controlled by adjusting the sampling method or related parameters of the noise vector.
[0069] The shrimp image is denoised using the following formula to obtain an enhanced image: ; in, It is noise from random sampling. This represents the variance reduction factor. Indicates the percentage of noise added. ,0< <1, 0< <1; Indicated based on time step Noisy images Given time step t and condition c, predict the amount to be added to the noisy image. The estimated value of noise in the data; In a certain intermediate generation step have: ; in, This represents a denoising operation based on random noise. Indicates based on guidance information The guided operation, It is in the steps A coefficient that controls the guiding intensity; The enhanced image in step s represents the image from... Gradually denoise and restore the enhanced image.
[0070] In this embodiment, the calculation method for shrimp images in the diffusion model's reverse denoising process is further refined through the reverse denoising formula. The parameters in the formula work closely together, playing a crucial role in generating the enhanced image that meets the requirements.
[0071] First, randomly sampled noise This introduces uncertainty and diversity into the denoising process. Different sampling values will result in variations in detail in the final enhanced image, which is crucial for enriching the diversity of training data. In practical applications, appropriately setting the sampling range and distribution of noise can simulate various possible noise interference scenarios, thereby enabling shrimp classification and recognition models to learn more robust features.
[0072] By comprehensively considering time steps, noisy images, and cue words, the denoising process can closely follow the biological characteristics of shrimp. For example, when a cue word emphasizes a unique texture of a shrimp, this noise estimation term guides the model to focus on restoring or generating image details that match that texture during the denoising process.
[0073] In the intermediate generation step, denoising is performed based on random noise. With boot operations based on boot information They work together. The denoising operation removes noise from the image, while the guidance operation adjusts the image based on cue words to better reflect the biological characteristics of shrimp. The coefficient controlling the guidance intensity balances these two processes. If... Larger values result in a more significant impact of the guided operation on image generation, producing images that more strictly conform to shrimp biological characteristics, but may limit image diversity to some extent; conversely, if... Smaller values result in a more pronounced effect of random noise in image generation, increasing image diversity, but may reduce the accuracy in matching shrimp biological characteristics. Therefore, in practical applications, adjustments need to be made based on specific needs and experimental results. The value is adjusted to generate enhanced images that are both diverse and consistent with the biological characteristics of shrimp.
[0074] The process of deriving the initial Gaussian noise based on the diffusion model is as follows: After obtaining the prompt word for a specific shrimp species, we input the prompt word into a trained and optimized diffusion model. As a deep learning architecture, the diffusion model's core function is to learn data distribution characteristics and generate data. This diffusion model is based on a progressive noise superposition algorithm to build its training system, and its inverse process aims to reconstruct the original information from the noise data. In this invention, the diffusion model can accurately reconstruct the initial Gaussian noise field required to generate the target shrimp image based on the input prompt word through an inverse deduction algorithm. Gaussian noise, as a random signal source with clear statistical characteristics, serves as the basic noise signal carrier in the field of image generation. By systematically learning the distribution characteristics of massive shrimp image data, the diffusion model can accurately analyze the Gaussian noise distribution pattern matching the morphological characteristics of shrimp during the inverse deduction process. During this process, the diffusion model continuously optimizes the noise distribution parameters to ensure that it strictly follows the generation rules of shrimp images, thereby ensuring that the final generated noise signal can be transformed into shrimp images with high fidelity and rich diversity after subsequent processing. The following is the derivation process of the relevant formulas for generating shrimp images by combining the initial Gaussian noise and the prompt word: Diffusion models are typically built upon Markov chains, progressively recovering the original Gaussian noise from a noisy image. Assume... In time step Noisy images, It is noise in the model estimation. and These are pre-set time-related parameters. This represents the variance reduction factor. This indicates the proportion of noise added. That is, the proportion of signal retained in each step is complementary to the proportion of noise added; The range of values for is: 0 < <1; usually It is a very small positive number (e.g., 0.0001 to 0.02), which increases with time step t. It can be monotonically increasing or decreasing (depending on the scheduling strategy). The range of values for is: 0 < <1; due to ,and ∈(0,1), therefore It is also located between (0, 1).
[0075] In the backdiffusion process (recovering the initial Gaussian noise from the noisy image), the update formula for each iteration is:
[0076] in, Is the model at time step Noise estimation, , They are at the time step , Noisy images, , These represent the time steps. , The variance reduction factor These are the parameters of the diffusion model. (Based on time steps...) Iterate gradually to Ultimately, the initial Gaussian noise can be obtained through inversion. .
[0077] Next, shrimp images are generated based on the cue words; in the diffusion model, the cue words are typically converted into text embedding vectors using text embedding techniques (such as CLIP's text encoder). Assume the generation process starts from the inverse of the initial Gaussian noise. Beginning, at every step In the generation and update process, text embedding vectors are combined. This guides the generation process.
[0078] The update formula for shrimp image denoising is extended as follows:
[0079] Here It combines cue word information (through text embedding vectors) Noise estimation after ( ).
[0080] In actual generation, some sampling techniques are also used, such as DDIM (Denoising Diffusion Implicit Model) sampling; the final denoising process update formula is:
[0081] in, , They are at the time step , Noisy images, , These represent the time steps. , The variance reduction factor These are the parameters of the diffusion model; This represents the variance reduction factor. Indicates the percentage of noise added. ,0< <1, 0< <1; In time step Noise estimation incorporating cue word information It is a noise vector sampled from a standard Gaussian distribution.
[0082] noise vector To increase the diversity of the generated data, by reducing the initial Gaussian noise... By iterating step by step according to the above formula, a rich variety of shrimp images can eventually be obtained. These shrimp images will provide ample data augmentation samples for subsequent shrimp identification.
[0083] Through guided inverse denoising, from Gradually recover to ,set up These are inverse denoising parameters. The noise is based on model estimation (affected by guiding information c), and the inverse denoising formula is: ; in, It is noise from random sampling. It is a parameter that controls the degree of influence of random noise.
[0084] Next, we will discuss the fusion of guidance information and noise. To ensure that guidance information and random noise are effectively fused during the generation process, a fusion function is defined. In a certain intermediate generation step have: ; in, This represents a denoising operation based on random noise. Indicates based on guidance information The guided operation, It is in the steps A coefficient that controls the guiding intensity; The enhanced image in step s represents the image from... Gradually denoise and restore the enhanced image.
[0085] The specific implementation scheme of the present invention is as follows:
[0086] After acquiring images of shrimp taken on-site, an image feature extraction algorithm is used to analyze the shrimp in detail. This algorithm focuses on identifying typical characteristics of the shrimp, such as their body shape and outline—whether they are slender (like the Chinese white shrimp) or relatively short and stout (like the Litopenaeus vannamei); the color and texture of the shrimp's body—the Chinese white shrimp is usually bluish-gray with fine longitudinal stripes, while the Litopenaeus vannamei is mostly light bluish-gray with less obvious horizontal stripes; and detailed features such as the shape of the shrimp's antennae and eyes.
[0087] Based on these extracted features, combined with a pre-built shrimp feature knowledge base containing typical feature descriptions of different shrimp species, we can generate specific cue words corresponding to the shrimp in the photo. The method for generating specific cue words is as follows: using a cue word inversion tool, such as Joy Caption, we first normalize the input image, then extract visual features using ViT, and finally input the visual features into the text decoder CLIP to invert the cue words. For example, for a photo of a Chinese white shrimp, the generated cue words might be "slender shrimp, bluish-gray body color, with fine longitudinal stripes, long and segmented antennae, and black round eyes," etc. These cue words accurately describe the features of the shrimp in the photo, providing clear guidance for subsequent diffusion model inversion.
[0088] Next, the feature-to-cue word mapping is performed. Specifically, a feature-to-cue word mapping model M is constructed, which maps the feature vector V to a vector representation T consisting of cue words. Here, a deep learning-based mapping network can be used, such as a Long Short-Term Memory (LSTM) network or a Transformer network.
[0089] Assuming we use a Transformer network, its core is the self-attention mechanism. The self-attention mechanism emphasizes important features by calculating attention scores between features at different locations. After processing through the self-attention mechanism and multiple layers of Transformer blocks, we obtain the feature-cue word mapping model, i.e., the vector representation T. Specifically, we input the features and their corresponding cue words into the Transformer network for training, resulting in a trained Transformer network, which serves as the feature-cue word mapping model.
[0090] Finally, a set of prompt words is generated. It can be generated by decoding and structuring the initial set of prompt words T.
[0091] This invention provides a novel technical approach for shrimp classification and recognition by applying the diffusion model image inversion method to shrimp image data augmentation. This method effectively addresses the shortcomings of traditional data augmentation methods and generative adversarial networks, generating high-quality, diverse, and biologically consistent shrimp images. This provides ample high-quality data support for training shrimp classification models and promotes the in-depth application of artificial intelligence technology in aquaculture.
[0092] The method provided by this invention will now be verified.
[0093] Table 1: Shrimp-related cue word inversion
[0094] refer to Figure 2 The dataset used in this experiment contains image data of Chinese shrimp of category 1, such as... Figure 2 As shown in (a), the dataset is divided into training, validation, and test sets in a 7:2:1 ratio. In the data augmentation experiment, the diffusion inversion data augmentation method of this scheme is applied to the training set to generate more training samples.
[0095] Baseline experiments (direct ResNet-18 recognition): The original dataset was fed into a pre-trained ResNet-18 model, all layers of the model were frozen, and only the last fully connected layer was used for fine-tuning. The classification accuracy of the model was evaluated on the test set.
[0096] Model fine-tuning experiment: Unfreeze the last few layers of the ResNet-18 model for fine-tuning, adjust hyperparameters such as the learning rate, train the model on the training set, and evaluate the accuracy on the test set.
[0097] Data augmentation experiments using this method: First, the diffusion inversion data augmentation method of this scheme is applied to the training set to generate new training samples. Then, the original training set and the augmented training set are merged to obtain the following result: Figure 2 The augmented dataset shown in (b) is input into the fine-tuned ResNet-18 model for training, and the model accuracy is finally evaluated on the test set.
[0098] Table 2: Comparison of Data Augmentation Classification Results
[0099] The results show that using the pre-trained ResNet-18 directly can achieve a certain classification accuracy, and the accuracy is improved after fine-tuning the model. However, after using the diffusion inversion data augmentation method proposed in this scheme, the classification accuracy of the model is further improved, indicating that this method has good effect in shrimp classification and recognition tasks.
[0100] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0101] Improving shrimp classification accuracy: By expanding the training set size and improving data quality, this invention can significantly improve the performance of the shrimp classification and recognition model, increase classification accuracy, and provide accurate technical support for shrimp seedling sorting, aquaculture management, and other related tasks.
[0102] Reduce farming costs: Automated and accurate classification of shrimp can replace manual sorting, improve sorting efficiency, reduce labor costs, and reduce economic losses caused by sorting errors.
[0103] Promoting the intelligent development of aquaculture: This invention introduces advanced artificial intelligence technology into the field of aquaculture, which helps to promote the intelligent and automated development of aquaculture, improve aquaculture efficiency and quality, and promote the upgrading and transformation of the aquaculture industry.
[0104] This invention improves the level of intelligence in shrimp farming, which helps to achieve green and efficient development of aquaculture.
[0105] and Figure 1 The corresponding method is referenced. Figure 3 This invention provides a data augmentation system based on a diffusion model, comprising: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.
[0106] It is evident that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0107] Furthermore, embodiments of the present invention also disclose a computer program product or computer program stored in a computer-readable storage medium. A processor of a computer device can read the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform the described method. Similarly, the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0108] It will be understood by those skilled in the art that all or some of the methods and systems disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0109] The above is a detailed description of the preferred embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.
Claims
1. A data augmentation method based on a diffusion model, characterized in that, The method includes the following steps: S100, Obtain the original dataset containing shrimp images, and generate a set of prompt words corresponding to the shrimp images; S200, based on the forward noise addition mechanism of the diffusion model, the shrimp image is subjected to progressive random noise superposition processing to obtain the latent spatial representation under different noise intensities, forming a noisy latent variable sequence; the latent variable sequence includes the noisy image at each time step; S300 uses a sequence of latent variables and a set of cue words as dual inputs. It embeds a cue word guidance mechanism in the reverse denoising process of the diffusion model and outputs an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp. S400 integrates the generated enhanced images into the original dataset to obtain the enhanced dataset.
2. The method according to claim 1, characterized in that, In step S100, obtaining the original dataset containing shrimp images and generating a set of prompt words corresponding to the shrimp images includes: S110, acquire shrimp images, identify typical features of shrimp in the shrimp images, and form a feature vector from the typical features of shrimp; the typical features include the shrimp's body outline, the color and texture of the shrimp's body, and the shape of the shrimp's antennae and eyes; S120, combining the typical features and the pre-built shrimp feature knowledge base, generate prompt words corresponding to the shrimp in the shrimp image; the shrimp feature knowledge base contains prompt words for different varieties of shrimp, and the prompt words are used to describe the typical features of the corresponding shrimp. S130, Based on the typical characteristics of shrimp and the corresponding prompt words, construct a feature-prompt word mapping model, and map the feature vector of shrimp to a vector representation composed of prompt words through the mapping model; S140, determine the feature score of each prompt word in the vector representation. When the feature score of a prompt word exceeds the feature score threshold, the prompt word is used as the prompt word corresponding to the shrimp image and included in the prompt word set.
3. The method according to claim 2, characterized in that, In S130, the construction of a feature-cue word mapping model based on typical characteristics of shrimp and corresponding cue words includes: S131, Form a feature vector from the typical characteristics of each shrimp species, and determine the prompt words corresponding to the feature vector; S132, input the feature vector and the corresponding prompt word into the Transformer network for training, and obtain the trained Transformer network as the feature-prompt word mapping model.
4. The method according to claim 1, characterized in that, In S200, the progressive random noise superposition processing of shrimp images based on the forward noise addition mechanism of the diffusion model, recording the latent spatial representation under different noise intensities, and forming a noisy latent variable sequence includes: S210, The prompt words in the prompt word set are input into the trained and optimized diffusion model. The diffusion model reconstructs the initial Gaussian noise required to generate the shrimp image based on the input prompt words through a reverse inference algorithm. S220, Based on the initial Gaussian noise field, the shrimp image is subjected to progressive random noise superposition processing to obtain noise images at each time step, and the noise images at each time step are sequentially arranged to form a latent variable sequence.
5. The method according to claim 4, characterized in that, The initial Gaussian noise required to reconstruct the shrimp image using the reverse deduction algorithm includes: The shrimp image is back-diffused using the following formula, by iterating the shrimp image from the maximum time step to the minimum time step, to obtain the initial Gaussian noise of the shrimp image: in, , They are at the time step , Noisy images, For time steps, For the maximum time step, , These represent the time steps. , The variance reduction factor These are the parameters of the diffusion model; Is the model at time step Noise estimation.
6. The method according to claim 1, characterized in that, In S300, the method of using the latent variable sequence and the cue word set as dual inputs, embedding a cue word guidance mechanism in the reverse denoising process of the diffusion model, and outputting an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp, includes: S310, which uses text embedding technology to transform the set of prompt words into text embedding vectors; S320, starting from the initial Gaussian noise obtained by inversion, the shrimp image is gradually denoised and updated. In the iterative update at each time step, the text embedding vector is combined to guide the generation process, and the output is an enhanced image that contains both random variations and conforms to the biological characteristics of shrimp.
7. The method according to claim 6, characterized in that, The iterative update formula for the shrimp image is: in, , They are at the time step , Noisy images, , These represent the time steps. , The variance reduction factor These are the parameters of the diffusion model; In time step Noise estimation incorporating cue word information It is a noise vector sampled from a standard Gaussian distribution; The shrimp image is denoised using the following formula to obtain an enhanced image: ; in, It is noise from random sampling. This represents the variance reduction factor. Indicates the percentage of noise added. ,0< <1, 0< <1; Indicated based on time step Noisy images Given time step t and condition c, predict the amount to be added to the noisy image. The estimated value of noise in the data; In a certain intermediate generation step have: ; in, This represents a denoising operation based on random noise. Indicates based on guidance information The guided operation, It is in the steps A coefficient that controls the guiding intensity; The enhanced image in step s represents the image from... Gradually denoise and restore the enhanced image.
8. A data augmentation system based on a diffusion model, characterized in that, The system includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1 to 7.