Data augmentation method, device and medium for class imbalance plankton image classification

CN122551334APending Publication Date: 2026-08-11ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

(1)传统数据增强方法(如几何变换、颜色抖动、噪声注入等):此类方法产出的样本往往与原始数据高度相关,可提供的有效增量信息非常有限,难以让模型学习到样本间更深层次的变化与不变性特征,对于模型性能的提升有限

Benefits of technology

[0019]与现有技术相比,本发明的有益效果是,本发明提出了一种基于改进潜在扩散模型(LDM)的数据增强方法(Plankton-LDM)。其中,全新的对去噪网络U-Net能够高效捕捉全局长距离像素依赖,聚合全局上下文信息,确保了所生成的浮游生物多变且复杂的宏观结构的完整性,增强了网络对细粒度特征的表达能力,聚焦高频特征重建,精准还原了浮游生物的复杂纹理与细长附肢等边缘细节。二阶精度的Heun采样策略,大幅降低了扩散去噪过程中的离散化误差,避免了关键细节的模糊与结构性伪影,在保证计算效率的同时,极大提高了生成影像的保真度和视觉感知质量。实验结果表明,本发明生成的高保真浮游生物影像在FID(21.12)、IS(2.68)和NIQE(6.33)等评价指标上均优于原始LDM与主流的GAN网络。在使用本发明增强后的数据集训练的多种主流分类模型上验证,这些分类模型的mAP50-95和F1分数分别平均提高了约3.7%和约3.0%,显著优于传统增强策略,证明了本发明的有效性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551334A_ABST
    Figure CN122551334A_ABST
Patent Text Reader

Abstract

This invention discloses a data augmentation method, device, and medium for classifying imbalanced plankton images. The method includes the following steps: acquiring a raw dataset of marine plankton images, dividing it into a training set, a validation set, and a test set, and labeling the plankton to obtain labeled training, validation, and test sets; constructing a Plankton-LDM network model based on a latent diffusion model; the Plankton-LDM network model includes a pre-trained autoencoder for data space transformation and an improved noise prediction network U-Net for performing denoising tasks; iteratively training the Plankton-LDM network model using the labeled training set to obtain a trained Plankton-LDM network model; using the trained Plankton-LDM network model, combined with a second-order Heun sampling strategy, performing a reverse denoising process to obtain a high-quality set of marine plankton images; and finally constructing a class-balanced augmented training set and inputting it into a deep learning image classification model for iterative training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image classification technology, and in particular to a data augmentation method, device and medium for classifying imbalanced planktonic images. Background Technology

[0002] Precise classification and community composition analysis of marine plankton are crucial for revealing the patterns of marine energy flow and predicting trends in marine environmental change (Anderson, S., Barton, A., Clayton, S., Dutkiewicz, S., & Rynearson, T. (2021). Marine phytoplankton functional types exhibit diverse responses to thermal change. Nature communications, 12(1), 6413.)(Sun, D., Huan, Y., Wang, S., Qiu, Z., Ling, Z., Mao, Z., & He, Y. (2019). Remotesensing of spatial and temporal patterns of phytoplankton assemblages in the Bohai Sea, Yellow Sea, and east China sea. Water Research, 157, 119-133.). Traditionally, marine plankton monitoring has relied primarily on trawls (Ratnarajah, L., Abu-Alhaija, R., Atkinson, A., Batten, S., Bax, NJ, Bernard, KS, Canonico, G., Cornils, A., Everett, JD, & Grigoratou, M. (2023). Monitoring and modelling marinezooplankton in a changing climate. Nature communications, 14(1), 564.) and the Niskin water sampler (Liu, Z., Stewart, G., Cochran, JK, Lee, C., Armstrong, RA, Hirschberg, DJ, Gasser, B., & Miquel, J.-C. (2005).Why do POC concentrations measured using Niskin bottle collections sometimes differ from those using in-situ pumps? Deep Sea Research Part I: Oceanographic Research Papers, 52(7), 1324-1344.) and other tools are used for on-site sampling and manual microscopic analysis (Engel, MS, Ceríaco, LM, Daniel, GM, Dellapé, PM, Löbl, I., Marinov, M., Reis, RE, Young, MT, Dubois, A., & Agarwal, I. (2021). The taxonomic impediment: a shortage of taxonomists, not the lack of technical approaches. In (Vol.193, pp. 381-387): Oxford University Press UK.), this method is time-consuming, labor-intensive, costly, and has obvious subjectivity and lag (Alfano, PD, Rando, M., Letizia, M., Odone, F., Rosasco, L., & Pastore, VP (2022). Efficient unsupervised learning for planktonimages. 2022 26th International Conference on Pattern Recognition (ICPR).) (Eerola, T., Batrakhanov, D., Barazandeh, NV, Kraft, K., Haraguchi, L.,Lensu, L., Suikkanen, S., Seppälä, J., Tamminen, T., & Kälviäinen, H. (2024). Survey of automatic plankton image recognition: challenges, existing solutions and future perspectives.Artificial Intelligence Review, 57(5), 114.). In recent years, deep learning-driven image processing technology has significantly improved the efficiency of automatic classification of plankton (Yang, Z., Li, J., Chen, T., Pu, Y., & Feng, Z. (2022). Contrastive learning-based image retrieval for automatic recognition of in situ marine plankton images. ICES Journal of Marine Science, 79(10), 2643-2655.). However, the high performance of deep learning models is highly dependent on massive, high-quality, and balanced training samples. In natural marine environments, the abundance of different planktonic groups varies greatly, leading to severe data scarcity and class imbalance in acquired image data. This easily causes overfitting in deep learning models, reducing the reliability and generalization ability of classification (Pala, A., Oleynik, A., Utseth, I., & Handegard, NO (2023). Addressing class imbalance in deep learning for acoustic target classification. ICES Journal of Marine Science, 80(10), 2530-2544.)(Sorochan, K., Vaswani, AR, Howard, A., Rühl, S., O'Grady, E., Möller, KO, & Johnson, CL (2026). Practical guidance on automated sorting of underwater images in plankton ecology research. ICESJournal of Marine Science, 83(1), fsaf212.).

[0003] To address the aforementioned issues, existing technologies typically employ data augmentation methods to provide sufficient data support for deep learning models, but these methods still suffer from the following drawbacks: (1) Traditional data augmentation methods (such as geometric transformation, color jitter, noise injection, etc.): The samples produced by these methods are often highly correlated with the original data, and the effective incremental information they can provide is very limited. It is difficult for the model to learn deeper changes and invariant features between samples, and the improvement of model performance is limited.

[0004] (2) Image data generation method based on variational autoencoder (VAE): This type of method takes minimizing reconstruction error as the main optimization goal. The generated samples are prone to blurring and averaging, and lack the ability to preserve the fine-grained morphological features of plankton.

[0005] (3) Image data generation methods based on generative adversarial networks (GANs): Although they can generate clearer samples compared to VAEs, the training process of their generators and discriminators is difficult to achieve a dynamic balance, which can easily lead to problems such as training instability and pattern collapse. Especially for species categories with only a small number of samples, it is difficult to learn the complete intra-class morphological distribution, which leads to a surge in the risk of generating invalid samples with artifacts and distortion.

[0006] (4) Existing standard diffusion model (DM): Although it performs well in the field of image generation, its default sampling strategy (DDIM) usually adopts a first-order approximation solution. When faced with marine plankton images with complex textures and slender appendages, it will introduce a large discretization error, resulting in blurring of key details and even structural artifacts. In addition, its original denoising network (U-Net) is limited by the local receptive field of convolution operation, making it difficult to establish global long-distance dependencies between image pixels and unable to accurately reconstruct the variable macroscopic structure of plankton.

[0007] In summary, existing image processing and generation technologies, limited by local receptive fields and lacking the ability to characterize complex nonlinear statistical dependencies, struggle to balance global structural consistency with the accurate reproduction of fine-grained features when processing images with intricate shapes and textures. Summary of the Invention

[0008] The purpose of this invention is to address the shortcomings of existing technologies by providing a data augmentation method, device, and medium for classifying imbalanced planktonic images.

[0009] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In a first aspect, the present invention provides a data augmentation method for classifying imbalanced planktonic images, comprising the following steps: S1 Data Preparation and Preprocessing: Obtain the original image dataset of marine plankton, divide it into training set, validation set and test set, and label the plankton in it to obtain training set, validation set and test set with label information; S2 Constructing the Plankton-LDM network model and latent space mapping: Constructing a Plankton-LDM network model based on the latent diffusion model; the Plankton-LDM network model includes a pre-trained autoencoder for data space transformation and an improved noise prediction network U-Net for performing denoising tasks; S3 Model Training: The Plankton-LDM network model is iteratively trained using a training set with labeled information to obtain a trained Plankton-LDM network model; S4 generates high-quality marine plankton images: using a trained Plankton-LDM network model, combined with a second-order Heun sampling strategy, a reverse denoising process is performed to obtain a high-quality marine plankton image set; S5 constructs a class-balanced augmented training set and inputs it into a deep learning image classification model for iterative training.

[0010] Further, in step S1, the annotation specifically involves: labeling the plankton in the image with rectangular boxes on the training set, validation set, and test set to obtain the training set, validation set, and test set with annotation information.

[0011] Furthermore, the labeling information includes the coordinates of four points (x1, y1), (x2, y2), (x3, y3), and (x4, y4) arranged clockwise within a rectangle, and the category ID of each plankton species.

[0012] Further, in step S2, the data space transformation specifically involves: inputting the labeled training set into a pre-trained autoencoder, wherein the labeled training set is a high-dimensional marine planktonic image, and the encoder compresses it into a low-dimensional latent spatial representation.

[0013] Further, in step S2, the improved noise prediction network U-Net specifically involves: introducing a VSS module to replace the standard convolutional block in the second stage of the downsampling path of the noise prediction network U-Net; and introducing HDRAB to replace the traditional residual block in the third and fourth stages of the downsampling path of the noise prediction network U-Net.

[0014] Further, in step S3, the iterative training of the Plankton-LDM network model specifically involves adding noise to the potential representation through a forward diffusion process and predicting the noise through an improved noise prediction network U-Net to complete the reverse denoising learning, thereby obtaining the trained Plankton-LDM network model.

[0015] Further, step S4 specifically involves: starting from random sampling noise that follows a standard normal distribution, performing inverse denoising in the latent space using a Heun sampler employing the second-order Runge-Kutta method. Within each denoising time step, a dual gradient evaluation is performed for prediction and correction, once for predicting the next latent representation state and once for correction. The weighted average of the two gradients is then used to perform the final latent representation update. The Heun sampler accurately approximates the true ordinary differential equation denoising trajectory through a second-order precision integral step, obtaining a progressively restored clear latent space representation. This progressively restored clear latent space representation is then mapped back to the image pixel space through the decoder of the pre-trained autoencoder, thereby generating the high-quality marine plankton image set.

[0016] Furthermore, the construction of the class-balanced augmented training set specifically involves: supplementing the training set with the high-quality marine plankton images for categories whose sample numbers are less than those of other categories; and expanding the sample numbers of each category to be consistent to obtain the class-balanced augmented training set.

[0017] In a second aspect, the present invention provides an electronic device, including a memory and a processor, characterized in that the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the data augmentation method for classifying imbalanced plankton images described above.

[0018] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the data augmentation method for classifying imbalanced planktonic images.

[0019] Compared with existing technologies, the beneficial effects of this invention are that it proposes a data augmentation method (Plankton-LDM) based on an improved latent diffusion model (LDM). The novel denoising network U-Net efficiently captures global long-range pixel dependencies and aggregates global contextual information, ensuring the integrity of the generated plankton's varied and complex macroscopic structure. It enhances the network's ability to express fine-grained features, focuses on high-frequency feature reconstruction, and accurately restores the complex textures and slender appendages of plankton. The second-order precision Heun sampling strategy significantly reduces discretization errors in the diffusion denoising process, avoiding blurring of key details and structural artifacts. While ensuring computational efficiency, it greatly improves the fidelity and visual perception quality of the generated image. Experimental results show that the high-fidelity plankton images generated by this invention outperform the original LDM and mainstream GAN networks in evaluation metrics such as FID (21.12), IS (2.68), and NIQE (6.33). The results were validated on several mainstream classification models trained using the dataset enhanced by this invention. The average mAP50-95 and F1 scores of these models were improved by approximately 3.7% and 3.0%, respectively, which are significantly better than traditional enhancement strategies, demonstrating the effectiveness of this invention.

[0020] This invention effectively solves the overfitting problem of deep learning classification models caused by the scarcity of marine planktonic image data and the imbalance of category distribution, and establishes a transferable data augmentation paradigm for cross-domain data-constrained scenarios. Attached Figure Description

[0021] Figure 1 This is a structural diagram of the improved noise prediction network U-Net; Figure 2 This is a structural diagram of the VSS module; Figure 3 This is a structural diagram of HDRAB; Figure 4 This is a structural diagram of Plankton-LDM; Figure 5 It is a supplementary training set of images; Figure 6 This is a graph showing the results of a comparative experiment; Figure 7 This is a structural diagram of the electronic device provided by the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] It should be noted that, unless otherwise specified, the features in the following embodiments and implementation methods can be combined with each other.

[0024] The core of this invention lies in proposing a data augmentation method based on an improved latent diffusion model, namely Plankton-LDM, to solve the class imbalance problem in deep learning-based plankton image classification. Specifically, this invention proposes a data augmentation framework based on an improved latent diffusion model, named Plankton-LDM (e.g., ...). Figure 4 By aligning feature-level distributions in the latent space through a pre-trained autoencoder and combining it with an improved denoising network, the problems of data scarcity and class imbalance are effectively solved, providing high-quality data support for downstream classification tasks. A denoising network, U-Net, integrating VSS and HDRAB, is proposed. It achieves the construction of global long-range pixel dependencies and accurate reconstruction of local high-frequency fine-grained features, ensuring the integrity of the variable structure of plankton and the restoration of complex textures. A high-fidelity image generation mechanism based on a second-order Heun sampling strategy is adopted. Through "prediction-correction" dual gradient evaluation, the accumulation of discretization errors and structural artifacts caused by traditional first-order solvers are effectively suppressed, significantly improving the fidelity of the synthesized image while ensuring generation efficiency.

[0025] Specifically, the technical solution of the present invention can be summarized as follows: This invention proposes a data augmentation framework based on an improved latent diffusion model, named Plankton-LDM (e.g., Figure 4 By performing feature-level distribution alignment in the latent space through a pre-trained autoencoder and combining it with an improved denoising network, the problems of data scarcity and class imbalance are effectively solved, providing high-quality data support for downstream classification tasks. A denoising network, U-Net, integrating VSS and HDRAB, is proposed. It achieves the construction of global long-range pixel dependencies and the accurate reconstruction of local high-frequency fine-grained features, ensuring the integrity of the variable structure of plankton and the restoration of complex textures. A high-fidelity image generation mechanism based on a second-order Heun sampling strategy is adopted. Through "prediction-correction" dual gradient evaluation, the accumulation of discretization errors and structural artifacts caused by traditional first-order solvers are effectively suppressed, and the fidelity of the synthesized image is greatly improved while ensuring generation efficiency.

[0026] In a first aspect, the present invention proposes a data augmentation method for classifying imbalanced planktonic images, comprising the following steps: S1 Data Preparation and Preprocessing: Obtain the original image dataset of marine plankton with class imbalance characteristics, divide it into training set, validation set and test set, and label the plankton in it to obtain the training set, validation set and test set with label information.

[0027] S1.1: Establish a raw image dataset of marine plankton and divide it into a training set, a validation set, and a test set in a ratio of 7:2:1, with no overlap between the three sets.

[0028] S1.2: Plankton Labeling. Using the image labeling software LabelImg, plankton within the images are sequentially labeled with rectangular bounding boxes on the training, validation, and test sets. The labeling results are saved as txt files, thus obtaining the training, validation, and test sets with labeled information. The labeled information includes the coordinates of four points of the rectangular bounding box arranged clockwise: (x1, y1), (x2, y2), (x3, y3), and (x4, y4), and the category ID of each plankton species.

[0029] S2. Construction of the Plankton-LDM Network Model and Latent Space Mapping: A Plankton-LDM network model based on the Latent Diffusion Model (LDM) is constructed. This network model includes a pre-trained autoencoder for data space transformation and an improved noise prediction network U-Net for performing denoising tasks. Specifically, the downsampling path of the improved noise prediction network U-Net incorporates a Visual State Space (VSS) module and a Hierarchical Expanded Residual Attention Block (HDRAB); the improved noise prediction network U-Net is as follows... Figure 1 As shown.

[0030] Input: The training set (high-dimensional marine planktonic images) output from step S1 and the original U-Net network architecture.

[0031] S2.1: Latent Space Representation Mapping (Autoencoder Part): The labeled training set obtained in step S1 is input into the pre-trained autoencoder. The labeled training set is a high-dimensional marine plankton image. The encoder compresses it into an information-dense low-dimensional latent space representation and performs diffusion and denoising processes in the latent space. While preserving the key semantic and structural information of the image, high-frequency redundant information at the pixel level is filtered out to reduce memory overhead and focus on the core semantic features of the image.

[0032] This step transforms the image data of the labeled training set from a high-dimensional pixel space to a low-dimensional latent space, providing a data foundation for subsequent diffusion and denoising processes in the latent space.

[0033] S2.2: Constructing the Improved Noise Prediction Network (Part 1); This step introduces the VSS module to improve the downsampling path of the noise prediction network U-Net in order to accurately capture the complex and varied morphological characteristics of plankton.

[0034] Specifically, a VSS module is introduced in the second stage of the downsampling path of the U-Net noise prediction network to replace the standard convolutional block. The VSS module can efficiently capture long-range dependencies between pixels in two-dimensional space with linear computational complexity, thereby aggregating global contextual information and ensuring the integrity of the generated image's macroscopic structure. The VSS module is as follows: Figure 2 As shown.

[0035] S2.3: The second part of constructing the improved noise prediction network; this step introduces HDRAB, combined with S2.2, by introducing HDRAB in the third and fourth levels of the downsampling path of the U-Net noise prediction network to replace the traditional residual blocks. HDRAB enhances feature extraction and representation capabilities by fusing multi-level residual connections and spatial attention mechanisms, making it more focused on the reconstruction of local fine high-frequency features such as complex textures and slender appendages on the surface of plankton. HDRAB is as follows... Figure 3 As shown.

[0036] The low-dimensional latent space representation output in this step serves as the direct object of operation for the forward diffusion noise addition process in the subsequent step S3. The improved noise prediction network U-Net structure, jointly constructed by S2.2 and S2.3, serves as the network carrier for the reverse denoising learning in the subsequent step S3. S2.1, S2.2, and S2.3 together constitute the Plankton-LDM to be trained.

[0037] S3 Model Training: In this step, the Plankton-LDM network model constructed in step S2 is iteratively trained using the training set with labeled information from step S1 to obtain the trained Plankton-LDM network model.

[0038] Noise is added to the latent representation through a forward diffusion process, and the noise is predicted by an improved noise prediction network, U-Net, to complete the inverse denoising learning. This yields the trained Plankton-LDM network model, whose structure is as follows: Figure 4 As shown.

[0039] S4 generates high-quality marine plankton images: Using a trained Plankton-LDM network model, combined with a second-order Heun sampling strategy, a reverse denoising process is performed to obtain a high-quality set of marine plankton images.

[0040] S4.1: Starting from random sampling noise that follows a standard normal distribution, the Heun sampler of the second-order Runge-Kutta method is used to perform inverse denoising in the latent space. Specifically, the Heun sampler is used as an ordinary differential equation solver to execute the inverse diffusion Markov chain.

[0041] S4.2: Within each denoising time step, a dual gradient evaluation is performed for prediction and correction, once for predicting the next potential representation state and once for correction; and a weighted average of the two gradients is used to perform the final potential representation update. The Heun sampler can accurately approximate the real ordinary differential equation denoising trajectory through a second-order precision integration step, significantly reducing discretization error and obtaining a clear potential space representation that is gradually restored.

[0042] S4.3: The gradually restored clear latent spatial representation is mapped back to the image pixel space through the decoder of the pre-trained autoencoder to generate a high-fidelity and morphologically rich high-quality marine planktonic image set.

[0043] S5 constructs a class-balanced augmented training set and inputs it into a deep learning image classification model for iterative training.

[0044] Samples are randomly selected from the high-quality marine plankton image set generated in step S4 and added to the pre-divided training set obtained in step S1.1, so that each category in the training set has 1500 samples. Figure 5 The image was labeled using the image annotation software LabelImg to obtain a class-balanced augmented training set, which was then input into the downstream classification model for training and evaluation.

[0045] Specifically, for categories with fewer samples than other categories in the training set, i.e., plankton with fewer than 1500 samples in the training set, high-quality marine plankton images of the corresponding category generated in step S4 are used to supplement the training set; the number of samples in each category is expanded to be consistent, eliminating the difference in abundance between categories, thereby constructing a category-balanced enhanced training set; the enhanced training set is input into the deep learning image classification model for iterative training to reduce the risk of overfitting caused by data scarcity and improve the generalization ability and classification accuracy of the classification model on the test set.

[0046] Example In this embodiment, the data comes from the publicly available dataset DYB-PlanktonNet, which was collected by an underwater dark-field imaging system. It includes 8 categories ( Penilia avirostris , Gammarids , Cumacea , Acartia , Caligus , Polychaeta , Macrura and Calanoid The dataset contains 5833 images, which exhibit significant class imbalance. To assess this, the original dataset was strictly divided into a training set (4083 images), a validation set (1167 images), and a test set (583 images) in a 7:2:1 ratio, as shown in Table 1.

[0047] Table 1. Details of the original dataset partitioning

[0048] The effectiveness of this embodiment can be further illustrated by the following experiment: The experimental environment and conditions for this invention are as follows: CPU: Intel Xeon Platinum 8358P @2.60GHz GPU: NVIDIA GeForce RTX 4090 (24GB) RAM: 43GB Software environment: PyTorch 2.0.0 + cu118, Python 3.8.10 Operating System: Ubuntu 20.04 The parameter settings for training the model in this invention are as follows: In the generation experiment, the total number of training epochs for the model's noise prediction network was set to 500, the image resolution to 256×256, the batch size to 32, and the initial learning rate to 0.0001, which was eventually decayed to 0 using cosine annealing. The optimizer was Adam. β 1 is set to 0.95. β The value was set to 0.999 to ensure the stability of model training. During forward diffusion, Gaussian noise was gradually added to the latent representation through a Markov chain. During backward diffusion, the model weights were updated by minimizing the error between the predicted noise and the actual added noise.

[0049] During classification model training, the total number of training epochs was set to 100, the number of warm-up epochs to 3, the batch size to 32, and the number of data loading threads to 32. Other hyperparameter settings were as follows: the initial learning rate was set to 0.01, the final learning rate decayed to 0.0001 using cosine annealing, the optimizer was stochastic gradient descent (SGD), the momentum parameter was 0.937, and the initial momentum during the warm-up phase was set to 0.8. During the inference evaluation phase, the intersection-over-union (IoU) threshold for non-maximum suppression was set to 0.6, and the confidence threshold was set to 0.6.

[0050] In this embodiment, the sampling step count is set to 20 steps to maintain high perceptual quality while also considering computational efficiency. After denoising, a clear latent spatial representation is obtained, which is finally mapped back to the image pixel space through the decoder of the autoencoder, generating a certain number of high-quality marine plankton image sets (e.g., 3000 images) for each category.

[0051] In this embodiment, by randomly selecting images, each classification image in the original training set is supplemented to 1500 images using generated images, resulting in a Plankton-LDM augmented dataset containing a total of 12000 images. This dataset is then input into mainstream deep learning classification networks (such as the YOLO series) for training, effectively addressing the risk of model overfitting due to data scarcity and significantly improving the accuracy of automated classification of marine plankton.

[0052] In view of the high variability of individual morphology and the complexity of texture features of marine planktonic organisms, this study selects FID, IS and NIQE as quality assessment indicators for generated images, and the calculation formulas are as follows.

[0053]

[0054] in, and These are the average vectors of the real image and the generated image, respectively. and These are the covariance matrices of the real image and the generated image, respectively. Let be the trace of the matrix.

[0055]

[0056] in, This represents the mathematical expectation of all generated images. This represents the image distribution induced by the generative model. For Kullback-Leibler divergence, This is the conditional probability prediction of the model for image x. This represents the marginal probability distribution of the generated image set.

[0057]

[0058] in, and Let represent the mean vectors of the MVG model for natural scenes and the MVG model for distorted images, respectively. and This corresponds to the covariance matrix of the two types of models.

[0059] The performance of the classification model was evaluated using mAP50-95 and F1 score, respectively. The calculation formulas are as follows:

[0060] In formula (5), N Represents the total number of categories. mAP The average accuracy is denoted as mAP50-95. This study uses mAP50-95 as the accuracy evaluation index, where the Intersection over Union (IoU) threshold increases from 50% to 95% in increments of 5%, for a total of ten thresholds. The mAP is calculated for each threshold, and the average of these ten thresholds is the mAP50-95. In formulas (6)-(8), TP , FP and FN These represent true positive, false positive, and false negative samples, respectively. P and R These represent precision and recall, respectively.

[0061] Experimental results show that the high-fidelity planktonic images generated by this invention outperform the original LDM and mainstream GAN networks in evaluation metrics such as FID (21.12), IS (2.68), and NIQE (6.33) (as shown in Table 2). Validation was performed on various mainstream classification models trained using the dataset enhanced by this invention. These models showed an average improvement of approximately 3.7% in mAP50-95 and approximately 3.0% in F1 scores (as shown in Tables 3 and 4), significantly outperforming traditional enhancement strategies and demonstrating the effectiveness of this invention.

[0062] This invention effectively solves the overfitting problem of deep learning classification models caused by the scarcity of marine planktonic image data and the imbalance of category distribution, and establishes a transferable data augmentation paradigm for cross-domain data-constrained scenarios.

[0063] Table 2 Performance Comparison of Different Generative Models

[0064] Note: In the table, PA, GA, CU, AC, CA, PO, MA, and CL correspond to plankton categories respectively. Penilia avirostris , Gammarids , Cumacea , Acartia , Caligus , Polychaeta , Macrura , Calanoid Table 3. Comparison of mAP50-95 for different classification models

[0065] Table 4. Comparison of F1 scores for different classification models

[0066] Using all the parameter values ​​listed in the specific implementation, we obtained Figure 6 The comparative experimental results shown demonstrate the effectiveness of the Plankton-LDM method of this invention: images generated by the original LDM and other GAN series models generally suffer from problems such as blurred details, artifacts, or structural distortion. For example, in categories PA and AC, the generated samples have blurred edges, and fine structures such as appendages are not effectively generated or displayed. In contrast, the images generated by Plankton-LDM show outstanding realism and restoration of biological details. For example, categories PA and CA have clear textures, categories GA and CU have detailed appendages, and all categories have realistic morphologies. This accurate restoration of details and morphological differences fully demonstrates the scientific validity and effectiveness of the method proposed in this study.

[0067] like Figure 7 As shown, this application provides an electronic device including a memory 101 for storing one or more programs and a processor 102. When the one or more programs are executed by the processor 102, they implement the method as described in any of the first aspects above.

[0068] The system also includes a communication interface 103. The memory 101, processor 102, and communication interface 103 are electrically connected directly or indirectly to each other to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, and the processor 102 executes various functional applications and data processing by executing the software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.

[0069] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0070] The processor 102 can be an integrated circuit chip with signal processing capabilities. The processor 102 can be a general-purpose processor 102, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0071] In the embodiments provided in this application, it should be understood that the disclosed methods and systems can also be implemented in other ways. The method and system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0072] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0073] On the other hand, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by processor 102, the computer program implements the methods described in any of the first aspects above. If the functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

Claims

1. A data augmentation method for class-imbalanced phytoplankton image classification, characterized in that, Includes the following steps: S1 Data Preparation and Preprocessing: Obtain the original image dataset of marine plankton, divide it into training set, validation set and test set, and label the plankton in it to obtain training set, validation set and test set with label information; S2 Constructing the Plankton-LDM network model and latent space mapping: Constructing a Plankton-LDM network model based on the latent diffusion model; the Plankton-LDM network model includes a pre-trained autoencoder for data space transformation and an improved noise prediction network U-Net for performing denoising tasks; S3 Model Training: The Plankton-LDM network model is iteratively trained using a training set with labeled information to obtain a trained Plankton-LDM network model; S4 generates high-quality marine plankton images: using a trained Plankton-LDM network model, combined with a second-order Heun sampling strategy, a reverse denoising process is performed to obtain a high-quality marine plankton image set; S5 constructs a class-balanced augmented training set and inputs it into a deep learning image classification model for iterative training.

2. The method according to claim 1, characterized in that, In step S1, the annotation specifically involves: labeling the plankton in the images on the training set, validation set, and test set with rectangular boxes, thereby obtaining the training set, validation set, and test set with annotation information.

3. The method according to claim 2, characterized in that, The labeling information includes the coordinates of four points (x1, y1), (x2, y2), (x3, y3), and (x4, y4) arranged clockwise within a rectangle, and the category ID of each plankton species.

4. The method according to claim 1, characterized in that, In step S2, the data space transformation specifically involves inputting the labeled training set into a pre-trained autoencoder. The labeled training set is a high-dimensional marine plankton image, and the encoder compresses it into a low-dimensional latent space representation.

5. The method according to claim 1, characterized in that, In step S2, the improved noise prediction network U-Net specifically involves: introducing a VSS module to replace the standard convolutional block in the second stage of the downsampling path of the noise prediction network U-Net; and introducing HDRAB to replace the traditional residual block in the third and fourth stages of the downsampling path of the noise prediction network U-Net.

6. The method according to claim 1, characterized in that, In step S3, the iterative training of the Plankton-LDM network model specifically involves adding noise to the potential representation through a forward diffusion process and predicting the noise through an improved noise prediction network U-Net to complete the reverse denoising learning, thereby obtaining the trained Plankton-LDM network model.

7. The method according to claim 1, characterized in that, Step S4 specifically involves: starting from random sampling noise that follows a standard normal distribution, a second-order Runge-Kutta method Heun sampler is used to perform inverse denoising in the latent space. Within each denoising time step, a dual gradient evaluation is performed for prediction and correction, once for predicting the next latent representation state and once for correction. A weighted average of the two gradients is then used to perform the final latent representation update. The Heun sampler accurately approximates the true ordinary differential equation denoising trajectory through a second-order precision integral step, obtaining a progressively restored clear latent space representation. This progressively restored clear latent space representation is then mapped back to the image pixel space through the decoder of the pre-trained autoencoder, thereby generating the high-quality marine plankton image set.

8. The method according to claim 1, characterized in that, The construction of the class-balanced augmented training set specifically involves: supplementing the training set with high-quality marine plankton images for categories whose sample numbers are less than those of other categories; and expanding the sample numbers of each category to be consistent to obtain the class-balanced augmented training set.

9. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the data augmentation method for classifying imbalanced planktonic images according to any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the data augmentation method for classifying imbalanced planktonic images as described in any one of claims 1-8.