A real-time underwater disease identification method for grouper based on PLDNET network
By constructing the PLDNET network, the problems of insufficient accuracy and poor real-time performance in the detection of diseases in East Star Grouper were solved, enabling accurate identification and efficient detection of small lesion areas and improving the efficiency of aquaculture management.
Patent Information
- Application Number
- CN202411678268.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-11-22
AI Technical Summary
Existing technologies for detecting diseases in East Star Grouper suffer from insufficient accuracy, poor real-time performance, and a lack of datasets, making it difficult to accurately identify and locate small lesion areas, especially in complex aquaculture environments, which makes it difficult to identify and manage early-stage diseases.
A real-time disease identification method for East Star Grouper based on the PLDNET network is constructed. The PLDD dataset is built through data augmentation and annotation. Multi-level deep convolution and focus modulation layer feature extraction are adopted, combined with an improved minimum point distance intersection-union loss function, to achieve accurate localization and classification of small target diseases.
It improved the accuracy of identifying minor diseases in the East Star Grouper, enhanced the model's generalization ability and real-time detection capability, significantly reduced the risk of early-stage diseases being overlooked, reduced economic losses, simplified the operation process, and improved detection efficiency.
Smart Images

Figure CN119559490B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aquaculture disease identification technology, and more specifically, to a real-time underwater disease identification method for grouper based on a PLDNET network. Background Technology
[0002] In aquaculture, early monitoring and accurate diagnosis of fish diseases are crucial for ensuring healthy fish populations and improving economic efficiency. Take the red grouper as an example; due to its high market and nutritional value, it is widely used in aquaculture. However, in high-density, mechanized farming environments, red grouper are highly susceptible to diseases such as vibriosis and parasitic infections, posing significant challenges to their management. Once a disease outbreak occurs, high mortality rates and a decline in market value can lead to severe economic losses.
[0003] In existing technologies, fish disease detection methods are mainly divided into traditional manual visual inspection and automated detection based on machine learning. Manual visual inspection relies on experts to diagnose by observation, which, while having some practical value, suffers from high subjectivity, time-consuming processes, low accuracy leading to misjudgments, and high labor costs, making it difficult to meet the needs of large-scale aquaculture. In recent years, with the continuous development of artificial intelligence, machine learning methods based on image processing technology have been gradually applied to fish disease detection. Automated detection methods based on machine learning improve the efficiency and accuracy of fish disease detection through steps such as image segmentation, feature extraction, and classification. However, automated detection methods based on machine learning still face some limitations: First, when dealing with smaller lesions such as parasitic diseases and early vibriosis, existing machine learning methods often lack sufficient detection accuracy, making it difficult to accurately identify and locate lesion areas, thus limiting the early identification capability of fish diseases. Second, due to the scarcity of disease image datasets for specific fish species, the generalization ability of existing models is limited, making it difficult to extend to different aquaculture environments. Finally, real-time performance remains a problem that urgently needs to be solved; most current methods focus on improving model detection accuracy, failing to meet the needs of rapid detection and real-time response in actual aquaculture environments. Therefore, in the aquaculture of East Star Grouper, given its small disease target and complex background,
[0004] First, existing machine learning-based automated detection technologies perform poorly in complex aquaculture environments (such as changes in light intensity and water quality fluctuations), struggling to accurately identify and locate small lesion areas, particularly in detecting early-stage vibriosis and parasitic diseases. This limitation leads to the failure to identify early-stage diseases, often allowing them to progress to later stages and ultimately causing greater economic losses to the aquaculture of red grouper. Second, due to the lack of high-quality disease image datasets specific to certain fish species (such as red grouper), existing target detection methods have significantly limited generalization capabilities, failing to effectively extend to specific red grouper aquaculture environments. This limitation creates a technological gap in the field of red grouper disease identification, necessitating the development of more adaptable detection solutions. Finally, while current technologies improve detection accuracy in some cases, they perform poorly in real-time performance, failing to respond promptly to actual aquaculture needs and thus delaying disease management. Summary of the Invention
[0005] The technical problem to be solved by this invention is to provide an underwater real-time disease identification method for red snapper based on PLDNET network. This method is efficient, accurate and has real-time detection capabilities to meet the needs of early disease detection and localization in red snapper, while accurately identifying the types of diseases, thereby effectively addressing the health management challenges in aquaculture.
[0006] The technical solution adopted by this invention to solve its technical problem is: to construct a real-time underwater disease identification method for grouper based on PLDNET network, including the following steps:
[0007] S1. Data Acquisition and Preprocessing: Acquire a dataset of disease images of the East Star Grouper fish and preprocess the data, including data augmentation, annotation, and dataset partitioning.
[0008] S2. Feature extraction: The preprocessed data of East Star Spot disease is input into the feature extraction module of the PLDD network model, and the feature extraction results are obtained after multi-level deep convolution.
[0009] S3. Feature enhancement and fusion: Input the data results of the extracted East Star Spot disease into the feature enhancement and fusion module. After the focus modulation layer and multi-scale feature fusion, the feature information of East Star Spot disease from different levels is integrated to enhance the performance of key information and obtain the enhanced fusion result.
[0010] S4. Change the loss function, replacing the traditional bounding box loss function with the improved minimum point distance intersection-union ratio loss function to accurately locate the small target area of the Eastern Star Spot disease.
[0011] S5. Disease prediction of East Star Grouper: The feature enhancement fusion result is input into the PLDNet feature prediction module for prediction, and multiple East Star Grouper disease prediction boxes and categories are output.
[0012] According to the above scheme, in step S1, data acquisition and preprocessing are achieved through the following method:
[0013] S101. Real-time acquisition of disease image datasets of East Star Grouper using underwater cameras in East Star Grouper fish farms;
[0014] S102. Perform various image enhancement operations on the images of East Star Spotted Fish Disease in the dataset, including mirroring, flipping, cropping, rotating, translating, blurring, adding noise, and adjusting brightness, to construct the PLDD East Star Spotted Fish Disease Dataset.
[0015] S103. Use Labelme annotation software to accurately annotate the target disease areas in the enhanced dataset images one by one, including manually drawing disease outlines using key point annotations and assigning corresponding disease category labels to each annotated area to ensure that each disease area can be accurately labeled.
[0016] S104. Divide the PLDD dataset into a training set, a test set, and a validation set in a ratio of 7:2:1. The training set is used for pre-training the PLDD model, the test set is used for evaluating the performance and tuning the parameters of the PLDD model, and the validation set is used for measuring the final performance of the model and evaluating its generalization ability.
[0017] According to the above scheme, in step S2, feature extraction is achieved through the following method:
[0018] S201. Input the preprocessed image data of the East Star Spot disease into the PLDNet feature extraction network. The feature extraction module consists of multiple convolutional layers and residual block layers to extract features at different levels in the East Star Spot image.
[0019] S202, Silence layer feature suppression: When the image of the Eastern Star Spot disease enters the feature extraction module, it first passes through the Silence layer, which is used to suppress certain noise or features unrelated to the disease in the input image of the Eastern Star Spot disease, and improve the effect of feature extraction in subsequent layers.
[0020] S203. Convolutional layer feature extraction: After the Silence layer, multiple deep convolution operations are used to extract various local features of disease edges, textures, and shapes layer by layer, gradually reducing the size of the East Star Spot disease image, increasing the depth of the feature map, and capturing more abstract features.
[0021] S204. Residual Block Layer Feature Optimization: In the feature extraction module, a residual block layer is attached after each convolutional layer to improve the nonlinear representation capability of the model. The residual block layer further extracts deep features of the East Star Spot disease image. Skip connections are used to reduce the gradient vanishing problem. Through multi-scale extraction of local features, the perception capability of the East Star Spot disease area is improved.
[0022] S205. Output feature map. Finally, output a high-dimensional feature map after processing with multiple convolutional layers and residual block layers.
[0023] According to the above scheme, in step S3, feature enhancement and fusion are achieved through the following method:
[0024] S301. Input the feature map output in step S2 into the feature fusion enhancement module and use the focus modulation layer mechanism to enhance the feature map of East Star Grouper disease. The focus modulation layer is divided into three parts: Query, Modulator, and Aggregate.
[0025] S302, Feature Fusion and Upsampling: Through multi-scale feature fusion and upsampling operations, feature maps at different levels are integrated and spatial resolution is gradually restored to ensure that the PLDNet network can process large, medium and small disease areas simultaneously.
[0026] S303. Output the enhanced fusion feature map after feature enhancement and feature fusion processing by the focus modulation layer, providing optimized input for subsequent prediction of disease areas in the East Star Grouper.
[0027] According to the above scheme, in step S301, the feature fusion enhancement module uses a focus modulation layer mechanism to enhance the features of the output East Star Grouper disease feature map, specifically as follows:
[0028] S301a. The feature map is input into the Query part. The Query part extracts key region features q(X) based on the input feature map X, which are used for subsequent focus modulation. The formula is as follows:
[0029] q(X)=f q (X)
[0030] Among them, f q (·) is a mapping function used to obtain key features of the input image;
[0031] S301b: The Modulator component aggregates and adaptively adjusts features at different levels. First, it extracts multi-level contextual information to form a feature set Z, where each level's feature is represented as Z. l Aggregate and weight them through linear transformation:
[0032] Zl =GeLU(Conv dw (Z l-1 ))
[0033] The final modulator output is obtained by weighted summation:
[0034]
[0035] Among them, G l Here, L represents the weights, and L represents the number of feature layers. This indicates element-wise multiplication, and the final output Z out The feature map after adaptive aggregation;
[0036] S301c, In the Aggregate section, the features output by the modulator are combined with the Query features through element-wise affine transformation to further enhance the salience of the diseased area:
[0037]
[0038] Among them, g i l z represents the weight. i l For the features of the l-th layer, h(·) is a linear transformation function, and y i This is the final enhanced feature.
[0039] According to the above scheme, the method for replacing the traditional bounding box loss function with the improved minimum point distance intersection-union ratio (MPDIoU) loss function in step S4 is as follows:
[0040] First, define the coordinates of the top-left and bottom-right corners of the ground truth bounding box and the predicted bounding box. Calculate the Euclidean distance between the corresponding vertices of these two bounding boxes. Simultaneously, calculate the overlap area and union area of the predicted and ground truth bounding boxes. Then, normalize the distance between vertices and divide it by the sum of the squares of the width and height of the bounding boxes. The formula for calculating MPDIoU is:
[0041]
[0042] The overlap area between the predicted bounding box and the ground truth bounding box is A. gt ∩A pre and the area of the union A gt ∪A pre Define the real bounding box A gt The coordinates of the top left corner are (x 1,gt ,y 1,gt The coordinates of the lower right corner point are (x 2,gt ,y 2,gt Similarly, prediction box A pre The coordinates of the top left corner are (x1,pre ,y 1,pre The coordinates of the lower right corner point are (x 2,pre ,y 2,pre For the corresponding vertices of the two bounding boxes, calculate the Euclidean distance between the top-left corner and the bottom-right corner respectively:
[0043] d1 2 =(x 1,pre -x 1,gt ) 2 +(y 1,pre -y 1,gt ) 2
[0044] d2 2 =(x 2,pre -x 2,gt ) 2 +(y 2,pre -y 2,gt ) 2
[0045] w and h are the width and height of the predicted bounding box or the ground truth bounding box, respectively;
[0046] MPDIoU loss function L MDPIoU The objective is to minimize the distance between the predicted bounding box and the ground truth bounding box, while maximizing their overlap area. The loss function is expressed as:
[0047] L MDPIoU =1-MDPIoU.
[0048] According to the above scheme, in step S5, the prediction of diseases in the grouper is achieved through the following method:
[0049] S501. Input the feature map after feature enhancement and fusion into the PLDNet feature prediction module to perform the localization and classification of disease areas in East Star Grouper.
[0050] The S502 and PLDNet feature prediction modules generate prediction boxes at three different scales based on the input feature map, namely P3 / 8, P4 / 16, and P5 / 32. Each prediction box is used to detect disease areas of small, medium, and large-sized grouper, ensuring that all types of disease areas can be accurately located.
[0051] S503. For each predicted box, the model outputs the location coordinates through the box regression head and assigns the predicted box a disease category label for East Star Grouper and a confidence level in the range of 0-1 through the cls classification head.
[0052] S504. The generated prediction boxes are optimized and filtered using the non-maximum suppression method to remove redundant and low-confidence prediction boxes, thereby obtaining high-confidence detection results for disease areas in East Star Grouper.
[0053] The underwater real-time disease identification method for grouper based on PLDNET network of the present invention has the following beneficial effects:
[0054] 1. The PLDNet (Plectropomus Leopardus Disease Detection Network) proposed in this invention enhances the recognition accuracy of early small lesion areas by introducing a FocalModulation layer, and improves the detection accuracy of lesion areas by optimizing the loss function through an improved Minimum Point Distance Intersection over Union (MPDIoU). It provides a new PLDD (Plectropomus Leopardus Disease Dataset) dataset, filling the gap in Plectropomus Leopardus disease image data, effectively improving the generalization ability of the model, and making up for the technical gap in the field of deep learning-based Plectropomus Leopardus disease recognition. The proposed PLDNet network architecture enables real-time detection of Plectropomus Leopardus diseases, can quickly respond to the occurrence of diseases, and meets the needs of efficient and accurate disease detection in aquaculture.
[0055] 2. This invention can improve the accuracy of identifying small disease areas in red grouper. Addressing the challenges of light variations and water quality fluctuations in complex aquaculture environments, PLDNet significantly improves the accuracy of identifying small lesion areas in red grouper by introducing multi-scale feature fusion technology and an improved loss function (MPDIoU). Experimental results show that PLDNet achieves an accuracy of 88.1% on the mAP@0.5 index, which is more than 8% higher than existing advanced target detection models. Especially in the detection of small lesions such as early-stage vibriosis and parasitic diseases, PLDNet exhibits higher stability and reliability, significantly reducing the risk of early-stage diseases being overlooked and effectively reducing economic losses in red grouper due to delayed diagnosis.
[0056] 3. This invention improves real-time performance and reduces the computational complexity of the model by optimizing the PLDNet network structure and algorithm, thereby significantly increasing the detection speed. In experiments, PLDNet can achieve efficient real-time detection in practical application scenarios, with a response speed that is about 15% faster than existing technologies. This rapid response capability not only meets the real-time disease monitoring needs of fish farmers, but also improves the efficiency of disease management for grouper and reduces further spread and losses caused by delayed detection.
[0057] 4. This invention enhances the generalization ability by constructing a high-quality image dataset of diseases in red snapper (PLDD), making up for the shortcomings of existing technologies in terms of dataset scarcity. The introduction of the PLDD dataset enables the PLDNet model to show stronger generalization ability in different aquaculture environments. In particular, PLDNet can maintain detection stability when facing different lighting conditions and complex water quality backgrounds. Compared with traditional models, PLDNet improves the recognition accuracy in different environments by about 10%, filling the technical gap in the detection of diseases in red snapper aquaculture.
[0058] 5. The detection method provided by this invention simplifies the operation process, is easy to operate and deploy, and reduces the reliance on professional knowledge, enabling farmers and technicians to apply it easily. Compared with existing technologies, PLDNet reduces the need for manual intervention through a highly automated real-time detection process, significantly improving the ease of operation and promoting the popularization and application of this technology in large-scale farms. This simplified process can help farmers make diagnostic decisions quickly, thereby improving the efficiency of the entire farming management. Attached Figure Description
[0059] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0060] Figure 1 This is a flowchart of the underwater real-time disease identification method for grouper based on PLDNET network according to the present invention;
[0061] Figure 2 This is an example image of the PLDD dataset of diseases in the Eastern Starfish of this invention;
[0062] Figure 3 This is an example diagram of the data enhancement for diseases in the East Star Grouper fish according to the present invention;
[0063] Figure 4 This is an example diagram of the PLDNet network structure of the present invention;
[0064] Figure 5 This is an example diagram of the Focal Modulation module structure of the present invention;
[0065] Figure 6 This is an example diagram illustrating the principle of MPDIoU (Intersection over Union Ratio Based on Minimum Point Distance) of the present invention;
[0066] Figure 7 This is an example diagram of the detection of grouper in the East Star Fish according to the present invention. Detailed Implementation
[0067] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0068] like Figure 1-7 As shown, the underwater real-time disease identification method for grouper based on PLDNET network of the present invention includes the following steps:
[0069] S1. Data Acquisition and Preprocessing: Acquire a dataset of disease images of the East Star Grouper fish and preprocess the data, including data augmentation, annotation, and dataset partitioning.
[0070] The following steps are specifically included in the example of Vibrio serrata infection:
[0071] S101. Real-time acquisition of disease image datasets of grouper from underwater cameras in the Dongxing Grouper fishery.
[0072] Images of Vibrio infection in grouper were captured under natural aquaculture conditions at a grouper farm. Both natural and artificial light sources were used to capture images of the disease under varying lighting conditions. Simultaneously, diseased grouper were meticulously observed and recorded, particularly focusing on specific locations and stages of Vibrio infection. During the collection process, each grouper was photographed from multiple angles—front, back, and side—to record the overall appearance of the Vibrio-infected area. Figure 2 As shown, the PLDD image dataset for diseases of the East Star Grouper contains 1041 images, mainly divided into three categories: normal East Star Grouper (a), Vibrio infection (b), and leech infection (c).
[0073] S102. Perform various image enhancement operations on the images of spotted grouper disease in the dataset, including mirroring, flipping, cropping, rotating, translating, blurring, adding noise, and adjusting brightness. For each image, three enhancement operations are randomly combined to generate three images with different enhancement operations for each original spotted grouper disease image, thus constructing the PLDD spotted grouper disease dataset.
[0074] Taking an original image of a spotted variegated ... Figure 3As shown, here are examples of data augmentation for diseases in the grouper: (a) is the original image of the grouper; (b) is the data-augmented image of (a) generated after flipping, adding noise, and cropping; (c) is the data-augmented image of (a) generated after flipping, translating, and adjusting brightness; and (d) is the data-augmented image of (a) generated after adjusting brightness, adding noise, and blurring.
[0075] S103. Use Labelme annotation software to accurately annotate each target disease area in the enhanced dataset image. This step includes manually drawing the disease outline using key point annotation and assigning a corresponding disease category label to each annotated area to ensure that each disease area can be accurately labeled.
[0076] For each image of Vibrio tumefaciens infection, keypoint annotation was used to manually draw the outline of the Vibrio lesion area, ensuring that all Vibrio lesions were accurately marked. Each lesion area was also assigned a corresponding category label, such as "Vibrio infection." Through this precise annotation method, the model can learn the different disease characteristics of the red grouper during training, thereby improving the model's detection performance for various fish diseases.
[0077] S104. Divide the PLDD dataset into a training set, a test set, and a validation set in a 7:2:1 ratio. The training set is used for pre-training the PLDD model, the test set is used for evaluating the performance and tuning the parameters of the PLDD model, and the validation set is used for measuring the final performance and evaluating the generalization ability of the model.
[0078] After image annotation of the PLDD (Polygonum aviculare) dataset, the entire dataset of images of Polygonum aviculare disease is divided into training, testing, and validation sets in a 7:2:1 ratio. For example, assuming there are 1000 annotated images of Vibrio aviculare infection, 700 images are used as the training set for the PLDD model, 200 images are used as the testing set to evaluate the initial performance of the PLDD model, and 100 images are used as the validation set to finally verify the generalization ability of the PLDD model. This reasonable data partitioning ensures that the sample distribution of each dataset is consistent, avoiding overfitting or underfitting of the model, thereby improving the model's prediction accuracy for unknown data.
[0079] S2. Feature Extraction: The preprocessed data on East Star Spot disease is input into the feature extraction module of the PLDD network model. After multi-level deep convolution, the feature extraction results are obtained. Specifically:
[0080] S201. Input the preprocessed image data of the Eastern Star Spot disease into the PLDNet feature extraction network. The feature extraction module consists of multiple convolutional layers (Conv) and residual block layers (ResNCSPELAN4) to extract features at different levels from the Eastern Star Spot image.
[0081] Image data of Vibrio infection in grouper fish, after data augmentation and standardization, is input into the feature extraction module. The PLDNet network consists of multiple convolutional layers (Conv) and residual block layers (ResNCSPELAN4). The PLDNet network structure is as follows: Figure 3 As shown, the convolutional layers first extract low-level features from the image, such as identifying the edges, textures, and basic shapes of Vibrio spp. infection. Taking Vibrio spp. infection as an example, after the input image is processed by the initial convolutional layers, the model can identify the Vibrio infection area on the fish body. These low-level features provide the foundation for subsequent deep analysis. The residual block layers, through skip connections, allow gradients to propagate deeper, preventing gradient vanishing in deep networks. Furthermore, the residual blocks further extract high-level features from the image, such as morphological changes and subtle pathological features of the diseased area. Taking leech infection as an example, the model, through this network structure, can identify tiny lesion areas on the fish body, especially in early symptom presentation.
[0082] S202, Silence layer feature suppression: When the image of the Eastern Star Spot disease enters the feature extraction module, it first passes through the Silence layer, which is used to suppress certain noise or features unrelated to the disease in the input Eastern Star Spot disease image, aiming to improve the effect of feature extraction in subsequent layers.
[0083] For Vibrio edulis infection of the red snapper, the Silence layer first performs preliminary processing on the input image of the red snapper to suppress noise and features unrelated to the disease. Using a high-pass filtering noise reduction algorithm, the Silence layer identifies and filters out elements that interfere with disease detection, such as light spots and suspended particles in the water, thus avoiding impact on subsequent detection. Furthermore, the Silence layer weakens features in areas unrelated to the small target disease in the red snapper (such as the scales of healthy fish or uninfected parts of the body), reducing interference from non-lesion areas. After these processing steps, the Silence layer enhances disease-related features, especially the small lesions caused by Vibrio edulis infection, to improve the feature extraction accuracy of subsequent convolutional layers. Finally, the Silence layer outputs an optimized feature map, providing a more accurate and effective foundation for subsequent detection.
[0084] S203. Convolutional layer feature extraction: After the Silence layer, multiple deep convolution (Conv) operations are used to extract various local features such as disease edges, textures, and shapes layer by layer, gradually reducing the size of the East Star Spot disease image while increasing the depth of the feature map, aiming to capture more abstract features.
[0085] During the convolutional layer feature extraction process, multiple convolutional layers are used to progressively extract features from the input image, with each layer capturing increasingly complex local features. The convolutional kernel moves across the Dongxingzhao image to detect edges, corners, and local texture features.
[0086] Taking Vibrio infection in the red grouper as an example, the convolutional kernel slides across the image, gradually extracting features of Vibrio infection through layer-by-layer convolution operations. The initial convolutional layer may only identify the edges or scale structure of the red grouper's body, establishing a basic geometric understanding for the network. In deeper convolutional layers, the kernel can focus on minute textures and color changes on the red grouper's scales, capturing local lesion features related to Vibrio infection, such as white spots and ulcers caused by Vibrio infection.
[0087] When dealing with vibrio infections, the network not only needs to identify the lesion area but also needs to remain robust under complex backgrounds and varying lighting conditions. Through progressive extraction via layers of convolution, the network can identify subtle differences in lesion areas and even the lesion's appearance under different lighting conditions. Deeper convolutional layers can capture higher-dimensional features, such as the localized abnormal spread and irregular morphological changes caused by vibrio infection, or the unique white spot morphology on spotted scales. This multi-layered feature extraction allows the model to distinguish subtle lesions even in complex backgrounds, improving the network's detection accuracy and stability in different environments.
[0088] S204. Residual Block Layer Feature Optimization: In the feature extraction module, a residual block layer is appended after each convolutional layer to enhance the model's nonlinear representation capability. The residual block layer (ResNCSPELAN4) further extracts deep features from the *Symplocos edulis* (Eastern Star Spot) disease image, using skip connections to reduce the gradient vanishing problem. Multi-scale extraction of local features aims to further improve the perception capability of the *Symplocos edulis* disease region.
[0089] The features extracted by the convolutional layers are optimized using ResNCSPELAN4 residual block layers. The core function of the residual block layer is to capture leech spot disease regions of different sizes and shapes through a multi-scale feature extraction mechanism. For example, the symptoms of leech infection manifest as extremely small lesion areas, which may be difficult to identify using traditional convolutional layers. However, the residual block layer, through deeper feature extraction and skip connections, ensures that the salient features of leech disease regions are fully learned, especially when the lesion area is small. This optimization improves the model's ability to detect small target diseases. In addition, by integrating information from different levels, the residual block ensures that no detail is lost in each convolutional feature layer, ultimately improving the model's stability and accuracy.
[0090] S205. Output Feature Maps: Finally, output high-dimensional feature maps processed by multi-layer convolution and residual block layers. These feature maps possess rich spatial and contextual information, aiming to provide a foundation for subsequent feature enhancement and disease prediction.
[0091] After multi-layer convolution and residual block processing, the network outputs high-dimensional feature maps, which contain rich spatial details and contextual information from the image. For example, in detecting disease in the spotted leech, the output high-dimensional feature maps can capture not only the main diseased areas on the fish's surface but also the surrounding healthy tissue, providing rich input information for subsequent feature enhancement and disease localization. Through multi-layer information fusion, these feature maps ensure that the model can accurately locate and identify diseased areas even under complex background or lighting conditions. Finally, these high-dimensional feature maps are passed to the feature enhancement module to further improve the accuracy of disease detection.
[0092] S3. Feature Enhancement and Fusion: The extracted data on East Asian star spot disease is input into the feature enhancement and fusion module. After focal modulation and multi-scale feature fusion, the feature information from different levels of East Asian star spot disease is integrated to enhance the representation of key information, resulting in an enhanced fusion result. Specifically:
[0093] S301. Input the feature map output in step S2 into the feature fusion and enhancement module to enhance the feature map of the East Star Grouper disease using the Focal Modulation mechanism. The Focal Modulation layer mainly consists of three parts: Query, Modulator, and Aggregate. Based on the importance of the diseased area, the Focal Modulation layer adaptively adjusts the feature weights, suppresses background noise, enhances the salience of the diseased area, and focuses on specific East Star Grouper image regions to process local details, especially small target lesions that are easily overlooked.
[0094] In an image of a grouper scale potentially infected with Vibrio, the focus modulation layer assigns higher weight to the infected area and lowers the weight to healthy tissue or the underwater background. For example, early symptoms of leech infection appear as tiny leech-infested areas on the fins; through focus modulation, these minute lesions can be enhanced in the feature map, ensuring that the model focuses on these small lesions even with a large field of view. Figure 5 The diagram shows an example of the FocalModulaiton module structure.
[0095] In the feature fusion enhancement module, the output feature map of the East Star Grouper is enhanced using a focal modulation layer mechanism. First, the feature map is input into the Query part, which extracts key region features q(X) from the input feature map X for subsequent focal modulation. The formula can be expressed as:
[0096] q(X)=f q (X)
[0097] Among them, f q (·) is a mapping function used to obtain key features of the input image.
[0098] Next, the Modulator part aggregates and adaptively adjusts features at different levels. First, multi-level contextual information is extracted to form a feature set Z, where each level's feature is represented as Zi. l Aggregate and weight them through linear transformation:
[0099] Z l =GeLU(Conv dw (Z l-1 ))
[0100] The final modulator output is obtained by weighted summation:
[0101]
[0102] Among them, G l Here, L represents the weights, and L represents the number of feature layers. This indicates element-wise multiplication. The final output Z out It is the feature map after adaptive aggregation.
[0103] Finally, in the Aggregate section, the features output by the modulator are combined with the Query features through element-wise affine transformation to further enhance the saliency of the diseased areas:
[0104]
[0105] Among them, g i l z represents the weight. i l For the features of the l-th layer, h(·) is a linear transformation function, and y i This is the final enhanced feature.
[0106] S302, Feature Fusion and Upsampling: Through multi-scale feature fusion and upsampling operations, feature maps at different levels are integrated to gradually restore spatial resolution. This aims to ensure that the PLDNet network can simultaneously handle large, medium, and small disease regions, improving multi-scale detection capabilities.
[0107] Taking leech infection and vibrio infection as examples, the differences in their pathological characteristics are quite significant. Leech infection typically presents as small and scattered lesions, with a small lesion area, making it difficult to accurately capture using conventional feature extraction methods. Vibrio infection, on the other hand, usually covers a larger area, exhibiting more spatially coherent large lesions. In the process of multi-scale feature fusion, the PLDNet model utilizes multi-layer convolutional neural networks to extract feature information at different scales, and combines a focus modulation mechanism to adaptively enhance features and adjust weights in important regions.
[0108] In the initial stage, the model extracts detailed features of small lesions caused by leech infection through deep convolution. These small lesions may initially appear as subtle local differences in the feature map; through layer-by-layer convolution, these minute lesions are gradually enhanced, thus capturing the edges and local details of the leech infection. For these small lesion areas, the model uses layer-by-layer feature extraction and downsampling operations to ensure that fine-grained features are not overwhelmed in the convolutional neural network. Simultaneously, a focus modulation mechanism adaptively suppresses background noise, further improving the detection capability of small lesions.
[0109] When dealing with Vibrio infections, the lesions are often large. Therefore, the PLDNet model, during feature extraction, not only needs to extract local details but also needs to restore higher spatial resolution through upsampling to ensure accurate capture of the overall lesion morphology. Upsampling, combined with feature information at different scales, can gradually reconstruct large-area Vibrio lesions. Furthermore, feature fusion organically combines low-level edge information with high-level contextual information, ensuring a complete display of the lesion region. In this process, the focus modulation module adaptively adjusts feature weights based on the importance of the lesion area, enhancing lesion saliency and suppressing redundant background, thereby effectively improving the detection performance of large-area lesions.
[0110] Through this fusion and enhancement of multi-scale features, the PLDNet model can accurately capture disease features of different sizes and shapes, whether it's a localized, minute leech infection lesion or a large-area Vibrio infection lesion. This is achieved by adaptively adjusting the receptive field and feature fusion process. This process not only improves the model's robustness in multi-scale disease detection but also enhances its generalization ability in complex environments, enabling it to cope with spatial scale variations in various lesion features and ensuring the accuracy and reliability of detection results.
[0111] S303. Output the enhanced fusion feature map after feature enhancement and feature fusion processing by the focus modulation layer, which aims to provide optimized input for subsequent prediction of disease areas in the East Star Grouper.
[0112] In an image of a red snapper containing multiple diseases, the enhanced feature map can clearly distinguish different diseased areas, such as vibrio infection on the scales and leech infection lesions on the fins. This not only allows the model to more accurately identify the type of each disease but also provides more reliable boundary information for locating diseased areas in the red snapper. Through this process, the model can reduce false alarm rates (such as misclassifying healthy red snapper areas as leech-infected areas) and improve the overall accuracy of disease detection.
[0113] S4. Change the loss function. Replace the traditional bounding box loss function with the improved minimum point distance intersection-union ratio (MPDIoU) loss function to accurately locate the small target area of the Eastern Star Spot disease.
[0114] Traditional IoU-type loss functions often fail to accurately locate diseases in small target areas, such as leech infestations, because the overlap between bounding boxes is small, especially when the lesion area has a complex shape or inaccurate location. MPDIoU, by introducing the concept of minimum point distance, not only considers the overlap area of the bounding boxes but also calculates the minimum distance between corresponding points between the predicted and ground truth boxes, thus providing a more accurate measure for fine-grained feature localization.
[0115] In the specific implementation, firstly, the coordinates of the top-left and bottom-right corners of the ground truth bounding box and the predicted bounding box are defined, and the Euclidean distance between the corresponding vertices of these two bounding boxes is calculated. Simultaneously, the overlap area and union area of the predicted and ground truth bounding boxes are calculated, and the distance between the vertices is normalized and divided by the sum of the squares of the width and height of the bounding boxes. The formula for calculating MPDIoU is:
[0116]
[0117] The overlap area between the predicted bounding box and the ground truth bounding box is A. gt ∩A pre and the area of the union A gt ∪A preDefine the real bounding box A gt The coordinates of the top left corner are (x 1,gt ,y 1,gt The coordinates of the lower right corner point are (x 2,gt ,y 2,gt Similarly, prediction box A pre The coordinates of the top left corner are (x 1,pre ,y 1,pre The coordinates of the lower right corner point are (x 2,pre ,y 2,pre For corresponding vertices of the two bounding boxes, calculate the Euclidean distance between the top-left corner and the bottom-right corner respectively:
[0118] d1 2 =(x 1,pre -x 1,gt ) 2 +(y 1,pre -y 1,gt ) 2
[0119] d2 2 =(x 2,pre -x 2,gt ) 2 +(y 2,pre -y 2,gt ) 2
[0120] w and h are the width and height of the predicted bounding box or the ground truth bounding box, respectively.
[0121] MPDIoU loss function L MDPIoU The goal is to minimize the distance between the predicted bounding box and the ground truth bounding box while maximizing their overlap area. The specific loss function expression is:
[0122] L MDPIoU =1-MDPIoU
[0123] By minimizing the MPDIoU loss function during training, the PLDNet model can more accurately locate small target disease areas in the red grouper, including subtle lesion features such as lesions caused by leech or vibriosis. This improvement not only optimizes the bounding box localization of disease areas in the red grouper but also enhances the robustness of the PLDNet model in handling small target diseases and diseases in complex conditions, thereby improving the detection performance of the PLDNet model.
[0124] S5. Disease prediction for the red snapper: The feature enhancement fusion results are input into the PLDNet feature prediction module for prediction, outputting multiple red snapper disease prediction boxes and categories. Specifically:
[0125] S501. Input the feature map after feature enhancement and fusion into the PLDNet feature prediction module to perform the localization and classification of disease areas in the East Star Grouper.
[0126] When detecting leech infection in grouper, after focus modulation and feature fusion processing, the feature map highlights the diseased areas and suppresses background noise. At this point, the important disease information in the feature map is transmitted to the feature prediction module, which will further process and analyze these diseased areas to generate predicted bounding boxes for diseased areas and classify them.
[0127] The S502 and PLDNet feature prediction modules generate prediction boxes at three different scales (P3 / 8, P4 / 16, P5 / 32) based on the input feature map. Each prediction box is used to detect disease areas in small, medium, and large-sized grouper, aiming to ensure that all types of disease areas can be accurately located.
[0128] The PLDNet feature prediction network generates prediction boxes at different scales (P3 / 8, P4 / 16, P5 / 32) based on the input feature map. Each prediction box corresponds to a different size of diseased area in the red snapper. P3 / 8 is used to detect smaller diseased areas, such as early lesions of white spot disease in red snapper, which are usually small and difficult to detect. P4 / 16 and P5 / 32 are used to handle medium and larger diseased areas, such as extensive infection areas of vibriosis, respectively. Through this multi-scale prediction, the model ensures that both small-scale white spot disease and large-area vibriosis can be accurately detected and corresponding prediction boxes can be generated.
[0129] S503. For each predicted box, the model outputs the location coordinates through the box (regression head) and assigns a disease category label for the grouper and a confidence level in the range of 0-1 through the cls (classification head). This aims to distinguish different types of grouper diseases and determine their confidence levels.
[0130] To improve the localization accuracy of the predicted bounding boxes, the model uses MPDIoU (Intersection over Union based on minimum vertex distance) as a metric for the similarity between the predicted and ground truth bounding boxes. Unlike the traditional IoU (Intersection over Union) method, MPDIoU calculates the minimum distance between the vertices of the predicted and ground truth bounding boxes, enabling precise alignment even for small-scale disease areas (such as lesions in early-stage Vibrio infection). This approach avoids the bias of traditional IoU when dealing with small targets, and significantly improves the accuracy of disease area localization, especially in the detection of diseases in grouper.
[0131] The purpose of MPDIoU is to optimize the regression loss function of the model, ensuring that the output predicted bounding box not only has a high overlap with the ground truth bounding box, but also minimizes the shape and positional differences between them. For example, in detecting Vibrio infection, traditional IoU may lead to inaccurate localization due to the irregular shape of the disease area, while MPDIoU can accurately capture the boundaries of these irregular shapes, allowing the predicted bounding box to be better aligned with the actual disease area. The principle of MPDIoU is as follows: Figure 5 As shown.
[0132] S504. The generated prediction boxes are optimized and filtered using the non-maximum suppression (NMS) method to remove redundant and low-confidence prediction boxes, thereby finally obtaining a set of high-confidence detection results for the disease areas of the East Star Grouper.
[0133] Some predicted bounding boxes may overlap due to factors such as lighting and angle. The NMS algorithm retains the box with the highest confidence while removing other redundant boxes in overlapping areas to reduce the problem of multiple detections. Furthermore, predicted boxes with low confidence (e.g., below 50%) are filtered out to ensure that the final output contains a small number of high-confidence disease area predictions. For example, after NMS processing, the leech-infected area on the fin of a red snapper may only have one high-precision predicted bounding box with a confidence level of 95%, thus providing a reliable basis for subsequent decision-making. Figure 5 The image shows an example of the predicted results for the detection of grouper.
[0134] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A real-time underwater disease identification method for grouper based on PLDNET network, characterized in that, Includes the following steps: S1. Collect a dataset of images of diseases in the East Star Grouper fish, and preprocess the data, including data augmentation, annotation, and dataset partitioning; Data acquisition and preprocessing are achieved through the following methods: S101. Real-time acquisition of disease image datasets of East Star Grouper using underwater cameras in East Star Grouper fish farms; S102. Perform various image enhancement operations on the images of East Star Spotted Fish Disease in the dataset, including mirroring, flipping, cropping, rotating, translating, blurring, adding noise, and adjusting brightness, to construct the PLDD East Star Spotted Fish Disease Dataset. S103. Use Labelme annotation software to accurately annotate the target disease areas in the enhanced dataset images one by one, including manually drawing disease outlines using key point annotations and assigning corresponding disease category labels to each annotated area to ensure that each disease area can be accurately labeled. S104. Divide the PLDD dataset into training set, test set and validation set in a ratio of 7:2:
1. The training set is used for pre-training of the PLDD model, the test set is used for evaluating the performance and tuning of the PLDD model, and the validation set is used for measuring the final performance of the model and evaluating its generalization ability. S2. Input the preprocessed data of East Star Spot disease into the feature extraction module of the PLDD network model, and obtain the feature extraction results after multi-level deep convolution. Feature extraction is achieved through the following method: S201. Input the preprocessed image data of the East Star Spot disease into the PLDNet feature extraction network. The feature extraction module consists of multiple convolutional layers and residual block layers to extract features at different levels in the East Star Spot image. S202, Silence layer feature suppression: When the image of the Eastern Star Spot disease enters the feature extraction module, it first passes through the Silence layer, which is used to suppress certain noise or features unrelated to the disease in the input image of the Eastern Star Spot disease, and improve the effect of feature extraction in subsequent layers. S203. Convolutional layer feature extraction: After the Silence layer, multiple deep convolution operations are used to extract various local features of disease edges, textures, and shapes layer by layer, gradually reducing the size of the East Star Spot disease image, increasing the depth of the feature map, and capturing more abstract features. S204. Residual Block Layer Feature Optimization: In the feature extraction module, a residual block layer is attached after each convolutional layer to improve the nonlinear representation capability of the model. The residual block layer further extracts deep features of the East Star Spot disease image. Skip connections are used to reduce the gradient vanishing problem. Through multi-scale extraction of local features, the perception capability of the East Star Spot disease area is improved. S205. Output feature map. Finally, output a high-dimensional feature map after processing with multiple convolutional layers and residual block layers. S3. Input the extracted data of East Star Spot disease into the feature enhancement and fusion module. After the focus modulation layer and multi-scale feature fusion, the feature information of East Star Spot disease from different levels is integrated to enhance the performance of key information and obtain the enhanced fusion result. Feature enhancement and fusion are achieved through the following methods: S301. Input the feature map output in step S2 into the feature fusion enhancement module and use the focus modulation layer mechanism to enhance the feature map of East Star Grouper disease. The focus modulation layer is divided into three parts: Query, Modulator, and Aggregate. In the feature fusion enhancement module, the output East Star Spotted Fish Disease Feature Map is enhanced using a focus modulation layer mechanism, specifically as follows: S301a. Input the feature map into the Query part. The Query part then processes the input feature map. Extracting key region features This is used for subsequent focus modulation, and its formula is expressed as: = in, It is a mapping function used to obtain key features of the input image; S301b: The Modulator component aggregates and adaptively adjusts features at different levels. First, it extracts multi-level contextual information to form a feature set. The features of each level are represented as Aggregate and weight them through linear transformation: The final modulator output is obtained by weighted summation: in, As weight, For the number of feature layers, This represents element-wise multiplication, and the final output is... The feature map after adaptive aggregation; S301c, In the Aggregate section, the features output by the modulator are combined with the Query features through element-wise affine transformation to further enhance the saliency of the diseased area: in, Indicates weight, For the first Features of the layer It is a linear transformation function. This is the final enhancement feature; S302, Feature Fusion and Upsampling: Through multi-scale feature fusion and upsampling operations, feature maps at different levels are integrated and spatial resolution is gradually restored to ensure that the PLDNet network can process large, medium and small disease areas simultaneously. S303. Output the enhanced fusion feature map after feature enhancement and feature fusion processing by the focus modulation layer, providing optimized input for subsequent prediction of disease areas in East Star Grouper. S4. Replace the traditional bounding box loss function with an improved minimum point distance intersection-union ratio loss function to accurately locate small target areas of East Star Spot disease. S5. Disease prediction of East Star Grouper: The feature enhancement fusion result is input into the PLDNet feature prediction module for prediction, and multiple East Star Grouper disease prediction boxes and categories are output.
2. The underwater real-time disease identification method for grouper based on PLDNET network according to claim 1, characterized in that, In step S4, the method for replacing the traditional bounding box loss function with the improved minimum point distance intersection-union ratio (MPDIoU) loss function is as follows: First, define the coordinates of the top-left and bottom-right corners of the ground truth bounding box and the predicted bounding box. Calculate the Euclidean distance between the corresponding vertices of these two bounding boxes. Simultaneously, calculate the overlap area and union area of the predicted and ground truth bounding boxes. Then, normalize the distance between vertices and divide it by the sum of the squares of the width and height of the bounding boxes. The formula for calculating MPDIoU is: The overlap area between the predicted bounding box and the ground truth bounding box is... and the area of union Define the real bounding box The coordinates of the top left corner are The coordinates of the bottom right corner are Similarly, prediction boxes The coordinates of the top left corner are The coordinates of the bottom right corner are For the corresponding vertices of the two bounding boxes, calculate the Euclidean distance between the top-left corner and the bottom-right corner respectively: These are the width and height of the predicted bounding box or the ground truth bounding box, respectively; MPDIoU Loss Function The objective is to minimize the distance between the predicted bounding box and the ground truth bounding box, while maximizing their overlap area. The loss function is expressed as: 。 3. The underwater real-time disease identification method for grouper based on PLDNET network according to claim 1, characterized in that, In step S5, the prediction of diseases in the grouper is achieved through the following method: S501. Input the feature map after feature enhancement and fusion into the PLDNet feature prediction module to perform the localization and classification of disease areas in East Star Grouper. The S502 and PLDNet feature prediction modules generate prediction boxes at three different scales based on the input feature map, namely P3 / 8, P4 / 16, and P5 / 32. Each prediction box is used to detect disease areas of small, medium, and large-sized grouper, ensuring that all types of disease areas can be accurately located. S503. For each predicted box, the model outputs the location coordinates through the box regression head and assigns the predicted box a disease category label for East Star Grouper and a confidence level in the range of 0-1 through the cls classification head. S504. The generated prediction boxes are optimized and filtered using the non-maximum suppression method to remove redundant and low-confidence prediction boxes, thereby obtaining high-confidence detection results for disease areas in East Star Grouper.
Citation Information
Patent Citations
Medical image segmentation method based on cross-enhanced self-attention
CN117495872A
Real-time measurement method for heartbeat of underwater living exopalaemon carinicauda and electronic equipment
CN118633920A