Fish disease identification method based on deep learning
By improving the YOLOv11 network structure, the multi-branch adaptive reparameter module CKDB, multi-scale adaptive feature pyramid network N-MAFPN and multi-scale enhanced attention module SM-Detect have been introduced, which has solved the problem of insufficient recognition of fish diseases, achieved efficient and accurate fish disease detection, and promoted the development of intelligent health monitoring of fish diseases.
Patent Information
- Application Number
- CN202510527760.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
In intensive aquaculture environment, the lack of identification of fish diseases makes it difficult to ensure the timeliness and accuracy of disease detection. Traditional methods rely on manual observation and are highly subjective, making it difficult to meet the needs of efficient and accurate automation.
Using a fish disease recognition method based on deep learning, by improving the YOLOv11 network structure, a multi-branch adaptive reparameter module CKDB, a multi-scale adaptive feature pyramid network N-MAFPN and a multi-scale enhanced attention module SM-Detect are introduced to enhance feature expression and disease detection accuracy.
The accuracy of fish disease detection was improved. The mAP@0.5 and mAP@0.5:0.95 of the CA-YOLOv11 model reached 99.10% and 87.60% on the crucian carp and eel data sets, respectively, which was significantly better than other target detection algorithms and realized intelligent health monitoring of fish disease.
Smart Images

Figure CN120451761A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to fish disease identification, and more specifically to a fish disease identification method based on deep learning. Background Art
[0002] Crucian carp (Carassius auratus), eel, salmon, trout, and grouper are among the most commonly farmed fish species in China. They are highly sought after in domestic and international markets for their rapid growth, adaptability, and delicious meat. With the rapid development of aquaculture, the scale of crucian carp farming has continued to expand. However, the widespread adoption of intensive aquaculture practices has led to increased density and complex and variable water quality (Gao et al., 2019), resulting in an increased incidence of crucian carp diseases (Gui et al., 2022). Common crucian carp diseases include bacterial, parasitic, and viral diseases (Lei et al., 2024, Preena et al., 2022). Outbreaks of these diseases can cause widespread mortality in fish stocks, resulting in significant economic losses for farmers (Gao et al., 2019, Tu et al., 2015).
[0003] Timely and accurate identification and prevention of fish diseases are crucial for the healthy development of the aquaculture industry (Dileepa et al., 2023, Mia et al., 2022). However, traditional fish disease identification relies primarily on the experience and manual observation of fish farmers (Li et al., 2022), which is labor-intensive and time-consuming. Furthermore, it is highly subjective and susceptible to personal experience and environmental factors, making it difficult to detect diseases in a timely and accurate manner (Park et al., 2007). Therefore, there is an urgent need for an efficient, accurate, and automated fish disease identification method to improve the intelligence level of aquaculture production and reduce the risks posed by diseases (Ahmed et al., 2022, Banan et al., 2020).
[0004] With advances in artificial intelligence (AI), deep learning, due to its powerful feature extraction capabilities, has been widely applied in various fields, including pest and disease detection, animal behavior recognition, agriculture, animal husbandry, and industry. Leveraging image recognition technology, deep learning models can quickly and accurately identify crucian carp diseases, significantly improving detection efficiency and accuracy. These applications play a vital role in disease prevention and control, bringing new possibilities to fishery production. To address the increasing incidence and mortality of shrimp diseases, Ruifeng et al. (2024) proposed an improved YOLOv8 network. By combining Farnberck optical flow and GLCM to extract shrimp motion and texture features, they constructed a dataset and used a SVM classifier to classify the data, enabling shrimp health detection. To address the problem of automated tilapia fillet trimming, He et al. (2024) proposed an improved TFDS-YOLOv8n model, combining a coordinate attention mechanism with a Slim-Neck architecture and employing MPDIoU to accelerate convergence. Experimental results show that this model reduces parameters by 0.29MB, improves accuracy and real-time detection speed, and enables the detection of fillet defects. To address the high computational cost and large number of parameters in fish recognition, Ruan et al. (2024) proposed a lightweight network, DeformableFishNet, which combines efficient global coordinate attention with a deformable convolutional network to extract multi-scale features of fish. Experimental results show that this model achieves a mean average performance (MAP) of 96.30% across various underwater scenarios. To address the early detection of salmon diseases in aquaculture, Shoaib Ahmed et al. (2021) used a dataset containing both enhanced and unenhanced images and used a support vector machine (SVM) algorithm to extract features for disease classification. Experimental results show that the enhanced SVM achieves an accuracy of 94.12%, exceeding that of the unenhanced version. To address the subjectivity, limited accuracy, and poor real-time performance of surface feature detection of abnormal fish in water, Zhang et al. (2024) proposed a real-time, accurate detection model based on an improved YOLOv5s. By optimizing intersection-over-union (IoU) and non-maximum suppression, the model extracts deep features from complex backgrounds by integrating omnidirectional dynamic convolution and attention modules. Experimental results show that the proposed model outperforms baseline models in precision, recall, and mAP on the validation set. (Xu et al., 2023) To address the economic losses caused by clustered diseases in sea cucumber aquaculture, a DT-YOLOv5 intelligent recognition model was proposed. This model enhances feature extraction capabilities by combining a coordinated attention mechanism with a bidirectional feature pyramid network. Experimental results show that the proposed model achieves precision, recall, and mAP@50:95 of 99.43%, 98.91%, and 84.89%, respectively. This research provides theoretical support for detecting abnormal behavior of aquatic animals in intensive aquaculture.(Xu et al., 2023) To address the serious impact of ammonia nitrogen accumulation on fish life in aquaculture, this paper proposes a new water quality monitoring method based on deep learning and three-dimensional motion trajectories. This method uses an improved YOLOv8 model combined with a Kalman filter, Kuhn-Munkres algorithm, and kernelized correlation filter algorithm to obtain three-dimensional position information of fish. In experiments, the model achieved precision, recall, mAP@0.5, and mAP@0.5:0.95 of 96.40%, 91.40%, 97.90%, and 60.20%, respectively, in an acute ammonia nitrogen stress recovery experiment on sturgeon, sea bass, and crucian carp. This study provides a new method and approach for studying aquatic animal behavior under ammonia nitrogen stress. (Li et al., 2023) To address the difficulty in detecting behavioral changes in fish caused by water pollution and disease in aquaculture, a DCN-YOLOv5 model was proposed. This model adapts to fish posture changes by adjusting the convolutional receptive field and accurately detects key behavioral characteristics. Comparative experiments with traditional models have shown that DCN-YOLOv5 achieves higher precision and recall, providing effective theoretical and practical support for intelligent aquaculture. (Ge et al., 2024) To address the problem of timely detection of skin ulcer syndrome in intensive sea cucumber aquaculture, a new one-stage multi-target detection and tracking algorithm, SUS-YOLOv5, was proposed. This algorithm optimizes the maximum suppression algorithm for overlapping regions of object detection boxes and introduces an SE-BiFPN feature fusion structure to improve the efficient transmission of feature information between deep and shallow layers of the network. Experimental results show that SUS-YOLOv5 achieves mAP@0.5 and mAP@0.5:0.95 of 95.40% and 83.80%, respectively, representing improvements of 3.30% and 4.10%, respectively, compared to the original YOLOv5. This research lays a foundation for the prediction of sea cucumber skin ulcer syndrome and has practical application value in improving the level of intelligent aquaculture.
[0005] In the field of fish disease identification, some researchers have made active explorations. To address the issue of mortality in carp aquaculture due to the susceptibility of carp to Aeromonas arteriitis, researchers (Nadja et al., 2024) employed a modified YOLOv8 algorithm to automatically detect wounds caused by Aeromonas hydrophila in carp, thereby acquiring wound data. Experimental results showed that at two different disease concentrations, the accuracy of wound infection identification reached 91.49% and 88.68%, respectively, while the accuracy of fish body identification reached 96.61% and 93.44%. To address the issue of early detection of infected fish in aquaculture, researchers (Shoaib Ahmed et al., 2021) combined image preprocessing and segmentation to reduce noise and amplify images, and used a support vector machine (SVM) algorithm to extract relevant features for disease classification. Experimental results showed that the SVM achieved an accuracy of 91.42% and 94.12%, respectively, with and without image enhancement. To address the challenges of lice and wounds in the fish farming industry, Gupta et al. (2022) constructed a deep convolutional neural network (DCNN) that efficiently detected wounds and lice in live salmon farm ecosystems, achieving a high accuracy of 96.70%. To address the impact of salmon lice proliferation on salmon welfare, Zhang et al. (2024) used machine learning to develop a rapid detection method for salmon lice at different larval stages in seawater, called RT-DETR. Experimental results showed that RT-DETR performed best among all tested models, achieving 97.20% accuracy and 97.80% recall. As can be seen from the above review and existing related research, existing research on aquaculture detection suffers from low accuracy, large number of model parameters, and slow detection speed. Summary of the Invention
[0006] The purpose of the present invention is to solve the problem of insufficient recognition of fish due to the diversity of intensive aquaculture environments and the rapidity of their movements, and to propose a fish disease recognition method based on deep learning.
[0007] The technical solution of the present invention is:
[0008] A fish disease identification method based on deep learning includes the following steps:
[0009] Step 1: Collect fish videos of the ecological environment, set the sampling frame rate to extract single-frame fish images, and create a fish disease dataset;
[0010] Step 2: Improve the network structure based on YOLOv11:
[0011] 1) Introducing a multi-branch adaptive reparameterization module (CKDB) into the Backbone network. This module uses multiple parallel branches during training to capture rich fish disease features. During inference, these complex branches are adaptively reparameterized to merge into an efficient convolutional layer, enhancing the model's feature representation and outputting multi-scale fish disease features.
[0012] 2) A multi-scale adaptive feature pyramid network (N-MAFPN) is used in the Neck network. An adaptive feature fusion strategy is used to effectively fuse features of different scales and extract multi-branch features to improve the model's detection accuracy for fish diseases and further output multi-branch fused disease features.
[0013] 3) Introducing a multi-scale enhanced attention module, SM-Detect, into the Head network. This module uses an adaptive attention mechanism on feature maps at different scales to dynamically adjust the features of the diseased area the model focuses on, ultimately locating and identifying fish diseases and outputting the disease location and category.
[0014] Step 3: Use the prepared fish disease dataset to perform iterative supervised training on the improved YOLOv11 network, so that the training is stopped when the iteration cutoff condition is met, and the ideal model CA-YOLOv11 is obtained for identifying fish diseases.
[0015] It also includes step 4, collecting fish videos in the ecological water environment on site, extracting frames and inputting them into the ideal model CA-YOLOv11, and automatically outputting the fish disease recognition results in the current image.
[0016] The preparation of the fish disease dataset comprises the following steps:
[0017] Screening images: Screen out clear and healthy images or images with fish diseases;
[0018] Image enhancement: Perform image enhancement on single-frame images to expand the fish disease dataset; image enhancement includes but is not limited to methods such as random rotation, horizontal and vertical flipping, scaling, and cropping;
[0019] Labeling: Experts select healthy or diseased fish based on the external signs of the disease type in a single-frame fish image and add disease category labels. The fish disease category labels are either the primary labels {healthy, abnormal, dead} or the secondary abnormal labels {Aeromonas disease, Edwardsiella disease, and mechanical damage disease}.
[0020] The dataset was divided into training, validation, and test sets at random proportions, with each set including six classes.
[0021] Outward signs of the disease types include:
[0022] Symptoms of Aeromonas disease include: skin ulcers, swollen internal organs, loss of appetite, and rapid breathing in fish;
[0023] Symptoms of Edwardsiella disease include: skin ulcers, bleeding, abdominal distension and liver enlargement in fish;
[0024] Mechanical damage can manifest as skin tears, fin damage, or internal bleeding.
[0025] The multi-branch adaptive re-parameterization module CKDB is as follows: the DeepDBB module is introduced into the C3K2 module of the YOLOv11 backbone network to generate CKDB; the CKDB introduces parallel branches, each branch is used to learn a different feature representation; through multi-branch training, richer feature information is captured, and the network is optimized for feature extraction of different fish diseases.
[0026] The multi-scale adaptive feature pyramid network N-MAFPN achieves efficient fusion of multi-scale features by introducing shallow auxiliary fusion SAF and advanced auxiliary fusion AAF.
[0027] The Head network introduces a multi-scale enhanced attention module SM-Detect, which first enhances low-level features and then further improves the performance of high-level features by performing multi-level and multi-angle reasoning on features of different scales. It is used to ensure that targets of different scales are effectively retained and enhanced in the fused feature map. In SM-Detect, the channel and spatial mixing module CSMM uses depthwise separable convolution to learn the correlation between feature space scales and channels of different scales.
[0028] A fish disease identification system based on deep learning, including a video acquisition camera, a cloud server, and a host computer;
[0029] The video acquisition camera is arranged above the natural ecological water environment and is used to collect long time-series videos of the ecological activities of fish;
[0030] The cloud server is provided with a memory and a processor, the memory is provided with a database and a program module, the database is used to store long time-series videos captured by the video capture camera; when the processor loads the program module, the method steps described above are executed to train and obtain the ideal model CA-YOLOv11, and the ideal model CA-YOLOv11 is used to automatically identify diseases in images captured on site;
[0031] The host computer is located in the remote control room and includes a front-end interface and a control background. The front-end interface is used to collect data, model parameters and collection instructions input by the user and send them to the control background. The control background exchanges data, model parameters and instructions with the cloud server to control the operation of the cloud server.
[0032] The video acquisition camera is mounted on a three-dimensional translation platform. The three-dimensional translation platform receives acquisition instructions forwarded by the cloud server and controls the video acquisition camera to move on the three-dimensional translation platform to change the shooting angle to avoid environmental interference factors.
[0033] Alternatively, the video acquisition camera is deployed on a drone, which receives an acquisition instruction forwarded by a cloud server, flies to the target water area, and adjusts the shooting angle to acquire images.
[0034] The front-end interface is also provided with a human-computer interaction area and a data display area. The human-computer interaction area is used to provide users with a drawing component; the data display area is used to visually display disease images and long-term disease statistical charts.
[0035] The beneficial effects of the present invention are:
[0036] 1. This paper proposes a new multi-branch adaptive reparameterization module CKDB: it uses multiple parallel branches in the training phase to capture rich fish disease features, and merges these complex branches into an efficient convolutional layer through adaptive reparameterization in the inference phase, effectively improving the feature expression ability of the model.
[0037] 2. The present invention proposes a novel multi-scale adaptive feature pyramid network N-MAFPN: This network adopts an adaptive feature fusion strategy, which improves the model's detection accuracy for fish diseases by effectively fusing features of different scales and extracting multi-branch features.
[0038] 3. This paper proposes to introduce a multi-scale enhanced attention module SM-Detect at the head of the network: by using an adaptive attention mechanism on feature maps of different scales to dynamically adjust the disease area features that the model focuses on, the model's ability to detect fish diseases is improved.
[0039] 4. The effectiveness of the proposed algorithm structure is demonstrated through ablation experiments. The experimental results show that the mAP@0.5 and mAP@0.5:0.95 of CA-YOLOv11 for fish disease detection reach 99.10% and 87.60%, respectively. Compared with Faster R-CNN, YOLOv5, YOLOv6, YOLOv7, YOLOv8, YOLOv9t, YOLOv10 and YOLOv11, the mAP@0.5 is improved by 2.40%, 5.40%, 6.20%, 6.60%, 5.00%, 4.30%, 4.20% and 4.60%, respectively, and the mAP@0.5:0.95 is improved by 19.40%, 9.20%, 9.90%, 13.70%, 7.10%, 6.60%, 9.40% and 5.20%, respectively.
[0040] 5. This invention solves the problem of insufficient identification of fish due to the diversity of intensive aquaculture environments and the rapidity of their movements, and is a key advancement in promoting intelligent health monitoring of fish diseases. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flow chart of fish disease detection according to an embodiment of the present invention.
[0042] Figure 2 4 is a network structure diagram of CA-YOLOv11 according to an embodiment of the present invention.
[0043] Figure 3 This is a network structure diagram of CKDB according to an embodiment of the present invention.
[0044] Figure 4 FIG. 4 is a network structure diagram of an N-MAFPN according to an embodiment of the present invention.
[0045] Figure 5 FIG. 4 is a network structure diagram of the SAF according to an embodiment of the present invention.
[0046] Figure 6 FIG. 4 is a network structure diagram of the AAF according to an embodiment of the present invention.
[0047] Figure 7 FIG. 4 is a network structure diagram of MS-Detect according to an embodiment of the present invention.
[0048] Figure 8 These are the training results of different models in the embodiments of the present invention.
[0049] Figure 9 This is an example of the crucian carp disease detection result of an embodiment of the present invention.
[0050] Figure 10 PR curves of eight groups of models according to the embodiment of the present invention.
[0051] Figure 11 This is an example of the eel disease detection result according to an embodiment of the present invention. DETAILED DESCRIPTION
[0052] To make the above-mentioned objects, features, and advantages of the present invention more readily apparent, the specific implementation methods of the present invention are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the invention. Therefore, the present invention is not limited to the specific implementation methods disclosed below.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this invention belongs. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The present invention will be further described in detail below with reference to the accompanying drawings and examples.
[0054] 1. Summary of the content of the present invention:
[0055] In order to achieve real-time and accurate detection of fish diseases in complex recirculating aquaculture environments, this paper proposes a deep learning-based fish disease detection model CA-YOLOv11. The main contributions and innovations of this work are summarized as follows: (1) A fish disease dataset in multiple scenes and different scales, including different lighting and angles, is established. This provides high-quality data for fish disease detection research; (2) First, a new adaptive reparameterization module CKDB is proposed. This module captures fish disease features through parallel branches and merges them into efficient convolution kernels through adaptive reparameterization in the inference stage, effectively improving the feature expression ability of the model. Secondly, a novel multi-scale adaptive feature pyramid network N-MAFPN is proposed. By effectively fusing features of different scales and extracting multi-branch features, the model improves the detection accuracy of fish diseases. Finally, in order to accurately locate the fish disease area, this paper introduces a multi-scale enhanced attention module SM-Detect at the head of the network. The adaptive attention mechanism dynamically adjusts the disease features of interest, thereby improving the detection ability of the model. (3) This algorithm, based on a comparative analysis of eight typical target detection algorithms, achieved a 99.10% accuracy rate on a 588 FPS crucian carp disease dataset. This algorithm supports precise identification and intelligent decision-making in aquaculture. The algorithm also achieved a 93.80% accuracy rate on an eel dataset.
[0056] 2. In order to describe the method and steps in detail, the method and steps are first described in detail using crucian carp as the object, and then briefly described using eel.
[0057] 1. Methods and Steps
[0058] 1.1 Data Collection and Dataset Creation
[0059] The crucian carp disease dataset used in this article was shot at the Yantai Coastal Zone Research Institute of the Chinese Academy of Sciences in Yantai, Shandong Province, China. The data acquisition equipment Hikvision camera (resolution 1920 pixels × 1080 pixels, frame rate 25f / s) was installed above the crucian carp pond. By screening the videos, 5 video clips with more crucian carp diseases were selected, and the length of each video ranged from 2 minutes to 5 minutes. In the absence of external interference, crucian carp swims slowly and has a small amplitude. Too short a time interval for taking frames can easily lead to high data repeatability. Therefore, in order to ensure the diversity of crucian carp disease data, video framing technology is used to capture one image every 5 seconds. The captured images will have poor image quality and no crucian carp disease targets in the images, and they will be manually discarded. The crucian carp disease detection process is as follows: Figure 1 shown.
[0060] from Figure 1 As can be seen in Figure 1, Step A: Constructing feature descriptions of different crucian carp diseases. The crucian carp disease dataset used in this paper includes four categories: healthy, aeromonas, Edwardsiella, and mechanical damage. Aeromonas is a bacterial infection caused by Aeromonas bacteria, manifested by skin ulcers, visceral swelling, loss of appetite, and rapid breathing in fish. It often occurs in environments with poor water quality or low immunity. Edwardsiella is a bacterial infection caused by Edwardsiella bacteria, manifested by skin ulcers, bleeding, abdominal distension, and liver enlargement in fish. It is common in environments with poor water quality or low immunity (Park et al., 2012). Mechanical damage is physical injury caused by external impact, friction, or collision, manifested by skin tears, fin damage, or internal bleeding. It is common during transportation, intensive farming, or harsh environments. Step B: Constructing a dataset of crucian carp disease images. To improve the robustness of the model in crucian carp disease detection, this study employed various data augmentation techniques. Specific methods include random image rotation, horizontal and vertical flipping, scaling, and cropping. Step C: Manually annotate the characteristics of different crucian carp diseases using the LabelImg image annotation tool. The dataset was randomly divided into training, validation, and test sets in a ratio of 7:2:1. Step D: Construct a deep learning model, CA-YOLOV11, for crucian carp diseases. Step E: Compare the performance of various deep learning models.
[0061] 1.2 Improved CA-YOLOv11 Network
[0062] YOLO, a widely used object detection algorithm, excels in real-time performance and detection accuracy, making it particularly suitable for applications requiring rapid response. Since its launch, the YOLO algorithm family has undergone multiple iterations, with the algorithm architecture progressively optimized from YOLOv1 to YOLOv11. Each version has been optimized for different needs and challenges, improving detection speed and accuracy, particularly when dealing with complex backgrounds and multi-target detection. YOLOv1: It first achieved real-time object detection, but faced challenges in detecting small objects. YOLOv2 and YOLOv3 introduced anchor boxes and batch normalization, improving small object detection capabilities. YOLOv4 and YOLOv5 made improvements in data augmentation and PyTorch implementation, increasing practicality and accessibility. YOLOv6 to YOLOv10 introduced more efficient architectures and training methods, such as the use of transformers and training without NMS (non-maximum suppression), improving performance and scalability on edge devices. YOLOv11 is the most advanced version in the YOLO series. It improves object detection performance through several innovations, including a Transformer-based backbone network that enhances small object detection; a dynamic head design that dynamically adjusts resource allocation based on image complexity to improve processing efficiency; NMS-free training that reduces inference time while maintaining accuracy; dual-label assignment to improve overlapping object detection; large kernel convolution to optimize feature extraction and reduce computational resource consumption; and a partial self-attention mechanism (PSA) that applies attention to partial regions of the feature map to improve global feature learning without significantly increasing computational burden. YOLOv11 is available in five versions: n, s, m, l, and x, based on the depth and width of the network architecture. These versions differ primarily in model size, computational resource requirements, detection accuracy, and speed. Table 1 shows the structural parameters of the five YOLOv11 versions.
[0063] Table 1 Structural parameters of the YOLOv11 series As can be seen from Table 1, in order to facilitate deployment on embedded devices and to minimize memory and computing requirements while ensuring reasonable accuracy, this study chooses YOLOv11n for optimization and improvement.
[0064] The YOLOv11 model consists of four key components: the input, the backbone, the neck, and the head. First, the input receives and preprocesses image data, including resizing images to a uniform size (e.g., 640x640), normalizing pixel values (scaling them to the range [0, 1]), and applying data augmentation techniques (such as random cropping, rotation, and color transformation) to improve model robustness. The backbone, as the core feature extraction component, extracts low-level and high-level image features through convolutional layers, activation functions, and pooling layers, generating a hierarchical feature representation. Its design incorporates the advanced Transformer network architecture to enhance the ability to extract detailed and semantic information from images. The neck further processes and fuses features from the backbone, employing structures such as FPN (Feature Pyramid Network) or PAN (Path Aggregation Network) to achieve multi-scale feature fusion. By optimizing feature transfer and cross-layer information flow, the neck improves the detection of objects of varying sizes, particularly small objects. Finally, the head network generates the final output of the model, namely the bounding box, object category, and confidence prediction. The head network predicts the bounding box of the object (including position and size) through regression, and performs category classification and confidence estimation for each bounding box. Through these outputs, YOLOv11 can efficiently perform target detection and use NMS-free technology to optimize the inference process to improve speed and accuracy. Overall, these four parts work closely together, enabling YOLOv11 to perform efficient and accurate target detection in various application scenarios. In order to better identify diseases of crucian carp, this study improved YOLOv11n and proposed CA-YOLOv11. The improved network structure is as follows Figure 2 shown.
[0065] from Figure 2 As can be seen from the figure, this paper proposes three improvement methods: first, a multi-branch adaptive re-parameterization module CKDB is proposed, which extracts rich crucian carp disease features through multiple parallel branches during training, and optimizes the convolution layer through re-parameterization during inference to enhance the feature expression ability of the model; second, a multi-scale adaptive feature pyramid network N-MAFPN is designed, which improves the detection accuracy of diseases of different sizes by effectively fusing multi-scale features; finally, a multi-scale enhanced attention module MS-Detect is adopted, which introduces the attention mechanism on feature maps of different scales to highlight the disease area and suppress background noise, thereby improving the detection ability of small target diseases.
[0066] 1.2.1 Enhanced backbone network module
[0067] In the YOLOv11 network, C3K2 uses the Botteneck module to reduce the feature dimension, but due to the convolution operation, some feature information may be lost, which will limit the model's feature representation ability and thus affect the model's performance. In the underwater biological target detection task, due to the richness of the feature scale, the original Botteneck module performs poorly in extracting features from the target, resulting in reduced model accuracy and performance. In order to enhance the feature extraction ability of YOLOv11 in the crucian carp disease detection task, this study introduced the DeepDiverseBranchBlock (DeepDBB) module in the C3K2 module of the YOLOv11 backbone network, and proposed CKDB. The CKDB module uses a multi-branch adaptive re-parameter module to enhance the network's ability to learn different crucian carp disease characteristics through multi-branch training. This method introduces parallel branches within the module, and each branch is responsible for learning different feature representations. Through multi-branch training, the network can capture richer feature information and improve the network's feature extraction ability for different crucian carp diseases. The network structure of CKDB is as follows Figure 3 shown.
[0068] from Figure 3 As can be seen from the figure, DeepDBB consists of a multi-branch topology structure with 1x1 convolution, sequential 1x1-KxK convolution, average pooling branch, BN, and branch addition. The design of separating training and inference in DeepDBB can make the model efficient during the inference phase. During the training phase, DeepDBB consists of convolution layers and average pooling layers of different sizes, which are arranged in parallel in a complex way and finally merge the outputs. The convolution operation is shown in formula (1). The continuous 1x1 and KxK convolutions can be converted into a single KxK convolution kernel using formula (2). After training, these complex structures will be converted into a single convolution layer, and the linear additivity property of the convolution module will be used to sum multiple convolution kernels to produce a single KxK convolution kernel for model inference and deployment. This additivity principle is shown in formula (3).
[0069]
[0070] Among them, O represents the output of the convolution kernel, I represents the input image or feature, F represents the convolution kernel, b is the bias after the convolution layer is re-parameterized, F′ represents the new convolution kernel obtained after some transformation or merging operation, F (1) Represents a 1×1 convolution kernel for linear combination between channels, F (2) Represents a K×K convolution kernel for spatial aggregation, and b′ represents the combination of b (1) and b (2) The new bias term obtained later, b (1) With the convolution kernel F(1) The associated bias term, b (2) Represents the convolution kernel F (2) The associated bias term, REP, indicates that the bias b is repeated or extended under a specific structure so that it can be applied to the corresponding channels in the convolution operation.
[0071] By decomposing and bundling different convolutional operations, the DeepDBB module enables the network to learn different feature representations and enhance its detection performance. In summary, DeepDBB, through its multi-branch topology and separation of training and inference, enhances the model's feature representation capabilities and detection accuracy while maintaining high efficiency in the inference phase, enabling CA-YOLOv11 to achieve higher accuracy in the crucian carp disease detection task.
[0072] 1.2.2 Improved PANet Enhanced Feature Fusion
[0073] In the identification of crucian carp diseases, effective feature fusion is crucial to improving detection accuracy and robustness. Traditional feature fusion methods, such as FPN, PAN, and BiFPN, have achieved good results in multi-scale target detection, but still suffer from problems such as information loss, difficulty in feature alignment, and high computational complexity. FPN processes objects of different scales by constructing a feature pyramid, but it is prone to losing low-level detail information; PAN strengthens information transfer but increases computational overhead; BiFPN optimizes information flow and fusion strategies, improving accuracy, but in extreme cases there are still problems with detail loss and high training complexity. In summary, when dealing with these problems, traditional target detection algorithms usually face the limitations of low small target detection accuracy and insufficient multi-scale feature fusion efficiency.
[0074] In order to solve the above problems, this study proposes a multi-scale adaptive feature pyramid network N-MAFPN. N-MAFPN achieves efficient fusion of multi-scale features by introducing shallow auxiliary fusion (SAF) and advanced auxiliary fusion (AAF). Its network structure is as follows Figure 4 shown.
[0075] from Figure 4 As can be seen in the figure, in the top-down path, the SAF module extracts edge and detail information from the multi-scale features output by the backbone network, enhancing the ability to identify small target disease areas. In the top-down path, the AAF module dynamically fuses shallow and deep features through cross-layer connections, improving the model's robustness in complex disease scenarios. The outputs of SAF and AAF are further integrated through the C3K2 module. C3K2's lightweight design not only reduces computational overhead but also optimizes feature expression, providing the detection head with richer and more diverse input features.
[0076] Crucian carp disease detection usually relies on edge and local detail information in shallow features, which is crucial for accurate disease identification. However, since shallow features are relatively primitive and easily affected by background interference, direct use may lead to instability in subsequent feature learning. To solve this problem, this study introduced SAF in N-MAFPN. The network structure is as follows: Figure 5 shown.
[0077] from Figure 5 As can be seen from the figure, the main goal of the SAF module is to combine high-resolution shallow features with deep features in the backbone network. Through multi-branch fusion and feature enhancement, the SAF module retains key local details (such as edge and texture information) in the shallow features while removing possible background noise. In addition, the SAF module provides the network with high-quality feature representations that contain both spatial details and semantic information by synergistically fusing shallow and deep features. In addition, this study uses 1×1 convolution to control the number of channels in the shallow information, ensuring that it occupies a smaller proportion in the Concat operation without affecting subsequent learning. This design effectively improves the model's ability to detect small target areas of crucian carp diseases, especially in complex backgrounds, and shows higher robustness. The output results after applying SAF are as follows:
[0078] P n ′=Concat(δ(C(Down(P n-1 ))),P n ,U(P′ n+1 ))(4)
[0079] Among them, P n-1 , P n and P n+1 are feature maps of different layers in the network, representing adjacent layers. Usually, they correspond to different resolutions, P n is the current layer, P n-1 It is the upper layer, P n+1 It is the next layer. n Represents the feature map of the current layer after operation, which can be considered as the new feature map obtained after fusion or processing. n+1 represents the high-resolution feature map generated by downstream operations. U(·) represents an upsampling operation. Down represents a downsampling operation, typically using 3×3 convolution and batch normalization to reduce the spatial dimension of the feature map. δ represents the SiLU activation function, and C represents a 1×1 convolution that controls the number of channels, typically used to adjust the number of channels or perform feature map transformations.
[0080] In order to further enhance the interactive utilization of feature layer information, this study introduces the advanced auxiliary fusion module (AAF) in the deeper layer of N-MAFPN. The network structure is as follows: Figure 6 The main goal of the AAF module is to improve the detection performance of the model for medium-sized disease targets by integrating multi-level features.
[0081] from Figure 6 It can be seen from the figure that when the AAF module integrates medium-resolution features, it can obtain the high-resolution features from the shallow layer (P′ n+1 ), shallow low-resolution layer (P′ n-1 ), the shallow layer of the same level (P′ n ) and the previous layer (P') n-1 ) for information aggregation. This multi-path feature interaction ensures that the final output feature map simultaneously contains detailed information, contextual information, and semantic information. By using 1×1 convolution to adjust the number of channels, the AAF module dynamically balances the impact of each layer's features on the result, significantly improving feature diversity and stability. The output after applying AAF is as follows:
[0082] P n =Concat(δ(C(Down(P′ n-1 ))),δ(C(Down(P” n-1 ))),P n ′,C(U(P′ n+1 ))) (5)
[0083] Among them, P n Represents the processed feature map of the current layer in N-MAFPN, and represents the output of the current feature map.
[0084] In addition, the AAF module uses a channel balancing strategy to ensure the same number of channels in shallow and deep layers, reducing information loss or conflict. In this study, the introduction of the AAF module significantly enhanced the detection performance of medium-target diseases, enabling the model to more accurately locate crucian carp diseases and demonstrate greater robustness to crucian carp disease characteristics in complex backgrounds.
[0085] 1.2.3 Improved Detection Head Network
[0086] In traditional target detection models, especially in the YOLO series, the detection head usually relies on convolutional feature maps at different levels to process targets. However, these methods often fail to fully utilize the advantages of multi-level features when dealing with small objects or targets of different sizes. In order to improve the performance of YOLOv11 in multi-scale target detection, this study proposed the SM-Detect module in the YOLOv11 detection head. The core idea of the SM-Detect module is to enhance the model's perception of targets of different scales through information exchange between multiple scale feature maps. Specifically, the SM-Detect module first enhances the low-level features, and then further improves the performance of high-level features by fusing features of different scales. This enhancement mechanism utilizes the enhancement of multi-scale features by convolution operations, ensuring that targets of different scales are effectively retained and enhanced in the fused feature maps. Its network structure is as follows: Figure 7 shown.
[0087] like Figure 7 As shown in the figure, the channel and spatial mixing module (CSMM) in SM-Detect uses depthwise separable convolution to learn the correlation between spatial scales and channels at different feature scales. First, the input image is segmented into patches of different sizes (3, 5, 7). These patches are initially processed through the Patch Embedding layer to generate feature representations. Second, depthwise separable convolution is used to learn the correlation between spatial dimensions and channels. This operation is divided into two steps: depthwise convolution performs convolution operations on each channel independently; and pointwise convolution performs 1x1 convolution on all channels to integrate information. Then, the GELU activation function is used after the convolution operation to introduce nonlinearity, and batch normalization is performed on the feature maps to stabilize the training process. Finally, the output features of the CSMM module are average pooled to fuse the feature representations of multi-scale features, enhancing the model's ability to capture information at different scales. In summary, SM-Detect can leverage features from different scales for multi-level and multi-angle reasoning during detection, thereby improving overall detection accuracy and robustness.
[0088] 2. Experimental Results and Analysis
[0089] 2.1 Experimental Environment and Evaluation Metrics
[0090] Experiments were conducted on a platform equipped with an Intel(R) Core(TM) i7-6800K CPU @ 3.40GHz, 16GB of RAM, and an NVIDIA GeForce RTX 3060 12GB GPU. The software environment was configured with CUDA 11.1.0, CUDNN 11.1, and Python 3.8.8. Considering the hardware performance and training results, this paper adopted batch training, dividing the training and validation processes into multiple batches. During model training, images were resized to 640×640 pixels as model input, and the StepLR mechanism was used to update the learning rate. SGD was the optimizer used, and other training parameters are shown in Table 2.
[0091] Table 2. Training parameter settings
[0092]
[0093] To objectively evaluate the model's performance in detecting crucian carp diseases, this study used a combination of the performance evaluation metrics provided in Table 3. The evaluation criteria primarily included precision, recall, mAP, FLOPs, and FPS. mAP was used as the primary performance metric for model evaluation.
[0094] Table 3 Performance evaluation indicators of target detection models
[0095]
[0096] Note: TP (True Positive): the number of samples correctly classified as positive by the model; TN (True Negative): the number of samples correctly classified as negative by the model; FP (False Positive): the number of samples incorrectly classified as positive by the model; FN (False Negative): the number of samples incorrectly classified as negative by the model; H×W represents the size of the output feature map; C out Indicates the output channel, C in represents the input channel, K represents the size of the convolution kernel, and T represents the time required for the model to process the image.
[0097] 2.2 Experimental results
[0098] In order to verify the effectiveness of YOLOv11 before and after improvement, this study uses P, R, Loss function, mAP@0.5, and mAP@0.5:0.9 to evaluate the performance of the model. The experimental results are as follows Figure 8 shown.
[0099] In order to effectively verify the accuracy and detection ability of CA-YOLOv11 in detecting crucian carp diseases, Figure 8As can be seen from (a) and (b), the accuracy and recall of CA-YOLOv11 are significantly higher than those of other models. In order to accurately reflect the detection accuracy of the model under different threshold conditions, this study introduces mAP@0.5 and mAP@0.5:0.95 to further evaluate the performance of the model. Figure 8 It can be seen from (c) and (d) that after about 280 iterations, the mAP@0.5 and mAP@0.5:0.95 of CA-YOLOv11 reached 99% and 87% respectively, and gradually stabilized, with the maximum values reaching 99.10% and 87.60% respectively. The experimental results show that compared with YOLOv11n, the mAP@0.5 and mAP@0.5:0.95 of CA-YOLOv11 increased by 4.6 percentage points and 5.2 percentage points respectively. The results show that CA-YOLOv11 can more accurately identify crucian carp diseases in complex circulating water environments. During the model training process, the loss function value can intuitively reflect whether the model can converge stably with the increase in the number of iterations. From Figure 8 As can be seen in (e) and (f), the training / loss and verification / loss curves of CA-YOLOv11 gradually converge with the number of iterations. After approximately 290 iterations, the loss function value stabilizes, indicating that the model has essentially converged. Compared with YOLOv11n, CA-YOLOv11 converges faster, indicating that the algorithm is able to better integrate the aspect ratio between the ground truth and predicted frames to better learn the characteristics of crucian carp diseases. In summary, the CA-YOLOv11 proposed in this study can more accurately identify crucian carp disease targets in complex environments.
[0100] Ablation Experiment
[0101] To more accurately identify the disease characteristics of crucian carp, this study conducted ablation experiments based on YOLOv11n to verify the effectiveness of the improved method. A total of eight groups of models were verified. The experimental results are shown in Table 4.
[0102] Table 4 Ablation experiment
[0103]
[0104]
[0105] As can be seen in Table 4, YOLOv11n has the lowest mAP value. However, in actual aquaculture, accurate identification of crucian carp diseases is key to intelligent health monitoring. Therefore, this paper conducts ablation experiments based on YOLOv11n using the proposed MS-Detect, CKDB, and N-MAFPN. The results of the ablation experiments show that each improvement improves the model's mAP. MS-Detect can effectively enhance the detection model's ability to detect objects of different scales, increasing the model's mAP@0.5 and mAP@0.5:0.95 by 2.20% and 2.40%, respectively. CKDB, by introducing a deeper branching structure, is able to capture richer feature information, especially in complex scenes and occlusion situations, enhancing the network's feature expression capabilities and increasing the model's mAP@0.5 and mAP@0.5:0.95 by 2.30% and 2.00%, respectively. N-MAFPN utilizes an adaptive feature pyramid structure to effectively fuse information at different scales, enabling the network to better handle objects of varying sizes. This performance is particularly strong for multi-scale object detection in complex scenes, improving the model's mAP@0.5 and mAP@0.5:0.95 by 2.50% and 1.60%, respectively. By integrating these three methods, this study comprehensively enhances CA-YOLOv11's strengths in multi-scale feature fusion, dynamic feature selection, detail capture, and robustness, resulting in a 4.60% increase in mAP, the highest ever. Furthermore, as shown in Table 4, CA-YOLOv11 also excels in precision and recall. Compared to the other seven models, the proposed CA-YOLOv11 demonstrates superior performance.
[0106] In order to detect carp diseases in complex environments, this study selected partially occluded, different time, different lighting and multi-target images to verify the effectiveness of CA-YOLOv11. Figure 9 shown.
[0107] from Figure 9 As can be seen from (a), YOLOv11 has the phenomenon of false detection and missed detection in crucian carp disease detection. The main reason is that the model's feature extraction ability is insufficient. Specifically, when processing crucian carp disease images, YOLOv11 is unable to fully capture the subtle features and complex pathological patterns of the disease, causing the model to mistake parts of normal fish bodies for diseased areas in some cases, or fail to accurately identify actual diseases. This limitation in feature extraction ability significantly affects the performance of YOLOv11 in high-precision disease detection tasks and limits its application in actual aquaculture management. Figure 9As can be seen in (b), CA-YOLOv11 can accurately identify crucian carp disease characteristics. This is primarily due to the introduction of the CKDB, N-MAFPN, and MS-Detect modules into the model's backbone, neck, and head, respectively. These innovative improvements collectively enhance the model's ability to extract features for crucian carp diseases, thereby overcoming the shortcomings of YOLOv11n in detection performance. Specifically, the CKDB module fully captures disease characteristics during training through a multi-branch structure, while maintaining the model's efficiency and lightweight through reparameterization during inference. N-MAFPN, through the adaptive fusion of multi-scale features, addresses the difficulty in uniformly processing disease features at different scales, ensuring the model's robustness and accuracy when handling diverse disease features. The introduction of the SM-Detect module enables the model to dynamically focus on key disease areas during detection, reducing background noise interference and improving detection accuracy and reliability. Based on the above improvements, the CA-YOLOv11 proposed in this paper shows significant advantages in the crucian carp disease detection task, which not only effectively reduces false detection and missed detection, but also improves the overall detection efficiency and accuracy.
[0108] In addition, this study uses the Precision-Recall (PR) curve to evaluate the performance of the proposed CA-YOLOv11 target detection model. The PR curve can provide a more comprehensive and accurate evaluation by showing the relationship between precision and recall. In target detection, precision reflects how many of the targets identified by the model are correct, while recall reflects how many actual targets the model can find. By adjusting the classification threshold of the model, different combinations of precision and recall can be obtained to draw a PR curve. The PR curve drawn by this study through ablation experiments is shown below. Figure 10 shown.
[0109] from Figure 10As can be seen from the PR curve, the area enclosed by the PR curve and the coordinate axis of the CA-YOLOv11 proposed in this study is the largest, indicating that the better the model performs on positive class detection, the higher the precision can be maintained at a higher recall rate. This is mainly due to the deep improvements to the backbone of the model. CKDB enhances the network's feature extraction capabilities in low-contrast and occlusion conditions; the MS-Detect attention mechanism selectively strengthens the response of different scales and different features, allowing the model to focus more on discriminative feature areas, thereby improving overall detection accuracy; the N-MAFPN structure improves the fusion of multi-scale information, allowing the model to exhibit higher precision and recall in the detection of targets of different sizes, thereby increasing the overall area of the PR curve. In summary, the CA-YOLOv11 proposed in this study significantly improves the performance in target detection tasks, especially in terms of the balance between precision and recall. Through the evaluation of the PR curve, this study verifies the effectiveness of the proposed method in handling imbalanced datasets, improving target detection accuracy, and reducing false positives.
[0110] 2.4. Comparison with other models
[0111] To further verify the superiority of CA-YOLOv11 in crucian carp disease recognition, this paper compared CA-YOLOv11 with some mainstream object recognition models. The comparison results of the models on the same dataset are shown in Table 5.
[0112] Table 5 Comparison with eight other target detection models
[0113]
[0114] As can be seen from Table 5, Faster R-CNN achieves an accuracy of 86.15% and a recall of 96.60%, demonstrating its good performance in capturing most crucian carp disease targets. However, its inference speed is slow and the model is large, making it unsuitable for real-time deployment. In addition, while its mAP@0.5 is 96.70%, its mAP@0.5:0.95 is only 68.20%, indicating weak detection accuracy at high IoU values. Compared to Faster R-CNN, the YOLO series of models performs better in terms of accuracy and inference speed. In particular, YOLOv5n and YOLOv8n maintain high detection accuracy while having very fast inference speed, making them suitable for real-time applications. YOLOv6n and YOLOv9t also perform well, with good detection accuracy and real-time performance. Despite their slower inference speed, they are still suitable for real-time detection tasks. YOLOv10n achieved optimal inference speed, with an accuracy of 93.00% and an mAP@0.5 of 94.90%. This model offers a very balanced performance in terms of accuracy and speed, making it suitable for scenarios requiring efficient detection, but its mAP@0.5:0.95 ratio is slightly lower. YOLOv11n outperformed other YOLO models in both precision and recall, achieving an mAP@0.5:0.95 ratio of 82.40%, demonstrating greater robustness against various crucian carp diseases. CA-YOLOv11 achieved optimal model performance with superior accuracy, recall, mAP@0.5, and mAP@0.5:0.95 ratios. Although CA-YOLOv11 has a slightly increased number of parameters and computational complexity compared to its predecessor, its mAP@0.5 and mAP@0.5:0.95 ratios increased by 4.60% and 5.20% respectively. CA-YOLOv11's ultra-high accuracy makes it particularly suitable for crucian carp disease detection tasks, which require extremely high detection accuracy. Furthermore, in actual testing, the average detection time for a single image was only 0.0017 seconds, meeting the requirements for real-time crucian carp disease detection. Furthermore, CA-YOLOv11 is a lightweight model, with a model size of only 11.60MB, making it easier to deploy on embedded devices.
[0115] 3. Conclusion
[0116] With the intensive and large-scale development of crucian carp aquaculture, accurate identification of crucian carp diseases has become particularly important. To this end, this study proposed a deep learning algorithm, CA-YOLOv11, to identify crucian carp diseases in circulating water monitoring scenarios. First, the CKDB module captures rich features through multi-branch parallelism during the training phase and adaptively merges them into efficient convolution kernels during the inference phase, enhancing feature expression. Second, the N-MAFPN network adopts a multi-scale adaptive feature fusion strategy to optimize the extraction and integration of features at different scales, thereby improving detection accuracy. Finally, the SM-Detect module introduces a multi-scale enhanced attention mechanism in the detection head to dynamically adjust the features of the diseased area of interest, further improving the accuracy of disease location and identification. Experimental results show that CA-YOLOv11 mAP@0.5 and mAP@0.5:0.95 reach 99.10% and 87.60% respectively. Compared with Faster R-CNN, YOLOv5, YOLOv6, YOLOv7, YOLOv8, YOLOv9t, YOLOv10 and YOLOv11, mAP@0.5 is improved by 2.40%, 5.40%, 6.20%, 6.60%, 5.00%, 4.30%, 4.20% and 4.60% respectively, and mAP@0.5:0.95 is improved by 19.40%, 9.20%, 9.90%, 13.70%, 7.10%, 6.60%, 9.40% and 5.20% respectively. The method proposed in this study can detect crucian carp diseases with high precision in complex scenarios, which has important research significance for further promoting large-scale crucian carp farming, improving animal health monitoring capabilities and animal welfare.
[0117] 4. The present invention uses eel as the research object for identification method steps as follows:
[0118] The data collection process is the same as that of crucian carp, and the data preprocessing steps are also the same.
[0119] Experts classify eels into three categories: healthy behavior, abnormal behavior, and death behavior.
[0120] Similarly, the network structure was improved based on YOLOv11, and iterative learning and training were performed using the collected eel dataset to obtain the ideal model CA-YOLOv11, which is used to identify diseases and further promote its application.
[0121] The recognition results are as follows Figure 11 As shown, the experimental results show that CA-YOLOv11 has an accuracy of 93.80%, a recall of 93.70%, a mAP@0.5 of 95.40%, and a mAP@0.5:0.95 of 80.60%.
[0122] In summary, based on the detailed steps for crucian carp and the simplified steps for eel, the experts further developed disease category labels, including primary labels {healthy, abnormal, dead} and secondary abnormality labels {Aeromonas, Edwardsiella, and mechanical damage}. Labeling is performed using either the primary or secondary label for evaluation.
[0123] 3. The present invention also provides another crucian carp disease identification system based on deep learning, including a video acquisition camera, a cloud server, and a host computer;
[0124] The video acquisition camera is set above the natural ecological water environment to collect long-time sequence videos of fish ecological activities;
[0125] The cloud server is provided with a memory and a processor. The memory is provided with a database and a program module. The database is used to store long time-series videos captured by the video capture camera. When the processor loads the program module, the method steps are executed to train and obtain an ideal model CA-YOLOv11, and the ideal model CA-YOLOv11 is used to automatically identify diseases in images captured on-site.
[0126] The host computer is located in the remote control room and includes a front-end interface and a control background. The front-end interface is used to collect user input data, model parameters and collection instructions and send them to the control background. The control background interacts with the cloud server for data, model parameters and instructions to control the operation of the cloud server.
[0127] The video acquisition camera is installed on a three-dimensional displacement platform, receives the acquisition instructions forwarded by the cloud server, and moves on the three-dimensional displacement platform to change the shooting angle to avoid environmental interference factors, including light angle and other factors; or the video acquisition camera is deployed on a drone, which receives the acquisition instructions forwarded by the cloud server, flies to the target water area and adjusts the shooting angle to capture images.
[0128] The front-end interface is also provided with a human-computer interaction area and a data display area. The human-computer interaction area is used to provide users with a marking component for experts or users to frame or classify disease images; the data display area is used to visually display disease images and long-term disease statistical charts.
[0129] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural change made to the above embodiment based on the technical essence of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A fish disease identification method based on deep learning, characterized in that: The following steps are involved: Step 1: Collect fish videos of the ecological environment, set the sampling frame rate to extract single-frame fish images, and create a fish disease dataset; Step 2: Improve the network structure based on YOLOv11: 1) Introducing the multi-branch adaptive re-parameterization module CKDB into the Backbone network; By using multiple parallel branches in the training phase to capture rich fish disease features, and adaptively reparameterizing these complex branches into an efficient convolutional layer in the inference phase, it is used to enhance the model's feature expression and output multi-scale fish disease features. 2) A multi-scale adaptive feature pyramid network (N-MAFPN) is used in the Neck network. An adaptive feature fusion strategy is used to effectively fuse features of different scales and extract multi-branch features to improve the model's detection accuracy for fish diseases and further output multi-branch fused disease features. 3) Introducing a multi-scale enhanced attention module, SM-Detect, into the Head network. This module uses an adaptive attention mechanism on feature maps at different scales to dynamically adjust the features of the diseased area the model focuses on, ultimately locating and identifying fish diseases and outputting the disease location and category. Step 3: Use the prepared fish disease dataset to perform iterative supervised training on the improved YOLOv11 network, so that the training is stopped when the iteration cutoff condition is met, and the ideal model CA-YOLOv11 is obtained for identifying fish diseases.
2. The fish disease identification method based on deep learning according to claim 1, characterized in that: It also includes step 4, collecting fish videos in the ecological water environment on site, extracting frames and inputting them into the ideal model CA-YOLOv11, and automatically outputting the fish disease recognition results in the current image.
3. The fish disease identification method based on deep learning according to claim 1, characterized in that: The preparation of the fish disease dataset comprises the following steps: Screening images: Screen out clear and healthy images or images with fish diseases; Image enhancement: Perform image enhancement on single-frame images to expand the fish disease dataset; image enhancement includes but is not limited to methods such as random rotation, horizontal and vertical flipping, scaling, and cropping; Labeling: Experts select healthy or diseased fish based on the external signs of the disease type in a single-frame fish image and add disease category labels. The fish disease category labels are either the primary labels {healthy, abnormal, dead} or the secondary abnormal labels {Aeromonas disease, Edwardsiella disease, and mechanical damage disease}. The dataset was divided into training, validation, and test sets at random proportions, with each set including six classes.
4. The fish disease identification method based on deep learning according to claim 1, characterized in that: Outward signs of the disease types include: Symptoms of Aeromonas disease include: skin ulcers, swollen internal organs, loss of appetite, and rapid breathing in fish; Symptoms of Edwardsiella disease include: skin ulcers, bleeding, abdominal distension and liver enlargement in fish; Mechanical damage can manifest as skin tears, fin damage, or internal bleeding.
5. The fish disease identification method based on deep learning according to claim 1, characterized in that: The multi-branch adaptive re-parameterization module CKDB is as follows: the DeepDBB module is introduced into the C3K2 module of the YOLOv11 backbone network to generate CKDB; the CKDB introduces parallel branches, each branch is used to learn a different feature representation; through multi-branch training, richer feature information is captured, and the network is optimized for feature extraction of different fish diseases.
6. The fish disease identification method based on deep learning according to claim 1, characterized in that: The multi-scale adaptive feature pyramid network N-MAFPN achieves efficient fusion of multi-scale features by introducing shallow auxiliary fusion SAF and advanced auxiliary fusion AAF.
7. The fish disease identification method based on deep learning according to claim 1, characterized in that: The Head network introduces a multi-scale enhanced attention module SM-Detect, which first enhances low-level features and then further improves the performance of high-level features by performing multi-level and multi-angle reasoning on features of different scales. It is used to ensure that targets of different scales are effectively retained and enhanced in the fused feature map; in SM-Detect, the channel and spatial mixing module CSMM uses depth-separable convolution to learn the correlation between feature space scales and channels of different scales.
8. The fish disease identification system based on deep learning according to claim 1, characterized in that: Including video acquisition camera, cloud server, and host computer; The video acquisition camera is arranged above the natural ecological water environment and is used to collect long time-series videos of the ecological activities of fish; The cloud server is provided with a memory and a processor, the memory is provided with a database and a program module, the database is used to store the long time series video captured by the video capture camera; when the processor loads the program module, the method steps according to any one of claims 1 to 7 are executed to train and obtain the ideal model CA-YOLOv11, and the ideal model CA-YOLOv11 is used to automatically identify the disease of the image collected on site; The host computer is located in the remote control room and includes a front-end interface and a control background. The front-end interface is used to collect data, model parameters and collection instructions input by the user and send them to the control background. The control background exchanges data, model parameters and instructions with the cloud server to control the operation of the cloud server.
9. The fish disease identification system based on deep learning according to claim 8, characterized in that: The video acquisition camera is mounted on a three-dimensional translation platform. The three-dimensional translation platform receives acquisition instructions forwarded by the cloud server and controls the video acquisition camera to move on the three-dimensional translation platform to change the shooting angle to avoid environmental interference factors. Alternatively, the video acquisition camera is deployed on a drone, which receives an acquisition instruction forwarded by a cloud server, flies to the target water area, and adjusts the shooting angle to acquire images.
10. The fish disease identification system based on deep learning according to claim 8, characterized in that: The front-end interface is also provided with a human-computer interaction area and a data display area. The human-computer interaction area is used to provide users with a drawing component; the data display area is used to visually display disease images and long-term disease statistical charts.
Citation Information
Cited By
Behavior specification detection method and device fusing semantic correlation between tags and medium
CN120747874A
Fry health state monitoring method based on image recognition
CN120787852A
Trachinotus ovatus image processing method and system based on machine vision
CN121236799A
A machine vision-based image processing method and system for jackfish
CN121236799B
Deep learning fish disease detection method, deep learning fish disease detection system and deep learning fish disease detection equipment based on unmanned aerial vehicle data fusion, and medium
CN121304686A