Method for detecting and classifying abnormal behaviors of fishes in underwater complex environment

By constructing a lightweight network structure adapted to underwater scenes, combined with a lightweight attention mechanism and hypergraph computing, the problems of image ambiguity and complex background interference in the detection and classification of abnormal underwater fish behavior are solved, achieving efficient and accurate monitoring of abnormal fish behavior and supporting the intelligent management of smart fisheries.

CN120673121APending Publication Date: 2025-09-19HUZHOU UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510660953.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies for detecting and classifying abnormal fish behavior in underwater environments suffer from image ambiguity, complex background interference, behavioral morphological diversity, and an imbalance between algorithm accuracy and efficiency, making it difficult to meet the real-time and precise monitoring needs of high-density aquaculture scenarios.

Method used

A lightweight attention mechanism and multi-scale feature fusion technology are used to construct a lightweight network structure. Combined with hypergraph computing and multi-scale motion perception, the feature extraction and classification accuracy are improved, and the YOLOv1n model is improved to adapt to underwater scenes.

Benefits of technology

It has achieved a leap from simple target positioning to behavioral semantic understanding, breaking through the limitations of traditional detection frameworks, providing a technical path for intelligent aquaculture monitoring systems, and improving the accuracy and real-time performance of abnormal fish behavior detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673121A_ABST
    Figure CN120673121A_ABST
Patent Text Reader

Abstract

The invention provides a method for detecting and classifying fish abnormal behaviors in an underwater complex environment, which comprises the following steps of: 1, performing data preprocessing and data enhancement on fish abnormal behavior images, and converting marked and classified fish abnormal behavior XML (Extensible Markup Language) files into normalized coordinate representation in a YOLO format; 2, a lightweight Mixed Local Channel Attention (MLCA) attention module is adopted, channel information and spatial information are considered at the same time, and local information and global information are combined to improve the expression effect of the network; 3, a C3k2-PKI Module module is adopted, so that local multi-scale details and global context information can be captured at the same time, features with higher discrimination capability are extracted, and the model is endowed with higher context adaptability; and 4, a HyperC2Net module and an MANet module are introduced, and complex high-order interaction is carried out between different layers and positions. The information flow is enriched; the high-order feature extraction capability of the neck is enhanced; and 5, the features of the underwater abnormal fishes are sent to YOLOV11n for fish abnormal behavior detection and classification. According to the method, the key problems of image fuzziness, complex background interference, behavioral form diversity, algorithm precision and efficiency imbalance and the like in the fish abnormal behavior detection and classification task are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer technology and aquaculture technology, in particular to the field of abnormal fish behavior detection and classification. Background Art

[0002] As the global aquaculture industry develops towards intensification and scale, fish health management faces significant challenges. Abnormal fish behavior can manifest early and visible signs, such as floating, capsizing, and agitation, due to environmental stressors such as insufficient dissolved oxygen, abnormal water temperature, and pH levels, pathogen infection, or physiological stressors such as insufficient food intake and metabolic disorders. Traditional manual monitoring methods rely on empirical judgment and suffer from low efficiency, poor real-time performance, and significant subjective bias, making them inadequate for modern high-density aquaculture scenarios.

[0003] In recent years, computer vision-based behavior recognition technology has demonstrated tremendous potential in aquatic life monitoring. Target detection algorithms, such as the YOLO family, have garnered significant attention due to their exceptional real-time performance. However, the unique characteristics of the underwater environment present multiple challenges for these algorithms: the absorption and scattering of light by water leads to blurred and low-contrast images; the dense movement of fish schools causes target overlap and occlusion; the scarcity of abnormal behavior samples leads to significant class imbalance; and, due to the limited computing resources of underwater equipment, conventional detection models struggle to meet practical requirements. Existing research often employs direct transfer of the basic YOLO architecture, often facing bottlenecks such as high small target miss rates, significant dynamic background interference, and insufficient behavioral feature extraction.

[0004] Therefore, intelligent fish abnormal behavior detection and classification technology has become the core research direction of the digital upgrade of aquaculture. Its goal is to achieve accurate, real-time and automated fish health monitoring through the integration of computer vision, deep learning and Internet of Things technologies, and provide technical support for disease prevention, water quality control and precision feeding. Summary of the Invention

[0005] Therefore, the technical problem to be solved by the present invention is the image fuzziness, complex background interference, behavioral morphological diversity, and imbalance between algorithm accuracy and efficiency that exist in the current fish abnormal behavior detection and classification tasks.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: a method for detecting and classifying abnormal behaviors of fish in complex underwater environments, which improves the expression effect of the network by deeply fusing the attention mechanism with multi-scale feature fusion technology, enhancing multi-scale motion perception and improving classification accuracy, and constructing a lightweight network structure adapted to underwater scenes. This method breaks through the limitations of traditional detection frameworks on image processing, and achieves a leap from simple target positioning to behavioral semantic understanding, providing a new technical path for building an intelligent aquaculture monitoring system, and has important practical value for preventing diseases and optimizing aquaculture management.

[0007] According to the above invention concept, the present invention adopts the following technical solutions:

[0008] A. Data Preprocessing: Label the fish abnormal behavior dataset using LabelImg software to generate an XML file. Read the XML file, convert the pixel coordinates of the target information in the file into normalized coordinates in the YOLO format, and write it to a TXT file.

[0009] B. Lightweight Attention Feature Fusion: Use the lightweight Mixed Local Channel Attention (MLCA) attention module, which considers both channel information and spatial information, and combines local information with global information to improve the expression effect of the network.

[0010] C. Enhanced multi-scale action perception and improved classification accuracy: The original C3k2 module in the YOLOv11n model backbone is replaced with the C3k2-PKIModule. The C3k2-PKIModule's multi-scale convolutional module can perceive behavioral features at different scales. This multi-scale feature extraction enables the model to adaptively capture abnormal features at different scales.

[0011] D. Hypergraph Computation-Based Feature Extraction: The YOLOv11n neck feature extraction network incorporates a hypergraph-based cross-level and cross-position representation network, HyperC2Net, and a hybrid aggregation network, MANet, to enhance feature extraction. HyperC2Net operates at five scales and breaks away from the traditional grid structure, allowing for complex, high-order interactions across levels and positions. MANet incorporates a cross-scale interaction module in the feature extraction stage, utilizing dilated and deformable convolutions to enhance spatial perception. It also incorporates a channel-attention mechanism to optimize the contribution weights of feature channels, effectively mitigating recognition biases caused by underwater image blur and target variations.

[0012] E. Use the YOLOv11n model for training and verification: Improve the hypergraph computing feature extraction method, adopt lightweight attention fusion, enhance multi-scale motion perception, and improve classification accuracy, and apply YOLOv11n to detect and classify abnormal fish behavior.

[0013] While the model is being trained, the model parameters are continuously adjusted to optimize its performance, ensuring that the model has both accuracy and speed in detecting fish anomalies. This invention method helps to prevent diseases and achieve smart fish farming, and is compared with other series of this model. At the same time, the proposed model is applied to cases such as underwater fish identification, water quality prediction, and fish feeding intensity analysis, indicating that the invention has a certain degree of feasibility in smart fisheries, and provides a practical technical support for the application of deep learning models in large-scale aquaculture scenarios. The specific network architecture is as follows: Figure 1 shown.

[0014] The beneficial effects of this invention include a novel lightweight network structure adapted for underwater scenarios, developed for detecting and classifying abnormal fish behavior in complex underwater environments. This system utilizes a deep integration of lightweight attention mechanisms, feature extraction techniques using hypergraph computing, enhanced multi-scale motion perception, and improved classification accuracy. This system overcomes the limitations of traditional detection frameworks for image processing, achieving a significant leap from simple target location to behavioral semantic understanding. This system provides a new technical approach for building intelligent aquaculture monitoring systems and has significant practical value for disease prevention and optimized aquaculture management. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a schematic diagram of the overall network architecture of a method for detecting and classifying abnormal behavior of fish in complex underwater environments provided by the present invention.

[0016] Figure 2 Schematic diagram of the overall network architecture of the method for detecting and classifying abnormal behavior of fish in complex underwater environments provided by the present invention Figure 1 Detailed description of the MLCA attention module.

[0017] Figure 3 Schematic diagram of the overall network architecture of the method for detecting and classifying abnormal behavior of fish in complex underwater environments provided by the present invention Figure 1 The detailed description diagram of the C3k2-PKIModule module.

[0018] Figure 4 Schematic diagram of the overall network architecture of the method for detecting and classifying abnormal behavior of fish in complex underwater environments provided by the present invention Figure 1 The following figure shows a detailed description of the HyperC2Net module.

[0019] Figure 5 Schematic diagram of the overall network architecture of the method for detecting and classifying abnormal behavior of fish in complex underwater environments provided by the present invention Figure 1 The detailed description diagram of the MANet module.

[0020] Figure 6This is the experimental result of a model for detecting and classifying abnormal behavior of fish in complex underwater environments provided by the present invention on a self-built data set. DETAILED DESCRIPTION

[0021] Reference Figures 1 to 5 ,This embodiment provides a method for detecting and classifying abnormal behavior of fish in complex underwater environments.

[0022] In the present invention, the machine running the experiment is a server cluster with an A100 graphics card, and the model is developed using the Pytorch development framework.

[0023] 1. Data preprocessing.

[0024] The dataset consists of videos of pufferfish. Labeled manually using LabelImg software, the labels were set to "health," "PH," "lowtemp," "hightemp," and "hypoxia," generating an XML file. Before model training, the XML data was converted to TXT format. Each row contains the target category ID, the center coordinates of the target's bounding box, and its width and height.

[0025] 2. Lightweight attention feature fusion.

[0026] The lightweight Mixed Local Channel Attention (MLCA) is an attention mechanism that fuses local channel interactions with spatial feature enhancements. It aims to improve the ability to capture both inter-channel dependencies and spatial details at a lower computational cost. Its core design includes a parallel branch structure.

[0027] In the channel dimension, local channel attention is used to dynamically assign weights to adjacent channels to avoid information loss caused by global channel compression. At the same time, lightweight one-dimensional convolution is used to achieve cross-channel interaction and enhance the ability to focus on key feature channels.

[0028] In the spatial dimension, lightweight depthwise separable convolutions are used to construct local spatial attention, capturing pixel-level spatial context and enhancing the representation of detailed features such as object edges and textures. The outputs of the two attention modules are adaptively weighted and fused to form a hybrid feature response map, collaboratively optimizing the representation of channel and spatial information while reducing the number of parameters.

[0029] In addition, MLCA further compresses computational complexity through group computing and parameter sharing mechanisms, allowing it to still run efficiently in embedded devices. This mechanism effectively solves the problem of local information loss caused by global channel compression in traditional attention models, and takes into account the balance between fine-grained feature extraction and computational efficiency. The overall structure of MLCA is as follows Figure 2 shown.

[0030] like Figure 2 The figure below illustrates the principle of MLCA. For the MLCA input feature vector, pooling is performed in two steps. The input is converted into a 1*C*ks*ks vector, and the first local pooling step extracts local spatial information. Building on the initial stage, the input is converted into a one-dimensional vector using two branches: the first branch contains global information, and the second branch contains local spatial information. After the one-dimensional convolution, unpooling is used to restore the original resolution of the two vectors, and then the information is fused to achieve mixed attention. Figure 2 Conv1d is a one-dimensional convolution, and the kernel size k is proportional to the channel dimension C. This means that when capturing local cross-channel interaction information, only the relationship between each channel and its k adjacent channels is considered. The choice of k is determined by the following formula.

[0031]

[0032] Where C is the number of channels, k is the size of the convolution kernel, γ and b are both hyperparameters with a default value of 2. odd means that k is only an odd number. If k is an even number, add 1.

[0033] 3. Enhance multi-scale motion perception and improve classification accuracy.

[0034] Abnormal fish behavior can manifest as changes in movement at different scales and frequencies. The C3k2-PKIModule's multi-scale convolutional module can perceive behavioral features at different scales. For example, small convolution kernel branches can capture subtle local movements, while large convolution kernel branches capture large-scale motion patterns. This multi-scale feature extraction makes the model more sensitive to behavioral changes of varying magnitudes and frequencies. Compared to traditional single-scale feature extraction, the PKI module can reduce the rate of missed detections of abnormal behavior. Furthermore, the C3k2-PKIModule combines local multi-scale features with global context, enabling richer extraction of fish morphology and motion trajectories. Furthermore, the CAA contextual attention mechanism in the C3k2-PKIModule integrates global information about the fish's surroundings and group behavior. The CAA module focuses on the overall movement patterns and environmental characteristics of the school of fish in the entire scene. When the behavior of a particular fish is inconsistent with the global pattern, the attention mechanism assigns higher weight to its features, highlighting the characteristic signals of abnormal behavior. As a result, the C3k2-PKI Module can extract more discriminative feature vectors, providing a more reliable basis for subsequent classification. The PKI module is an Inception module that consists of a small convolution kernel to obtain local information, followed by a set of depth-wise separable convolutions DWConv to capture contextual information across multiple scales. The PKI module can be expressed as follows:

[0035]

[0036] in, It is k s ×k s Local features extracted by convolution, is the mth k (m) ×k (m) ) The context features extracted by deep convolution DWConv. In this invention, we set k s =3,k (m) =(m+1)×2+1. The PKI module does not use dilated convolution to prevent the extraction of overly sparse feature representations. Local and contextual features are fused through 1×1 convolution. The formula is as follows:

[0037]

[0038] in Represents the output feature. 1×1 convolution is used as a channel fusion mechanism to integrate features with different acceptance sizes. In this way, our PKI module can capture a wide range of contextual information without affecting the integrity of local texture features. The C3k2-PKIModule module structure is shown in the attached figure. Figure 3 shown.

[0039] 4. Feature extraction based on hypergraph computing.

[0040] (1) HyperC2Net module

[0041] Due to the complexity of the underwater environment, such as water refraction, reflection, and changes in lighting conditions, the captured fish images are often blurred, further weakening the ability to identify the fish's outline and motion trajectory. Therefore, this study introduces the HyperC2Net module into the feature extraction network to adjust the network parameters in real time to cope with lighting and background interference, while capturing local action details and global motion trends, and fully integrating cross-level and cross-position information from the backbone network. The HyperC2Net module structure is shown in the attached figure. Figure 4 shown.

[0042] (2) MANet module

[0043] The architecture synergistically fuses three typical convolution variants: 1×1 bypass convolution for channel feature recalibration, depthwise separable convolution (DSConv) for efficient spatial feature processing, and C2f module for enhanced feature hierarchical integration. This fusion produces a more diverse and rich gradient flow during the training phase, which significantly amplifies the semantic depth encapsulated in each basic feature in the five key stages. Our MANet can be formulated as follows:

[0044]

[0045] where X mid The channel number is 2c. And X1, X2, ..., X 4+n The feature channel count is c. Finally, we fuse and compress the semantic information of these three types of features through a concatenation operation, and the 1×1 convolution generates X with the number of channels. out。 The MANet module structure diagram is as follows Figure 5 shown.

[0046] In summary, the present invention's method for detecting and classifying abnormal fish behavior in complex underwater environments introduces attention feature fusion technology to enhance multi-scale motion perception and improve classification accuracy, improves feature extraction methods based on hypergraph computing, and establishes a more accurate model for detecting and classifying abnormal fish behavior. This model can effectively address key issues in the detection and classification of abnormal fish behavior, such as image ambiguity, complex background interference, behavioral morphological diversity, and imbalances between algorithm accuracy and efficiency. This invention provides a new technical approach for building intelligent aquaculture monitoring systems and has important practical value for preventing diseases and optimizing aquaculture management.

Claims

1. A method for detecting and classifying abnormal behavior of fish in complex underwater environments, the method comprising the following steps: A. Perform data preprocessing and data enhancement on fish abnormal behavior images, and convert the marked and classified fish abnormal behavior XML files into normalized coordinate representations in YOLO format; B. Attention feature fusion; C. Enhance multi-scale motion perception and improve classification accuracy; D. Feature extraction based on hypergraph computing; E. Use the YOLOV11n model for training and verification.

2. The method according to claim 1, wherein the preprocessing step includes preprocessing the underwater fish image. The data preprocessing and data enhancement steps are as follows: A. Use LabelImg software to label unusual fish and generate an XML file. Read the XML file, convert the pixel coordinates of the target information in the file into normalized coordinates in YOLO format, and write it to a TXT file. B. Using the deep convolutional generative adversarial network DCGAN, the pooling layer is eliminated and replaced with strided convolution, making the overall network model differentiable.

3. According to the method in claim 1, a lightweight Mixed Local Channel Attention (MLCA) attention module is introduced to simultaneously consider channel information and spatial information, and combine local information and global information to improve the expression effect of the network.

4. According to the method described in claim 1, the backbone part adopts the C3k2-PKI Module module, which can simultaneously capture local multi-scale details and global context information, extract more discriminative features, and give the model stronger context adaptability.

5. According to the method of claim 1, the neck part introduces HyperC2Net and MANet modules, which perform complex high-order interactions between different layers and positions. This enriches the information flow and enhances the ability to extract high-order features in the neck.

6. According to the method of claim 1, the characteristics of abnormal underwater fish are sent to YOLO V11n to detect and classify abnormal fish behavior.

Citation Information

Cited By

  • Information processing method, device and equipment for fish hemorrhagic disease detection and medium

    CN120876481A