A method of monitoring, describing and diagnosing abnormal behavior of fish

By using a backbone network-hybrid encoder-decoder and detection head connected network architecture, the problems of feature fusion and scale adaptation in the detection of abnormal behavior of underwater fish are solved, and the accurate detection and stable monitoring of abnormal fish behavior are achieved.

CN120783398BActive Publication Date: 2025-11-11ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511299312.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-11-11
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing methods for monitoring abnormal fish behavior struggle to accurately detect such behavior in underwater environments with unclear and complex imaging, especially for fish with irregular body shapes and complex underwater contours. Furthermore, they cannot effectively integrate coarse-grained and fine-grained features, resulting in poor model generalization and low detection accuracy.

Method used

A backbone network-hybrid encoder-decoder and detection head cascaded network architecture is adopted to construct an abnormal behavior feature extraction network with rich branch aggregation pixel perception. A detail enhancement module with parallel pixel fusion attention and a cross-scale abnormal small target decoupling detection head with dynamic weight allocation are introduced to improve feature representation capability and robustness.

Benefits of technology

It achieves accurate detection of abnormal fish behavior, improves the stability and robustness of the model in complex underwater environments, and can adapt to the detection of abnormal fish behavior at different scales and locations. It is particularly suitable for complex behavior perception tasks with multiple scales and attributes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783398B_ABST
    Figure CN120783398B_ABST
Patent Text Reader

Abstract

This invention discloses a method for monitoring, describing, and diagnosing abnormal fish behavior. The invention utilizes a rich-branch, pixel-aware abnormal behavior feature extraction network in the backbone network, improving its ability to represent features of irregularly shaped parts of the fish body and complex underwater contours. A detail enhancement module with parallel pixel fusion attention is introduced into the hybrid encoder to independently enhance important features of each sub-region, while ensuring natural transitions between adjacent blocks during sub-feature map fusion. A dynamically weighted, cross-scale decoupled detection head for abnormal small targets is constructed in the decoder and detection head modules, enabling accurate detection of abnormal targets with significantly off-center underwater shooting angles and immature, small-sized fish species. This invention is suitable for the rapid capture of abnormal behavioral features in fish, such as rollover, rapid convulsions, and disease, successfully improving the speed and accuracy of abnormal fish behavior monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision detection technology, and more specifically, to a method for monitoring, describing, and diagnosing abnormal behavior in fish. Background Technology

[0002] As an integral part of animal consumption products, fish have always occupied an important position. Their flesh is delicious and rich in protein and other nutrients, offering significant benefits to human health. Currently, fish farming is mainly based on large-scale, high-density factory farming. In the context of large-scale farming, timely detection of abnormal behaviors in farmed fish, such as rollover, rapid convulsions, and disease, is crucial for managing potential fish infectious diseases and diagnosing whether the chemical composition and physical conditions of the farming environment are suitable.

[0003] Currently, deep learning technology has been widely applied in several key fields such as computer vision and speech recognition. In computer vision, object detection is considered a core and challenging task. This task aims to accurately identify all objects of interest from static images or dynamic videos. Object detection not only needs to classify objects within an image but also accurately predict their location. This comprehensive task places high demands on both the accuracy and efficiency of the algorithms.

[0004] Due to the unclear and complex underwater environment imaging, and the diversity of behavior and body size of the fish being examined, using deep learning technology to monitor, describe, and diagnose abnormal fish behavior still faces many challenges, such as: ① Most machine vision detection methods often rely on single visual features, monitoring only the fish's movement trajectory, speed, or direction, ignoring external features such as surface damage and body distortion. Furthermore, these machine vision detection methods only detect abnormal fish behavior without providing accurate descriptions and diagnostic strategies; ② Abnormal fish behavior detection tasks often involve unbalanced or uncoordinated swimming postures, and deviation from the normal activity layer. Coarse-grained features and changes in tail fin wagging frequency, as well as fine-grained features such as dull skin, spots, and edema, are not typically distinguished or combined by ordinary object detection models. This leads to poor model generalization and limitations in practical applications. ③ Different species of farmed fish have significant differences in body size. For the same disease, variations in morbidity and the degree of change in affected areas can cause differences in the scale of the features to be detected. Ordinary object detection models cannot dynamically adjust the receptive field or sampling density based on the image. For large targets, this can easily lead to prediction box shifting or breakage, while for small targets, they are easily lost in convolutional layers near the input. Summary of the Invention

[0005] 1. The technical problem that the invention aims to solve

[0006] In view of the shortcomings of the existing technology, this invention provides a method for monitoring, describing, and diagnosing abnormal fish behavior. This invention employs a backbone network-hybrid encoder-decoder-detector connected network architecture with a detection head, and designs a target detection model for abnormal fish behavior. An abnormal behavior feature extraction network using rich branch aggregation pixels is constructed in the backbone network, improving the backbone network's feature representation ability for irregularly shaped parts of the fish body surface and complex underwater fish contours. A detail enhancement module with parallel pixel fusion attention is introduced into the hybrid encoder. By independently enhancing the important features of each sub-region, and ensuring a natural transition of features between adjacent blocks during sub-feature map fusion, this attention mechanism can fully reference all spatial and channel features, thereby improving the network's ability to recognize fish behavior in different environmental locations and its monitoring stability in complex underwater environments. A decoupled detection head for abnormal small targets across scales was constructed by dynamically weighting the decoder and detection head modules, maintaining the end-to-end semantic mechanism. At the same time, when performing feature fusion after the decoder structure, learnable parameters were used to adjust the weights for all feature fusion layers. The cross-scale connection method was used to enhance the effect of high-low layer feature fusion. This enabled the model to accurately detect abnormal targets with relatively deviated underwater shooting angles and immature, small-sized fish species, improving the robustness and effectiveness of the model in real-world application environments.

[0007] 2. Technical Solution

[0008] To achieve the above objectives, the technical solution provided by the present invention is as follows:

[0009] The present invention provides a method for monitoring, describing, and diagnosing abnormal behavior in fish, comprising the following steps:

[0010] S110: Acquire low-resolution images of fish and perform image preprocessing. Label the preprocessed images with individual fish exhibiting abnormal behavior, establish an abnormal fish dataset, and divide it into training set, validation set, and test set.

[0011] S120: Construct an abnormal fish behavior target detection model, and input the abnormal fish dataset into the target detection model for model training; the target detection model is obtained by concatenating a backbone network, a hybrid encoder, a decoder, and a detection head;

[0012] S130: Input the real-time acquired images into the trained fish abnormal behavior target detection model, and use the model output information as data samples for visual and text prompts. Input the fish abnormal behavior description and diagnostic language model for processing, and output the fish abnormal behavior description and diagnostic language.

[0013] Furthermore, in step S120, in the target detection model:

[0014] A rich branch aggregation pixel-aware abnormal behavior feature extraction network is used as the backbone network;

[0015] The hybrid encoder includes two feature recovery modules, two feature fusion modules, multiple reparameterized convolutional modules, and multiple parallel pixel fusion attention detail enhancement modules. The two feature recovery modules in the hybrid encoder are respectively connected to the output layer of the backbone network. The two feature recovery modules are connected in series. The outputs of the two feature recovery modules are respectively fed into the two series-connected feature fusion modules. The final outputs of the two series-connected feature fusion modules and the individual outputs of each feature fusion module are fed into the reparameterized convolutional modules. The reparameterized convolutional modules are then connected in series with the detail enhancement modules.

[0016] The hybrid encoder and decoder are spliced ​​together, and the output of the decoder is mapped to shallow feature maps of different resolutions through linear transformation and reconstruction operations, which are then fed into the cross-scale abnormal small target decoupling detection head with dynamic weight allocation.

[0017] Furthermore, the abnormal behavior feature extraction network includes multiple cascaded processing units, each of which contains a segmentation padding layer and multiple cascaded gated aggregation attention modules; layer normalization layers are set between adjacent processing units; and the hybrid encoder receives the output of the layer normalization layers as its input.

[0018] Furthermore, the segmentation filling layer is obtained by concatenating an edge verification layer, a convolutional layer, and a normalization layer, wherein the edge verification layer performs position encoding operations;

[0019] The gated attention aggregation module includes an input layer, a layer normalization layer, a dual-path attention aggregation module, a residual connection, a layer normalization layer, a gated convolution module, a residual connection, and an output layer, which are connected in series.

[0020] Furthermore, the dual-path attention aggregation module is designed with two types of path branches, including a convolutional branch and an embedding-scaling attention branch; the convolutional branch uses a depthwise separable convolution and dynamic encoding scheme that includes depthwise convolution and pointwise convolution; the embedding-scaling attention branch uses attention computation that includes a length scaling factor and a temperature scaling factor.

[0021] The gated convolutional module includes an input layer, a dual-parallel branch structure, a gated fusion layer, a linear transformation layer, a residual connection, and an output layer, all connected in series. The dual-parallel branch structure consists of a left branch and a right branch. The left branch extracts local spatial features by passing through a linear transformation layer, a depthwise separable convolutional layer, and an activation function layer. The right branch contains only a linear transformation layer. The outputs of the two paths are fused through the gated fusion layer.

[0022] Furthermore, the detail enhancement module performs attention-weighted fusion of the two input feature maps, including: fusing the input feature maps element-wise to obtain an initial feature map; using a channel-oriented attention module to model the initial feature map in terms of channel dimensions to generate a channel attention map; using a spatial attention module to generate a spatial attention map by combining average pooling and max pooling; adding the channel attention map and the spatial attention map and broadcasting them to form a joint attention map; inputting the joint attention map to a pixel-oriented attention module to generate a pixel-wise fusion weight map; achieving dynamic weighted fusion of the input feature maps through the fusion weight map; and performing spatially enhanced convolution processing through a slice fusion convolution module.

[0023] Furthermore, the channel-oriented attention module includes an input layer, a global average pooling layer, a dimensionality-reduced 1x1 convolutional layer, a ReLU activation function layer, an up-dimensional 1x1 convolutional layer, and an output layer, which are connected in series.

[0024] The spatial attention-oriented module includes an input layer, a channel average pooling layer, a channel max pooling layer, a splicing layer, a 7x7 convolutional layer with reflection fill, and an output layer, which are connected in series.

[0025] The pixel-oriented attention module includes an input layer, a splicing layer, a dimension transformation layer, a 7x7 depthwise convolutional layer, and an output layer, which are connected in series.

[0026] Furthermore, the slice fusion convolutional module includes a convolutional layer, a batch normalization layer, a SiLU activation function, and a feature space slice fusion module connected in series; wherein: the specific processing steps of the feature space slice fusion module on the input feature map are as follows:

[0027] The input feature map is divided into four sub-blocks along the spatial dimension. The average pixel value of each sub-block is calculated and the variance is normalized to obtain the standardized feature distribution. The normalization result is mapped to the range [0,1] using the Sigmoid function to obtain the attention weight of the sub-block. Each sub-block is multiplied by its corresponding attention weight to obtain the weighted sub-block. Finally, all the weighted sub-blocks are concatenated by tensors according to their original spatial positions and recombined into the output feature map.

[0028] Furthermore, the dynamic weight allocation cross-scale abnormal small target decoupling detection head includes a spatial dynamic fusion and aggregation module and a multi-scale feature selection fusion detection head. The spatial dynamic fusion and aggregation module includes a spatial dynamic weighted fusion module. The spatial dynamic fusion and aggregation module first performs a spatial dynamic weighted fusion on the input first shallow feature map and the upsampled result of the second shallow feature map, and the downsampled result of the first shallow feature map and the second shallow feature map, respectively. Then, it performs another spatial dynamic weighted fusion on the third shallow feature map and the spatial dynamic weighted fusion result of the two parts. Finally, it performs a residual connection operation on the three spatial dynamic weighted fusion results to output a first deep feature map, a second deep feature map, and a third deep feature map.

[0029] The classification head of the multi-scale feature selection fusion detection head is used to predict the category, and the regression head is used to predict the offset of the target detection anchor box. The two receive the fusion result as input, and the final output is the concatenation result of the outputs of the classification head and the regression head.

[0030] Furthermore, in step S130, the open-source Llama 3.1 large language model is used as the language model for describing and diagnosing abnormal fish behavior. Based on the results of the target detection model, the user-uploaded problem descriptions and professional fish diseases, the language model breaks down the abnormal fish diagnosis and treatment strategy into question-answer pairs, segments long texts according to semantics or paragraphs, integrates the answers, and stores the generated question-answer pairs and text fragments as knowledge units in the PostgreSQL database as the knowledge basis for subsequent answers. At the same time, it connects to the knowledge base of professional aquaculture personnel and the knowledge base of fish disease diagnosis and treatment experts to generate a knowledge base for describing and diagnosing abnormal fish behavior. The model is trained through reinforcement learning using feedback from experts in the field of fishery diseases to optimize the output quality of the model.

[0031] 3. Beneficial effects

[0032] Compared with existing known technologies, the technical solution provided by this invention has the following significant advantages:

[0033] (1) The rich branch aggregation pixel perception abnormal behavior feature extraction network proposed in this invention has a gated aggregation attention module, which includes a dual-path attention aggregation module and a gated mechanism convolution module. The dual-path attention aggregation module is designed with convolution branches and embedding-scaling attention branches, which effectively integrate different attention branches and learnable convolution modules, enabling the model to competitively select between fine-grained and coarse-grained features, realizing the joint detection of coarse-grained behavior and fine-grained lesions in fish, and improving the accuracy, stability and cross-fish species generalization ability of abnormal behavior detection.

[0034] (2) The parallel pixel fusion attention detail enhancement module proposed in this invention performs attention-weighted fusion of two input feature maps. It uses a channel-oriented attention module to enhance useful channels and suppress useless channels, thereby improving the model's ability to perceive the location features of abnormal fish behavior. It also uses a spatial attention module to compress channel information and only uses spatial location information. It splices feature maps compressed by different pooling methods to analyze complete spatial structure information. It extracts spatial attention in a larger receptive field through 7×7 convolution, enabling the model to capture the spatial changes of fish group movement patterns and effectively suppress the response of irrelevant regions such as background interference and noise regions. The pixel-oriented attention module performs channel dimension splicing and depth convolution to extract local context relationships between feature maps and joint attention. It uses Sigmoid mapping to generate pixel-wise fusion weights, which realizes the fine discrimination of selecting the better feature map from the two source features for each pixel. The fused features are then processed by a slice fusion convolution module for spatial enhancement. While independently enhancing the important features of each sub-region, it ensures the natural transition of features between adjacent blocks when the sub-feature maps are fused, thereby improving the model's ability to perceive multi-scale features of abnormal fish behavior targets.

[0035] (3) This invention constructs a cross-scale abnormal small target decoupling detection head with dynamic weight allocation. Its spatial dynamic weighted fusion module first compresses the feature maps of each scale to obtain the attention weight base vector. After concatenating the attention weight base vector, the fusion weight is generated. By spatial attention weighted fusion of features at different scales, the model can adaptively focus on the key areas and key scales of abnormal fish behavior. The multi-scale feature selection fusion detection head includes multiple regression heads and classification heads, which effectively improves the robustness and accuracy of abnormal fish behavior detection for different sizes, locations and types of abnormalities. It is particularly suitable for complex behavior perception tasks with multiple scales and attributes.

[0036] (4) This invention constructs a language model for describing and diagnosing abnormal fish behavior. The detection results of the trained fish abnormal behavior target detection model are used as data samples for visual and textual prompts of the language model. The output results of the language model are fine-tuned using a supervised expert scoring method, and the visual detection results are converted into natural language descriptions. It can better infer the potential causes behind abnormal fish behavior based on multimodal prompts, realize automated, interpretable, and expert-level behavior diagnosis, and realize the connection between perception and decision-making in the intelligent fish farming system. Attached Figure Description

[0037] Figure 1 This is a network structure diagram of the abnormal behavior target detection model for fish in this invention;

[0038] Figure 2This is a block diagram of the structure of a single processing unit in the abnormal behavior feature extraction network of this invention;

[0039] Figure 3 This is a structural block diagram of the dual-path attention aggregation module in the abnormal behavior feature extraction network of the present invention;

[0040] Figure 4 This is a structural block diagram of the gated mechanism convolutional module in the abnormal behavior feature extraction network of the present invention;

[0041] Figure 5 This is a structural block diagram of the detail enhancement module for parallel pixel fusion attention according to the present invention;

[0042] Figure 6 This is a structural block diagram of the channel-oriented attention module in the detail enhancement module of the present invention;

[0043] Figure 7 This is a structural block diagram of the spatial attention module in the detail enhancement module of the present invention;

[0044] Figure 8 This is a structural block diagram of the pixel-oriented attention module in the detail enhancement module of the present invention;

[0045] Figure 9 This is a schematic diagram of the feature space slice fusion module in the slice fusion convolution module of the detail enhancement module of the present invention;

[0046] Figure 10 This is a structural block diagram of the cross-scale abnormal small target decoupling detection head with dynamic weight allocation in this invention;

[0047] Figure 11 This is a flowchart illustrating the method for monitoring, describing, and diagnosing abnormal fish behavior in this invention. Detailed Implementation

[0048] To address the unclear and complex imaging of underwater environments and the diversity of fish behavior and body size, this invention provides a method for monitoring, describing, and diagnosing abnormal fish behavior. This invention employs a backbone network-hybrid encoder-decoder cascaded network architecture, designing a target detection model for abnormal fish behavior. An abnormal behavior feature extraction network using rich branch aggregation pixels is constructed within the backbone network, improving its feature representation capabilities for irregularly shaped parts of the fish body and complex underwater fish contours. A parallel pixel fusion attention detail enhancement module is introduced into the hybrid encoder. By independently enhancing important features of each sub-region, and ensuring natural transitions between adjacent blocks during sub-feature map fusion, this attention mechanism can fully reference all spatial and channel features, thereby improving the network's ability to identify fish behavior in different environmental locations and enhancing the monitoring stability in complex underwater environments. A decoupled detection head for abnormal small targets across scales is constructed by dynamically weighting the decoder and detection head modules, maintaining the end-to-end semantic mechanism. At the same time, when performing feature fusion after the decoder structure, learnable parameters are used to adjust the weights for all feature fusion layers. The cross-scale connection method enhances the effect of high-low layer feature fusion, enabling the model to accurately detect abnormal targets with relatively deviated underwater shooting angles and immature, small-sized fish species, thus improving the robustness and effectiveness of the model in real-world application environments.

[0049] This invention is applicable to the rapid capture of abnormal behavioral characteristics in fish, such as side-rolling, rapid convulsions, and disease, and successfully improves the speed and accuracy of monitoring abnormal fish behavior. It also features high efficiency and accuracy in outputting detection results for low-precision fish images taken in turbid water.

[0050] To further understand the content of this invention, a detailed description of the invention will be provided in conjunction with the accompanying drawings and embodiments.

[0051] Example 1

[0052] This embodiment of a method for monitoring, describing, and diagnosing abnormal fish behavior includes the following steps:

[0053] S110: Acquire low-resolution images of farmed fish from a monitoring perspective and perform image preprocessing, including scaling the images to 640×640 pixels, padding the remaining edges of the images with 114 grayscale values, and normalizing the padded images. Label the preprocessed images with individual fish exhibiting abnormal behavior, establish an abnormal fish dataset, and divide it into training, validation, and test sets.

[0054] Specifically, the low-resolution images of farmed fish from the monitoring perspective are captured underwater by an underwater aquaculture monitoring IP camera. Farmed fish include, but are not limited to, crucian carp, grass carp, and mandarin fish. In this embodiment, crucian carp in a farming environment is used as an example. Image scaling refers to scaling the longest side of the image to the target size of 640 pixels, while scaling the short side at the same ratio, maintaining the original aspect ratio. The filling step refers to placing the scaled image at the center of a canvas of the target size (640×640). The remaining space on both sides of the short side is filled with gray (RGB values ​​are typically (114, 114, 114)). Image normalization refers to normalizing the image size and grayscale values. The preprocessed image has a length of 640 pixels, a width of 640 pixels, and 3 color channels. The aforementioned abnormal fish dataset refers to dividing the preprocessed images into a training set, a validation set, and a test set in a 7:2:1 ratio.

[0055] S120: Construct an abnormal fish behavior target detection model. Input the aforementioned abnormal fish dataset into the target detection model for model training.

[0056] Combination Figure 1 The target detection model is obtained by connecting a backbone network, a hybrid encoder, a decoder, and a detection head in series.

[0057] Specifically, the constructed rich-branch aggregated pixel-aware abnormal behavior feature extraction network is used as the backbone network. The backbone network uses a multi-layered, cascaded gated aggregation attention module and a segmentation padding layer as a processing unit, with layer normalization layers between adjacent processing units, progressively extracting features from local pixels to deep abstract feature representations. The gated convolutional module and dual-path attention aggregation module within the aforementioned gated aggregation attention module are used to process hierarchical, multi-scale information.

[0058] Let the input feature map size be H×W×C, and the numbers of the cascaded gated aggregation attention module and segmentation padding layer are as follows: i Then the size of the feature map after the change for each processing unit is:

[0059]

[0060] Total output , , The feature map is used as the feature extraction map output by the backbone network, and as the input feature map of the hybrid encoder.

[0061] The hybrid encoder includes two feature recovery modules with reparameterized convolutional modules and residual connections, two feature fusion modules with reparameterized convolutional modules, multiple reparameterized convolutional modules, and multiple parallel pixel fusion attention detail enhancement modules. The two feature recovery modules in the hybrid encoder are connected to the layer normalization layer in the backbone network via a 1x1 convolutional layer. The two feature recovery modules are connected in series, and their outputs are fed into the reparameterized convolutional layers and concatenation layers of the two series-connected feature fusion modules, respectively. Finally, the outputs of the two series-connected feature fusion modules and their respective reparameterized convolutional layers are fed into three parallel reparameterized convolutional modules. These three parallel reparameterized convolutional modules are then connected in series with three parallel detail enhancement modules, forming the entire hybrid encoder structure.

[0062] The output of the reparameterized convolution module can be expressed as:

[0063]

[0064] in, This represents a positional addition operation of feature maps. Conv(·) represents a 1×1 convolution operation, and RepConv(·) represents a reparameterized convolution operation. The training formula for this reparameterized convolution model is as follows:

[0065]

[0066] in, , These are the weights of the convolutional module. , , The learnable parameters in the batch normalization layer are initialized using a normal distribution and updated during backpropagation. The weights of the reparameterized convolutional module and the learnable parameters in the batch normalization layer are initialized using a normal distribution.

[0067] The feature recovery module contains a reparameterized convolution module structure, which can be represented as follows:

[0068]

[0069] in, X 1, X 2 represents the layer normalized output received from the backbone network and the input received from the upper layer of the feature recovery module, respectively.

[0070] The feature fusion module contains a reparameterized convolution module structure, which can be represented as follows:

[0071]

[0072] in, X 1,X 2 represents the output received from the 1×1 convolution module of the feature recovery module and the input received from the upper layer of the feature fusion module, respectively.

[0073] The hybrid encoder introduces a parallel pixel fusion attention detail enhancement module at the output position of the three reparameterized convolutional modules. The reparameterized convolutional modules can deeply mine the detailed information in the input feature map through multiple convolution operations. The detail enhancement module uses the reparameterized convolutional modules to integrate the output feature information from different scales, thereby enhancing the model's ability to perceive abnormal fish movements at different scales.

[0074] The hybrid encoder and decoder are spliced ​​together, and the result is fed into the constructed dynamic weight allocation cross-scale abnormal small target decoupled detection head. This allows features at different levels to automatically adjust their fusion weights according to the content of the input feature map, reducing information loss and enhancing the detection capability of smaller targets, especially detailed parts such as fish fry or fins.

[0075] The completed target detection model combines a backbone network, a hybrid encoder, a decoder, and a detection head structure. By utilizing a gated aggregation attention module and a reparameterized convolution module, it achieves accurate detection of abnormal fish behaviors (such as fry aggregation and abnormal fin wagging). Furthermore, it optimizes small target detection through dynamic weight adjustment, providing an efficient method for monitoring abnormal behaviors.

[0076] The structure of the core module and image processing flow of the target detection model described in this embodiment are described in detail below:

[0077] Combination Figure 2 The abnormal behavior feature extraction network comprises a gated aggregation attention module and a segmentation and padding layer. The segmentation and padding layer divides the feature image into small blocks and maps each block to a single vector. The segmentation and padding layer is obtained by concatenating an edge verification layer, a convolutional layer, and a normalization layer, where the edge verification layer performs positional encoding. The specific structure of the gated aggregation attention module is as follows: input layer → layer normalization → dual-path attention aggregation module → residual connection → layer normalization → gated convolutional module → residual connection → output layer. This gated aggregation attention module effectively integrates different attention branches and learnable convolutional modules, enabling the model to competitively select between fine-grained and coarse-grained features, and effectively handle multi-level and multi-scale information.

[0078] Combination Figure 3 The dual-path attention aggregation module is designed with two types of path branches: convolutional branches and embedding-scaling attention branches. The module structure of the convolutional branch is as follows:

[0079]

[0080] in, It is the pixel coordinate information of the input feature map. It is the kernel size. It is a depthwise separable convolution, which includes depthwise convolution and pointwise convolution. SiLU is the selected activation function. It is a dynamic encoding scheme. These are the feature values ​​of the obtained convolutional branch feature map.

[0081] The formula for the dynamic coding scheme is as follows:

[0082]

[0083]

[0084] in, The total number of input positions is given by the following formula:

[0085]

[0086] The total number of pooling positions is given by the following formula:

[0087]

[0088] Where r refers to the space reduction scaling factor, in this embodiment ; , These are the normalized relative distances of the currently processed pixel in the height and width directions, respectively. It is a normalized relative distance coordinate table that includes all directions. This refers to the multilayer perceptron transformation, where the activation function used is ReLU; This refers to the feature space lifting matrix. For the first layer bias, , ; It is a convolutional attention head-specific projector. For the second layer bias, , ;in The hyperparameter is set to 4 in this embodiment to represent the number of attention heads for convolution. and The random initialization is performed by truncating the normal distribution, and the optimization and update are carried out during the backpropagation of the convolutional branch. It is the output of the dynamic encoding scheme.

[0089] The module structure formula for the embedded-scale attention branch is:

[0090]

[0091] in, This represents the element-wise multiplication of two matrices (or tensors) of the same dimension. express Norm normalization, This indicates the use of the Softmax function. The function indicates the use of the Softplus function. It is a temperature scaling factor. It is a learnable parameter vector, initialized to S is the length scaling factor. Indicates the kernel size. This represents the total number of pooling positions. It is the output of the dynamic encoding scheme. The output feature values ​​for the module structure of the embedding-scale attention branch. Represents the original query vector. This represents the original key vector. It is a learnable query embedding factor.

[0092] Combination Figure 4 The gated convolutional module structure is as follows: input layer Dual parallel branch structure Gated Fusion Layer Linear Transformation Layer Residual connection The output layer features a dual-parallel branch structure consisting of a left branch and a right branch. The left branch sequentially passes through a linear transformation layer, a depthwise separable convolutional layer, and an activation function layer using GeLU to extract local spatial features. The right branch contains only a single linear transformation layer. The outputs of the two paths are fused through a gated fusion layer, which performs element-wise multiplication on the spatial features of the left branch output and the dynamic weights of the right branch output, thereby enhancing the key features.

[0093] Combination Figure 5 The detail enhancement module introduced in the hybrid encoder performs attention-weighted fusion of the two input feature maps and further enhances feature representation capabilities through attention-guided convolution. Its input consists of two feature maps, where x is a shallow feature map and y is a deep feature map. The output is a fused and enhanced feature map. This module first fuses the two input feature maps element-wise to obtain an initial feature map. Subsequently, a channel-oriented attention module was used to... F 0. Perform channel-dimensional modeling to generate channel attention maps. It uses a spatial attention module to generate a spatial attention graph by combining average pooling and max pooling. The two attention maps mentioned above are added together and broadcast to the same dimension, forming a joint attention map. The data is then input into the pixel-oriented attention module, where a fine-grained fusion weight map is generated through pixel-by-pixel modeling. Each pixel value is normalized to 1 Range. Next, dynamic weighted fusion of the input feature maps is achieved by fusing weight maps, and the calculation formula is as follows:

[0094]

[0095] in The feature responses from shallow feature maps were emphasized. Information from the deep feature maps is preserved. Finally, the fused features... Spatially enhanced convolution processing is performed using a slice fusion convolution module.

[0096] Specifically, the detail enhancement modules include:

[0097] The channel-oriented attention module compresses the feature map space to capture the importance of each channel, then generates attention weights for each channel, enhancing useful channels and suppressing useless channels, thereby improving the model's ability to perceive the location features of abnormal fish behavior. For example... Figure 6 As shown, the channel-oriented attention module structure is as follows: Input layer → Global average pooling layer → 1×1 convolutional layer (reduced dimensionality) → ReLU → 1×1 convolutional layer (upgraded dimensionality) → Output layer. This module outputs the attention value for each channel.

[0098] The spatial attention module compresses channel information, utilizing only spatial location information. It concatenates feature maps compressed using different pooling methods to analyze complete spatial structure information. Spatial attention is extracted over a larger receptive field using 7×7 convolutions, enabling the model to capture spatial variations in fish school movement patterns and effectively suppress responses from irrelevant regions such as background interference and noise. Figure 7 As shown, the spatial attention module structure is as follows: input layer → channel average pooling layer → channel max pooling layer → concatenation → 7×7 convolutional layer with reflection padding → output layer. This module outputs the attention value for each spatial pixel.

[0099] The pixel-oriented attention module performs channel-dimensional concatenation and depthwise convolution to extract local contextual relationships between feature maps and joint attention. It then uses a sigmoid map to generate pixel-wise fusion weights, enabling fine-grained discrimination to select the superior feature map from two source features for each pixel. For example... Figure 8As shown, the pixel-oriented attention module structure is as follows: input layer (containing two input feature maps) → concatenation → dimensionality transformation layer → 7×7 depthwise convolutional layer → output layer. The module outputs a pixel-wise generated fused weight map.

[0100] The operation steps for the feature map of each channel in this embodiment are as follows:

[0101]

[0102] in, It is the original input feature map. It is the fused channel-oriented attention graph and space-oriented attention graph, A p It is the final output feature map for pixel-oriented attention; The symbol represents the Concat concatenation operation, Rearrange represents the dimension transformation operation, and Sigmoid represents the Sigmoid activation function.

[0103] The slice fusion convolutional module consists of convolutional layers, batch normalization layers, SiLU activation functions, and a feature space slice fusion module. First, convolutional layers are used to extract spatial features from the input X; batch normalization stabilizes training and accelerates convergence; the SiLU activation function introduces non-linearity, enhancing the model's ability to express complex patterns; finally, the feature space slice fusion module further fuses and strengthens channel and spatial information, improving the quality of the feature response. The specific operation can be represented as follows:

[0104]

[0105] Where Y represents the feature map after processing by the slice fusion convolution module, Lightnon_paraAtten(·) is the feature space slice fusion module, SiLU(·) is the SiLU activation function, and BN is the batch normalization operation.

[0106] like Figure 9 As shown, the feature space slicing fusion module independently enhances the important features of each sub-region while ensuring a natural transition of features between adjacent blocks during sub-feature map fusion, thereby improving the model's ability to perceive multi-scale features of abnormal fish behavior targets.

[0107] The specific processing steps of the feature space slicing and fusion module for the input feature map are as follows, assuming the input feature map tensor is... Where: B is This hyperparameter is typically set in the range [8, 64], and in this embodiment, it is set to 16. C is the number of channels, H is the number of vertical pixels in the feature image, and W is the number of horizontal pixels in the feature image. The input features... Figure X Divide the space into four sub-blocks on an even scale, with each sub-block having the following spatial dimensions:

[0108]

[0109] Define each The sub-blocks are as follows:

[0110]

[0111]

[0112]

[0113]

[0114] Next, each feature map sub-block is processed by the feature space slicing and fusion module. Specifically: First, the input feature map is divided into four sub-blocks, and the average value of all pixels within each sub-block is calculated to obtain the mean of that sub-block. Then, the variance of the sub-blocks is normalized to obtain the standardized feature distribution. The normalization result is mapped to the range [0,1] using the Sigmoid function to obtain the attention weight of that sub-block. Each sub-block is multiplied by its corresponding attention weight to obtain the weighted sub-block output. Finally, all weighted sub-blocks are concatenated according to their original position tensors and recombined into the output feature map. The result is a lightweight, parameter-free attention-enhanced sub-map. The tensor concatenation operation refers to joining multiple tensors end-to-end in dimensions such as channel, height, and width to merge them into a new tensor.

[0115] Combination Figure 10 In this embodiment, the dynamic weight allocation cross-scale abnormal small target decoupling detection head consists of a spatial dynamic fusion aggregation module and a multi-scale feature selection fusion detection head. In the decoder and detection head module, the decoder first uses the reparameterized convolution module and the detail enhancement module with parallel pixel fusion attention to output the constructed query vector. This vector is then mapped to three layers of feature maps with different resolutions through linear transformation and reconstruction operations, named the first shallow feature map, the second shallow feature map, and the third shallow feature map, respectively. The feature map sizes are (batch_size, 256, 80, 80), (batch_size, 256, 40, 40), and (batch_size, 256, 20, 20), respectively. The three layers of feature maps at different resolutions are fused using a spatial dynamic fusion aggregation module to perform multi-scale feature fusion, resulting in three fused feature maps of size [size missing]. The three fused feature maps are named the first deep feature map, the second deep feature map, and the third deep feature map, respectively, and denoted as... The three fused feature maps are then fed into a multi-scale feature selection fusion detection head. The `num_class` hyperparameter is set to 1, and the `reg_max` hyperparameter is set to 16. The final sizes of the three generated feature maps are as follows: .

[0116] The spatial dynamic fusion and aggregation module includes a spatial dynamic weighted fusion module. This module first compresses feature maps at each scale to obtain attention weight base vectors. These base vectors are then concatenated to generate fusion weights. These weights are then used to weight and fuse the inputs at the three scales to obtain a preliminary fusion result. Finally, this preliminary fusion result is integrated using a convolution to obtain the output of the spatial dynamic weighted fusion module. The specific formula is as follows:

[0117]

[0118] in, x i The input features are for the i-th layer. In this embodiment, there are a total of 3 layers. w i For the Softmax weights of each layer, y out Output results for the spatial dynamic weighted fusion module. This represents a 3×3 convolution operation. This represents the element-wise multiplication of two matrices (or tensors) of the same dimension.

[0119] The multi-scale feature selection fusion detection head contains multiple classification heads and regression heads; the classification head is used to predict the category, and the regression head is used to predict the offset of the target detection anchor box. Both receive the final fusion result as input, and the final output is the result of splicing the outputs of the classification head and the regression head.

[0120] Let the first shallow feature map, the second shallow feature map, and the third shallow feature map be represented as follows:

[0121]

[0122]

[0123]

[0124] The two-way fusion process of the first shallow feature map and the second shallow feature map can be represented as follows:

[0125]

[0126]

[0127] The two-way fusion process of the obtained two-way fusion and the third shallow feature map can be represented as:

[0128]

[0129] The final fusion result obtained after residual connection processing is shown below, where ResBlock represents the residual connection operation:

[0130]

[0131]

[0132]

[0133] For the regression head branch, it receives data at each scale. As input for predicting anchor boxes, it can be represented as:

[0134]

[0135] in It is the final output of the regression head branch, the final predicted bounding box. It can be represented as:

[0136]

[0137] Where i represents the scale level. Anchors are the coordinates of the center point of each location on the feature map, DFL(·) is the distributed regression operation, and dist2bbox(·) is the transformation that converts the distance prediction value into the actual bounding box coordinates based on the center point of the anchor box.

[0138]

[0139] Where s represents the step size, , Here, H represents the offset of the distributed regression in the x and y directions, H is the number of vertical pixels in the feature image, and W is the number of horizontal pixels in the feature image. , It is a scaling factor for distributed regression in the horizontal and vertical directions, which is initialized using a normal distribution and updated during backpropagation. , , , These represent the x-coordinate of the anchor frame's center point, the y-coordinate of the anchor frame's center point, the predefined width ratio of the anchor frame to the entire image, and the predefined height ratio of the anchor frame to the entire image, respectively.

[0140] For the classification head branch, each scale is received. As input for predicting the category, it can be represented as:

[0141]

[0142] in, This represents a 3×3 convolution operation. This indicates a 1×1 convolution operation, and Sigmoid indicates the use of the Sigmoid activation function.

[0143] The steps of obtaining the regression and classification header results, concatenating them, and then obtaining the final output can be represented as:

[0144]

[0145] in , , These are the outputs of the regression head. , , These are the outputs of the classification header, and Concat is the tensor concatenation operation.

[0146] The specific process for training the fish abnormal behavior target detection model is as follows: The `num_class` parameter is set to 1, representing abnormal fish. In this embodiment, the abnormal state is defined as 200 epochs. The SGD optimizer is used, and the `batch_size` is set to 8. Mosaic enhancement is used in the last 10 epochs of training. The initial learning rate is set to 0.01, and the momentum parameter is set to 0.937. The abnormal fish dataset (1988 images in total) obtained and segmented according to the steps described in S110 is input into the fish abnormal behavior target detection model for training.

[0147] S130: Use the detection results from the fish abnormal behavior target detection model trained above as data samples for visual and textual cues. Specific details are as follows: Figure 11As shown, this includes the number and abnormal state of abnormal individuals in the current detection frame. In this embodiment, the open-source Llama 3.1 large language model is used, along with the structured data samples described above, for language model processing. This language model processing includes breaking down the user-uploaded question description and professional fish disease and abnormal fish treatment strategies into question-answer pairs based on the target detection model's output, and segmenting the long text according to semantics or paragraphs. The answers are then integrated, and the generated question-answer pairs and text fragments are stored as knowledge units in a PostgreSQL database as the knowledge base for subsequent answers. Simultaneously, external knowledge bases such as professional aquaculture personnel knowledge bases and fish disease treatment expert knowledge bases are accessed to generate a fish abnormal behavior description and diagnosis knowledge base, thus constructing a fish abnormal behavior description and diagnosis language model. The model uses supervised fine-tuning to train the initial model, and then uses feedback from experts in the field of fishery diseases. For example, when the model answers, "There are 20 fish in this aquaculture area, of which 3 are rolling over. This may be due to low dissolved oxygen in the pond, viral or bacterial infection, etc. It is recommended to test the water first, then observe the appearance of the fish, analyze feeding records and water temperature changes, quickly determine the cause, and then prioritize treatment," the model is scored. The scoring results are used as reinforcement feedback to carry out reinforcement learning training and optimize the output quality of the model.

[0148] In this embodiment, the model was evaluated using a dataset of crucian carp images with white spot disease processed in step S110. The trained target detection model achieved an average precision (mAP0.5) of 91.7% at a threshold of 0.5 on the validation set. Experiments demonstrate that this invention is suitable for the rapid capture of abnormal behavioral features in fish, such as side-rolling and disease, successfully improving the speed and accuracy of abnormal fish behavior monitoring. It also features high efficiency and accuracy in outputting detection results for low-precision fish images taken in turbid water.

Claims

1. A method for monitoring, describing, and diagnosing abnormal behavior in fish, characterized in that, Includes the following steps: S110: Acquire low-resolution images of fish and perform image preprocessing. Label the preprocessed images with individual fish exhibiting abnormal behavior, establish an abnormal fish dataset, and divide it into training set, validation set, and test set. S120: Construct an abnormal fish behavior target detection model, and input the abnormal fish dataset into the target detection model for model training; the target detection model is obtained by concatenating a backbone network, a hybrid encoder, a decoder, and a detection head; in the target detection model: A rich branch aggregation pixel-aware abnormal behavior feature extraction network is used as the backbone network; The hybrid encoder includes two feature recovery modules, two feature fusion modules, multiple reparameterized convolutional modules, and multiple parallel pixel fusion attention detail enhancement modules. The two feature recovery modules in the hybrid encoder are respectively connected to the output layer of the backbone network. The two feature recovery modules are connected in series. The outputs of the two feature recovery modules are respectively fed into the two series-connected feature fusion modules. The final outputs of the two series-connected feature fusion modules and the individual outputs of each feature fusion module are fed into the reparameterized convolutional modules. The reparameterized convolutional modules are then connected in series with the detail enhancement modules. The hybrid encoder and decoder are spliced ​​together, and the output of the decoder is mapped to shallow feature maps of different resolutions through linear transformation and reconstruction operations, which are then fed into the cross-scale abnormal small target decoupling detection head with dynamic weight allocation. S130: Input the real-time acquired images into the trained fish abnormal behavior target detection model, and use the model output information as data samples for visual and text prompts. Input the fish abnormal behavior description and diagnostic language model for processing, and output the fish abnormal behavior description and diagnostic language.

2. The method for monitoring, describing, and diagnosing abnormal fish behavior according to claim 1, characterized in that: The abnormal behavior feature extraction network includes multiple cascaded processing units. Each processing unit contains a segmentation and padding layer and multiple cascaded gated aggregation attention modules. Layer normalization layers are set between adjacent processing units. The hybrid encoder receives the output of the layer normalization layers as its input.

3. The method for monitoring, describing, and diagnosing abnormal fish behavior according to claim 2, characterized in that: The segmentation padding layer is obtained by concatenating an edge verification layer, a convolutional layer, and a normalization layer, wherein the edge verification layer performs position encoding operations. The gated attention aggregation module includes an input layer, a layer normalization layer, a dual-path attention aggregation module, a residual connection, a layer normalization layer, a gated convolution module, a residual connection, and an output layer, which are connected in series.

4. The method for monitoring, describing, and diagnosing abnormal fish behavior according to claim 3, characterized in that: The dual-path attention aggregation module is designed with two types of path branches, including a convolutional branch and an embedding-scaling attention branch. The convolutional branch uses a depthwise separable convolution and dynamic encoding scheme that includes depthwise convolution and pointwise convolution. The embedding-scaling attention branch uses attention computation that includes a length scaling factor and a temperature scaling factor. The gated convolutional module includes an input layer, a dual-parallel branch structure, a gated fusion layer, a linear transformation layer, a residual connection, and an output layer, all connected in series. The dual-parallel branch structure consists of a left branch and a right branch. The left branch extracts local spatial features by passing through a linear transformation layer, a depthwise separable convolutional layer, and an activation function layer. The right branch contains only a linear transformation layer. The outputs of the two paths are fused through the gated fusion layer.

5. A method for monitoring, describing, and diagnosing abnormal fish behavior according to any one of claims 1-4, characterized in that: The detail enhancement module performs attention-weighted fusion of two input feature maps, including: fusing the input feature maps element-wise to obtain an initial feature map; using a channel-oriented attention module to model the channel dimension of the initial feature map and generate a channel attention map; using a spatial attention module to generate a spatial attention map by combining average pooling and max pooling; adding the channel attention map and the spatial attention map and broadcasting them to form a joint attention map; inputting the joint attention map to a pixel-oriented attention module to generate a pixel-wise fusion weight map; achieving dynamic weighted fusion of the input feature maps through the fusion weight map; and performing spatially enhanced convolution processing through a slice fusion convolution module.

6. The method for monitoring, describing, and diagnosing abnormal fish behavior according to claim 5, characterized in that: The channel-oriented attention module includes an input layer, a global average pooling layer, a dimensionality-reduced 1x1 convolutional layer, a ReLU activation function layer, an up-dimensional 1x1 convolutional layer, and an output layer, which are connected in series. The spatial attention-oriented module includes an input layer, a channel average pooling layer, a channel max pooling layer, a splicing layer, a 7x7 convolutional layer with reflection fill, and an output layer, which are connected in series. The pixel-oriented attention module includes an input layer, a splicing layer, a dimension transformation layer, a 7x7 depthwise convolutional layer, and an output layer, which are connected in series.

7. The method for monitoring, describing, and diagnosing abnormal fish behavior according to claim 6, characterized in that: The slice fusion convolutional module includes a convolutional layer, a batch normalization layer, a SiLU activation function, and a feature space slice fusion module connected in series; wherein, the specific processing steps of the feature space slice fusion module on the input feature map are as follows: The input feature map is divided into four sub-blocks along the spatial dimension. The average pixel value of each sub-block is calculated and the variance is normalized to obtain the standardized feature distribution. The normalization result is mapped to the range [0,1] using the Sigmoid function to obtain the attention weight of the sub-block. Each sub-block is multiplied by its corresponding attention weight to obtain the weighted sub-block. Finally, all the weighted sub-blocks are concatenated by tensors according to their original spatial positions and recombined into the output feature map.

8. The method for monitoring, describing, and diagnosing abnormal fish behavior according to claim 7, characterized in that: The dynamic weight allocation cross-scale abnormal small target decoupling detection head includes a spatial dynamic fusion and aggregation module and a multi-scale feature selection fusion detection head. The spatial dynamic fusion and aggregation module includes a spatial dynamic weighted fusion module. The spatial dynamic fusion and aggregation module first performs spatial dynamic weighted fusion on the input first shallow feature map and the upsampled result of the second shallow feature map, and the downsampled result of the first shallow feature map and the second shallow feature map, respectively. Then, it performs spatial dynamic weighted fusion again on the third shallow feature map and the spatial dynamic weighted fusion result of the two parts. Finally, it performs residual connection operation on the three spatial dynamic weighted fusion results respectively to output the first deep feature map, the second deep feature map, and the third deep feature map. The classification head of the multi-scale feature selection fusion detection head is used to predict the category, and the regression head is used to predict the offset of the target detection anchor box. The two receive the fusion result as input, and the final output is the concatenation result of the outputs of the classification head and the regression head.

9. A method for monitoring, describing, and diagnosing abnormal fish behavior according to claim 8, characterized in that: In step S130, the open-source Llama 3.1 large language model is used as the language model for describing and diagnosing abnormal fish behavior. Based on the output of the target detection model, the user-uploaded problem descriptions and professional fish diseases, the language model breaks down the abnormal fish diagnosis and treatment strategy into question-answer pairs, segments long texts according to semantics or paragraphs, integrates the answers, and stores the generated question-answer pairs and text fragments as knowledge units in the PostgreSQL database as the knowledge basis for subsequent answers. At the same time, it connects to the knowledge base of professional aquaculture personnel and the knowledge base of fish disease diagnosis and treatment experts to generate a knowledge base for describing and diagnosing abnormal fish behavior. The model is trained by reinforcement learning using feedback from experts in the field of fishery diseases to optimize the output quality of the model.

Citation Information

Patent Citations

  • Medical image small target segmentation method based on double-branch feature fusion attention

    CN116681679A

  • Transform-based human body posture estimation method and system

    CN118298016A