Fish body key point detection method and system based on smart fishery
Through the improved fish body key point detection model, combined with the parallel structure of Basic Block and Mamba Block, feature extraction of local details and global associations is achieved, which solves the problem of accurate positioning of fish body key point detection in smart fisheries, improves detection accuracy and achieves lightweight model.
Patent Information
- Application Number
- CN202511144095.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies in the field of smart fisheries lack the ability to accurately locate the key points of fish bodies. Especially when faced with complex environments such as large changes in fish postures and frequent occlusions, existing models find it difficult to achieve high-precision detection of key points on fish bodies.
An improved fish body key point detection model is adopted. All Basic Blocks in the HRNet model are replaced by BM Blocks. The BM Block is composed of Basic Block and Mamba Block in parallel. Cross-resolution feature interaction is achieved through a multi-resolution feature processing module. The feature extraction of local details and global correlation is combined, and the loss function is optimized to improve the detection accuracy.
The accuracy and robustness of fish body key point detection are significantly improved, increasing the key point detection accuracy by 20%. At the same time, the model is lightweight, reducing video memory usage, and is suitable for resource-constrained devices.
Smart Images

Figure CN120635950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a fish body key point detection method and system based on smart fishery. Background Art
[0002] Keypoint detection is a key research area in computer vision, widely used in human pose estimation, object tracking, biometric recognition, and the recently emerging fields of smart agriculture and smart fisheries. Traditional keypoint detection methods rely primarily on handcrafted features and machine learning algorithms, such as the Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), and Histogram of Oriented Gradients (HOG). These methods perform well for specific tasks and controlled environments, but their robustness and generalization capabilities are limited in complex environments such as scale variations, occlusion, and illumination changes.
[0003] With the development of deep learning, keypoint detection methods based on convolutional neural networks (CNNs) have rapidly replaced traditional methods based on image processing and handcrafted features, significantly improving detection accuracy and generalization. The DeepPose framework proposed by Toshev et al. was the first to model human pose estimation as a direct regression problem of keypoint coordinates. Using a deep neural network architecture, it effectively mines high-level semantic information in images, laying the foundation for subsequent deep learning-based keypoint detection methods.
[0004] On this basis, a number of classic architectures have been proposed, including the Stacked Hourglass Network, Cascaded Pyramid Network, AlphaPose, Simple Baselines and other series of optimization networks. These models effectively improve the accuracy of key point prediction and the convergence speed of the model by introducing mechanisms such as multi-scale fusion, pyramid structure, and residual module, promoting continuous progress in this field.
[0005] At the same time, lightweight architectures are gaining popularity to meet the demands of edge computing and real-time detection. The YOLO series of models, due to their end-to-end structural design, excellent real-time performance, and ease of deployment, are widely used in real-time pose estimation tasks. YOLOv8n-Pose, the latest lightweight keypoint detection model in the series, has demonstrated excellent performance in tasks such as human pose estimation, cattle body shape recognition, and strawberry stem detection and picking. Furthermore, models such as YOLOv7 and PointMap achieve a good balance between keypoint positioning accuracy and speed, providing a reliable technical foundation for deployment on resource-constrained terminal devices.
[0006] The introduction of the high-resolution network (HRNet) is of great significance in high-resolution image processing. By maintaining feature maps at multiple resolutions in parallel and enabling cross-scale information exchange, this network effectively preserves spatial details and structural information, making it particularly suitable for tasks requiring precise keypoint positions. HRNet has been successfully applied to a variety of fields, including human pose estimation, keypoint detection in cattle, and cervical spine localization in medical imaging, demonstrating its advantages in modeling fine structures. Despite its excellent performance, its large number of model parameters and high computational overhead limit its deployment in resource-constrained scenarios. Consequently, researchers have conducted extensive research on lightweighting and structural pruning optimization of its network structure to improve its practicality.
[0007] At the same time, the Transformer architecture has attracted widespread attention in computer vision for its global modeling capabilities and self-attention mechanism. Its introduction into keypoint detection enables the model to handle long-range dependencies and complex spatial structures, significantly improving detection accuracy and robustness. In recent years, architectures such as the MMSF-Transformer, which integrates multi-scale and multi-level semantics, and the Keypoint Former, which combines dynamic convolution with a sparse token mechanism, and its lightweight version, Keypoint Former-S, have emerged. These methods combine the local feature extraction advantages of CNNs with the global feature modeling capabilities of the Transformer, demonstrating excellent adaptability in complex pose and multi-object tasks.
[0008] The aforementioned research by domestic and international scholars has primarily focused on keypoint detection on humans or large animals. In the field of smart fisheries, there is relatively little research on accurately locating keypoints on fish bodies. Furthermore, existing models struggle to directly transfer keypoints to fish due to their large posture variations and frequent occlusions. Furthermore, existing fish keypoint detection methods primarily focus on three annotated points: the head, abdomen, and tail. This limits the number of detection points and makes the task relatively simple.
[0009] To this end, it is urgently necessary to combine lightweight design, high-precision network structure and time series modeling capabilities to build a dedicated key point detection model that adapts to the complex scenario needs of smart fisheries. Summary of the Invention
[0010] The main purpose of the present invention is to provide a fish body key point detection method and system based on smart fishery, aiming to overcome the defect of lack of accurate positioning of fish body key points in the field of smart fishery.
[0011] To achieve the above objectives, the present invention provides a fish body key point detection method based on smart fishery, comprising the following steps: Acquire fish body images; The fish body image is input into the fish body key point detection model, and multiple key points are detected in the fish body image; wherein, the fish body key point detection model replaces all Basic Blocks in the HRNet model with BMBlock, and the BM Block is composed of Basic Block and Mamba Block in parallel; after the fish body key point detection model extracts the basic features of the fish body image through the initial two convolution modules, it goes through multiple stages of multi-resolution feature processing modules, and the BM Blocks inside each multi-resolution feature processing module process feature maps of different scales in parallel and realize cross-resolution feature interaction through the fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image.
[0012] Furthermore, the Basic Block in the BM Block uses convolution and normalization operations to enhance local feature expression; the Mamba Block leverages its long sequence modeling advantage to explore global dependencies between features. The Basic Block and Mamba Block operate in parallel and then fuse their outputs, allowing the features to have both local details and global associations.
[0013] Furthermore, upsampling and downsampling operations are performed between different branches in the multi-resolution feature processing module to adjust the sizes of feature maps with different resolutions so as to adapt and fuse them.
[0014] Furthermore, each of the multi-resolution feature processing modules includes at least one BM Block.
[0015] Furthermore, the Basic Block sequentially includes a first convolutional layer, a first batch of normalization layers, a first activation function layer, a second convolutional layer, a second batch of normalization layers, and a second activation function layer; In the Basic Block, the input features are first extracted through the first convolution layer to extract local spatial features, normalized by the first batch of normalization layers, activated by the first activation function layer, and then adjusted through the second convolution layer to adjust the channels and refine the features, and normalized again by the second batch of normalization layers; finally, they are added to the original input features through a residual connection, activated by the second activation function layer, and the features processed by the branch are output.
[0016] Furthermore, the Mamba Block backbone path is sequentially composed of a root mean square normalization layer, a depthwise separable 1D convolutional layer, a first SiLU linear unit, a first linear layer, and an S6 module; In the Mamba Block, the input features are first normalized by the RMS normalization layer, and then enter the depthwise separable 1D convolutional layer to extract sequence or time-related features. They are activated by the first SiLU linear unit and linearly transformed by the first linear layer. The output of the first linear layer is multiplied by the output of the S6 module to enhance the features. At the same time, the Mamba Block trunk path branches out a branch path after the root mean square normalization layer, which is the connected second SiLU linear unit and the second linear layer. The trunk path and the branch path are finally integrated through the residual connection, and the features processed by the branch path are added to the features on the trunk path to output the integrated features.
[0017] Furthermore, the features after the Basic Block and Mamba Block operations are concatenated or fused at the end of the BM Block, and the fused features that fuse the local spatial features and the long-sequence global features are output and passed to the subsequent network layers.
[0018] Furthermore, the loss function expression of the fish body key point detection model is:
[0019] in, represents the regression loss of the key parts, It represents the structural proportional error term introduced, which is used to measure whether the predicted point satisfies the proportional stability between the segments of the tail fin; To adjust the hyperparameters, for balancing and The importance of two parts of loss.
[0020] Furthermore, the The calculation formula is:
[0021] in, represents the true length of the i-th key point pair, is the corresponding predicted length; 、 They respectively refer to the true length and predicted length of another key point pair that is uniquely corresponding to the key point pair; and are the structural importance weights of the key point pairs.
[0022] The present invention also provides a fish body key point detection system based on smart fishery, comprising: An acquisition module, used for acquiring fish body images; A detection module is provided, which is used to input the fish body image into a fish body key point detection model and detect multiple key points in the fish body image; wherein, the fish body key point detection model replaces all Basic Blocks in the HRNet model with BM Blocks, and the BM Blocks are composed of Basic Blocks and Mamba Blocks in parallel; after the fish body key point detection model extracts the basic features of the fish body image through the initial two convolution modules, the fish body key point detection model passes through multiple stages of multi-resolution feature processing modules, and the BM Blocks within each multi-resolution feature processing module process feature maps of different scales in parallel and realize cross-resolution feature interaction through a fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image.
[0023] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0024] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0025] The present invention provides a fish body key point detection method and system based on smart fishery, comprising: obtaining a fish body image; inputting the fish body image into a fish body key point detection model, and detecting multiple key points in the fish body image; wherein the fish body key point detection model replaces all Basic Blocks in the HRNet model with BM Blocks, and the BM Blocks are composed of Basic Blocks and Mamba Blocks in parallel; after the fish body key point detection model extracts the basic features of the fish body image through two initial convolution modules, the fish body key point detection model passes through multiple stages of multi-resolution feature processing modules, and the BM Blocks within each multi-resolution feature processing module parallelly process feature maps of different scales and realize cross-resolution feature interaction through a fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image. In the present invention, the Basic Block in the HRNet model is completely replaced with the BM Block through the fish body key point detection model. The BM Block is composed of the Basic Block and the Mamba Block in parallel, which combines local details with global associations, improves the accuracy of key point detection in fish body images, and overcomes the defect of lack of accurate positioning of fish body key points in the field of smart fishery. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 1 is a schematic diagram of the steps of a fish body key point detection method based on smart fishery in one embodiment of the present invention; Figure 2 This is a structural block diagram of a fish body key point detection system based on smart fishery in one embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.
[0027] The implementation, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0029] Reference Figure 1 In one embodiment of the present invention, a method for detecting key points of a fish body based on smart fishery is provided, comprising the following steps: Step S1, obtaining a fish body image; Step S2: input the fish body image into the fish body key point detection model, and detect multiple key points in the fish body image; wherein, the fish body key point detection model replaces all Basic Blocks in the HRNet model with BM Blocks, and the BM Blocks are composed of Basic Blocks and Mamba Blocks in parallel; after the fish body key point detection model extracts the basic features of the fish body image through the initial two convolution modules, it passes through multiple stages of multi-resolution feature processing modules, and the BM Blocks inside each multi-resolution feature processing module process different scale feature maps in parallel and realize cross-resolution feature interaction through the fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image.
[0030] In this embodiment, as described in step S1 above, images containing fish bodies can be acquired using underwater surveillance cameras, drone photography equipment, or other image acquisition devices. These images can be sourced from natural water environments, such as lakes and rivers, or artificial breeding environments, such as fish ponds and aquariums. The acquired fish images can also have specific dimensions, such as a height of 1216 pixels and a width of 800 pixels, and use an RGB three-channel color mode. This mode can provide rich visual information for subsequent key point detection of the fish body.
[0031] In practical applications, to improve image quality and make it more suitable for subsequent analysis and processing, a series of preprocessing operations can be performed on the acquired raw images. These preprocessing operations include, but are not limited to, image denoising, brightness adjustment, and contrast enhancement to remove interference factors in the image and highlight the characteristics of the fish.
[0032] As described in step S2 above, the fish key point detection model used in this step innovatively improves the HRNet (High-Resolution Network) model. It replaces all Basic Blocks in the HRNet model with BM Blocks (Basic Mamba Blocks). The BM Blocks utilize a unique parallel structure, consisting of a traditional Basic Block and an innovative Mamba Block in parallel. While the traditional Basic Block can efficiently extract local detail features of the fish, the Mamba Block excels at capturing long-range dependencies between fish features. This design makes the model more adaptable when handling fish pose estimation tasks.
[0033] The Mamba Block, a key innovation of the model, builds on the Mamba sequential model architecture and constructs a backbone path consisting of "RMSNorm→DWConv1d→SiLU→Linear→S6." This path effectively enables dynamic modeling of long sequences. Furthermore, the DWConv1d (depthwise separable convolution) in this path preserves the temporal characteristics of fish body features while reducing the model's computational complexity. This makes it particularly suitable for handling large flexible deformations in areas such as the tail.
[0034] Specifically, when a fish image is input into the fish keypoint detection model, it first passes through two initial convolutional modules. These modules perform preliminary feature extraction on the fish image, thereby obtaining basic features. These basic features contain some basic contours and local features of the fish, laying the foundation for subsequent in-depth feature processing.
[0035] Based on the extracted basic features, the model further processes them through multiple stages of multi-resolution feature processing modules. Within each multi-resolution feature processing module, the BM Block processes feature maps of different scales in parallel. These feature maps at different scales can respectively capture local details and global structural features of the fish. Furthermore, the model implements cross-resolution feature interaction through a fusion module, allowing features at different scales to complement and fuse with each other, thereby improving the model's ability to detect key points on the fish.
[0036] After multiple stages of multi-resolution feature processing, the final multi-resolution feature processing module concatenates all feature maps. This concatenation allows the model to integrate feature information at different scales and levels, forming a more comprehensive and rich feature representation. Based on this fused feature representation, the model then outputs coordinate information containing multiple key points. This coordinate information is initially in the model's feature space. To accurately locate the key points of the fish in the original image, the model maps this coordinate information back to the original fish image, ultimately achieving precise detection of the fish's key points.
[0037] Through the above steps, this method can accurately locate multiple key points in fish images. These key points include important parts such as the eyes, fins (including dorsal, pectoral, pelvic, anal, and caudal fins), head, and tail. The precise detection of these key points provides strong support for subsequent refined individual behavior recognition tasks in intelligent aquaculture systems, such as fish posture analysis, behavior recognition, and growth assessment. This helps achieve efficient management and scientific decision-making in smart fisheries.
[0038] In one embodiment, the Basic Block in the BM Block uses convolution and normalization operations to enhance local feature expression; the Mamba Block leverages its long sequence modeling advantage to explore global dependencies between features. The Basic Block and the Mamba Block operate in parallel and then fuse their outputs, so that the features have both local details and global associations.
[0039] In this embodiment, in the task of detecting key points of fish bodies in smart fisheries, the innovative design of BM Block perfectly realizes the collaborative modeling of local features and global dependencies through a dual-branch parallel architecture.
[0040] Basic Block (traditional convolution branch): uses the standard 3×3 convolution → BN → ReLU → 1×1 convolution → BN structure → ReLU, where the 3×3 convolution kernel is responsible for extracting local texture features of the fish body (such as scale structure and fin details), and the 1×1 convolution is used for channel dimensionality reduction and feature reorganization.
[0041] The residual connection (x + F(x)) effectively retains the original input information, alleviates the gradient vanishing problem of deep networks, and is particularly suitable for capturing the stable features of the rigid structure of the fish body (such as the fish head and fish eyes).
[0042] Mamba Block (sequence modeling branch): It significantly reduces the number of parameters through group convolution while preserving the temporal correlation of sequence features. With its long sequence modeling advantage, it mines the global dependency between features. The outputs of the two branches are fused through channel concatenation or element-wise summation to form a composite feature representation that combines spatial detail with temporal correlation. For example, when detecting a fish turning, the Basic Block provides local morphological features of the tail fin, while the Mamba Block captures the long-range temporal relationship between the tail fin's oscillation and the body's movement. The combination of the two allows for precise calculation of the turn angle and amplitude.
[0043] In one embodiment, upsampling and downsampling operations are performed between different branches in the multi-resolution feature processing module to adjust the sizes of feature maps with different resolutions so as to adapt and fuse them.
[0044] In one embodiment, each of the multi-resolution feature processing modules includes at least one BM Block.
[0045] In one embodiment, the Basic Block sequentially includes a first convolution layer, a first batch of normalization layers, a first activation function layer, a second convolution layer, a second batch of normalization layers, and a second activation function layer; In the Basic Block, the input features are first extracted through the first convolution layer to extract local spatial features, normalized by the first batch of normalization layers, activated by the first activation function layer, and then adjusted through the second convolution layer to adjust the channels and refine the features, and normalized again by the second batch of normalization layers; finally, they are added to the original input features through a residual connection, activated by the second activation function layer, and the features processed by the branch are output.
[0046] In this embodiment, the Basic Block, as a traditional convolutional branch within the BM Block, is responsible for extracting fine local features of the fish. Its structural design and operational logic strictly adhere to the best practices of modern convolutional neural networks. Through multi-stage processing and a residual connection mechanism, it efficiently extracts and enhances local features of the fish.
[0047] The first convolutional layer (Conv_3×3) uses a 3×3 convolution kernel with a stride of 1 and padding of 1 to maintain the feature map size. This layer is primarily responsible for extracting basic local features of the fish, such as scale texture, fin edges, eye contours, and other spatial details. The number of convolution kernels (output channels) is typically 64 or 128, depending on the network design stage.
[0048] Batch Normalization Layer (BN1): Batch normalization is performed on the output of the first convolutional layer. This normalization operation accelerates network convergence, enhances model stability, and alleviates the problem of internal covariate shift.
[0049] The first activation function layer (ReLU) and the second activation function layer: Apply the Rectified Linear Unit (Rectified Linear Unit) activation function to introduce nonlinear characteristics, enabling the network to learn complex function mapping relationships and enhance feature expression capabilities.
[0050] The second convolutional layer (Conv_1×1) uses a 1×1 convolution kernel to resize the feature map's channel dimensions and reorganize features. This layer's primary function is to reduce computational complexity (by reducing the number of channels) and integrate cross-channel information. The number of output channels is typically the same as the number of input channels to maintain dimensionality matching for residual connections.
[0051] Second batch normalization layer (BN2): Normalizes the features again to further stabilize the network training process.
[0052] Residual connection mechanism: After the input feature x passes through the five layers mentioned above, the output feature F(x) is obtained. Through the residual connection, the original input x is added to the processed feature F(x). This design effectively alleviates the vanishing gradient problem in deep networks and ensures that the network can learn more subtle feature differences.
[0053] The role of this structure in detecting fish keypoints: The spatial locality of the 3×3 convolution kernel makes it particularly suitable for capturing fine structures in the fish body, such as the bifurcation of the tail fin and the swing angle of the pectoral fin. Feature hierarchy construction: By cascading two layers of convolution, a hierarchical representation is formed, from basic features to combined features. For example, from edge detection to texture combination, higher-level semantic features are gradually constructed.
[0054] The synergistic mechanism complements the features of the Mamba Block: The local features extracted by the Basic Block provide Mamba Block with precise spatial positioning information, while the global dependencies captured by the Mamba Block help Basic Block better distinguish similar local structures (such as the key points of different fins). Dynamic fusion: At the output of the BM Block, the two features are fused through channel concatenation or attention weighting to form a composite feature representation that combines spatial details and temporal correlations.
[0055] In one embodiment, the Mamba Block backbone path is sequentially a root mean square normalization layer, a depthwise separable 1D convolutional layer, a first SiLU linear unit, a first linear layer, and an S6 module; In the Mamba Block, the input features are first normalized by the RMS normalization layer, and then enter the depthwise separable 1D convolutional layer to extract sequence or time-related features. They are activated by the first SiLU linear unit and linearly transformed by the first linear layer. The output of the first linear layer is multiplied by the output of the S6 module to enhance the features. At the same time, the Mamba Block trunk path branches out a branch path after the root mean square normalization layer, which is the connected second SiLU linear unit and the second linear layer. The trunk path and the branch path are finally integrated through the residual connection, and the features processed by the branch path are added to the features on the trunk path to output the integrated features.
[0056] In this embodiment, in the fish key point detection model for smart fisheries, the Mamba Block is the core module for implementing long-range feature association. Through the coordinated design of trunk and branch paths, it efficiently captures the temporal dependencies in fish movement. The specific workflow is as follows: The input fish features first pass through a root mean square normalization layer. This layer acts like a feature calibration, ensuring that feature values in different ranges do not interfere with each other, making model training more stable. This normalization is particularly helpful when processing features with varying movement amplitudes, preventing large features from dominating the model learning.
[0057] Next comes the depthwise separable 1D convolutional layer, which is key to capturing temporal relationships. This layer slides the convolution kernel along the sequence dimension (e.g., the temporal order of video frames) to specifically extract the motion patterns of key points on the fish at different time points. For example, it can identify the order in which the tail fin sways across consecutive frames. The depthwise separable design reduces computational effort while preserving temporal features.
[0058] The extracted temporal features are activated by the SiLU function, which can retain the positive information in the features while smoothing the negative values, allowing the model to better express the complex patterns of fish movement, such as the different states of the tail fin bending to the left or right.
[0059] The activated features then pass through the first linear layer, which further integrates information from different channels, combines temporal features with inter-channel dependencies, and forms a more abstract feature representation, preparing for subsequent processing of complex motion patterns.
[0060] Finally, the S6 module (enhanced activation function) uses multiple layers of nonlinear transformations and residual connections to enhance feature extraction for fish in violent motion or complex postures. For example, when a fish suddenly turns, this module can capture the coordinated motion of multiple parts, while residual connections ensure that the original feature information is not lost.
[0061] The branch path begins with the RMS-normalized features and independently passes through a second SiLU linear unit. This step is equivalent to opening an additional processing channel for the features, extracting details that the main path may overlook through different nonlinear transformations, such as the subtle vibrations of the fish's fins. The branch path features are then adjusted to a dimension compatible with the main path through a second linear layer, and then added and fused with the output of the main path. This design, through the collaboration of the two, with the main path processing the main temporal relationships and the branch path supplementing the detailed features, provides the model with a more comprehensive perception of the fish's movement.
[0062] In the backbone path, the output of the S6 module is added to the initial input features through a residual connection. Even if some information is lost during deep processing, the original features can be recovered through the residual connection, preventing performance degradation due to increased model depth.
[0063] The output of the Mamba Block is fused with the local features of the Basic Block. For rigid parts like the fish's head, the Basic Block's local positioning capabilities are relied upon, while for flexible parts like the tail fin, the Mamba Block's temporal modeling results are prioritized. This dynamic fusion enables the model to both accurately locate key points and understand their kinematic relationships.
[0064] When detecting tail fins, Mamba Block captures the fin's swing sequence over several consecutive frames. Combined with processing by a specialized module, it accurately calculates the swing amplitude and frequency, which is crucial for determining fish health (sick fish often exhibit abnormal tail swings). In multi-fish scenarios, Mamba Block can analyze the coordinated motion of key points across the fish, such as identifying the synchronized movement of heads and tails as a group of fish turn, enabling intelligent identification of group behavior.
[0065] The design of depth-wise separable convolution and root mean square normalization allows Mamba Block to reduce computational complexity compared to traditional modules while maintaining powerful modeling capabilities.
[0066] Mamba Block cleverly combines local features and global dependencies through the design of trunk timing modeling + branch detail supplementation + residual information retention. It can not only clearly see the position of each key point of the fish body, but also understand their movement correlation in the time dimension, providing core technical support for accurate detection of smart fisheries.
[0067] In one embodiment, the features after the Basic Block and Mamba Block operations are concatenated or fused at the end of the BM Block, and the fused features that fuse the local spatial features and the long-sequence global features are output and passed to the subsequent network layer.
[0068] In this example, the local features extracted by the Basic Block (such as the caudal fin bifurcation) provide a precise spatial positioning benchmark for the Mamba Block, helping it more accurately capture temporal changes. The long-range temporal relationships captured by the Mamba Block (such as the correlation between caudal fin swinging and body turning) in turn guide the Basic Block in more accurately distinguishing similar local structures (such as key points on different fins).
[0069] In one embodiment, the loss function expression of the fish body key point detection model is:
[0070] in, represents the regression loss of the key parts, It represents the structural proportional error term introduced, which is used to measure whether the predicted point satisfies the proportional stability between the segments of the tail fin; To adjust the hyperparameters, for balancing and The importance of two parts of loss. It is initially set to 1e3 to emphasize a greater penalty when its prediction is wrong.
[0071] In this embodiment, to effectively incorporate the proportional stability assumption from biomorphology, the present invention introduces a scale-aware weighted loss (SAWL) into the loss function. This introduction of a proportional error term allows the model to overcome the length measurement bias caused by individual differences in traditional methods. Based on the density and importance of key points, the model progressively simulates the compliance domain (the length range that adapts to individual morphological variations), thereby improving the accuracy of the prediction of the length region limit and achieving decentralized and precise positioning of the prediction points.
[0072] To address the problem of prediction point stacking in dense keypoint detection, an innovative scaling mechanism is introduced to optimize the loss function. By modeling the spatial distribution and biological characteristics of keypoints in the fish body structure, dynamically assigning weights to different keypoints and imposing higher penalties on prediction errors in dense areas, the model's ability to perceive fine-grained regions is enhanced, effectively alleviating the problem of prediction point clustering and improving the accuracy of dense keypoint detection.
[0073] In one embodiment, the The calculation formula is:
[0074] in, represents the true length of the i-th key point pair, is the corresponding predicted length; 、 They respectively refer to the true length and predicted length of another key point pair that is uniquely corresponding to the key point pair; and are the structural importance weights of the key point pairs.
[0075] In the above embodiment, the following technical advances are achieved: Optimizing the loss function to mitigate keypoint deviation: Targeted improvements to the loss function effectively mitigate the clustering and positional offset issues that often occur with predicted keypoints. This improvement directly improves the spatial distribution rationality of keypoint positioning and the accuracy of individual positions.
[0076] Significantly improved keypoint detection accuracy: Thanks to the aforementioned model architecture enhancements and loss function optimization, a significant improvement in the core performance metric (keypoint detection accuracy) was achieved. Testing on a benchmark dataset demonstrated that, compared to the original solution, the proposed method achieved an average improvement of nearly 20% in keypoint detection accuracy (reducing positioning error by an average of 20% and increasing the proportion of keypoints meeting the preset accuracy threshold by 20%).
[0077] Significant lightweighting: Through model compression or specific configuration optimization, this invention can further achieve a lightweight version that reduces video memory usage by up to 12%, while maintaining keypoint detection accuracy that is superior to, or at least comparable to, the original model. This significantly improves the feasibility of deploying this solution on resource-constrained devices, such as mobile devices and embedded systems.
[0078] Reference Figure 2 In another embodiment of the present invention, a fish body key point detection system based on smart fishery is provided, comprising: An acquisition module, used for acquiring fish body images; A detection module is provided, which is used to input the fish body image into a fish body key point detection model and detect multiple key points in the fish body image; wherein, the fish body key point detection model replaces all Basic Blocks in the HRNet model with BM Blocks, and the BM Blocks are composed of Basic Blocks and Mamba Blocks in parallel; after the fish body key point detection model extracts the basic features of the fish body image through the initial two convolution modules, the fish body key point detection model passes through multiple stages of multi-resolution feature processing modules, and the BM Blocks within each multi-resolution feature processing module process feature maps of different scales in parallel and realize cross-resolution feature interaction through a fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image.
[0079] In this embodiment, for the specific implementation of each module in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.
[0080] Reference Figure 3In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.
[0081] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.
[0082] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0083] In summary, the fish body key point detection method and system based on smart fishery provided in the embodiment of the present invention include: obtaining a fish body image; inputting the fish body image into a fish body key point detection model, and detecting multiple key points in the fish body image; wherein, the fish body key point detection model replaces all Basic Blocks in the HRNet model with BM Blocks, and the BM Blocks are composed of Basic Blocks and Mamba Blocks in parallel; after the fish body key point detection model extracts the basic features of the fish body image through the initial two convolution modules, the BM Blocks in each multi-resolution feature processing module process different scale feature maps in parallel through multiple stages and realize cross-resolution feature interaction through the fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image. In the present invention, the Basic Block in the HRNet model is completely replaced with the BM Block through the fish body key point detection model. The BM Block is composed of the Basic Block and the Mamba Block in parallel, which combines local details with global associations, improves the accuracy of key point detection in fish body images, and overcomes the defect of lack of accurate positioning of fish body key points in the field of smart fishery.
[0084] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.
[0085] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.
[0086] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A fish body key point detection method based on smart fishery, characterized in that: The following steps are involved: Acquire fish body images; The fish body image is input into the fish body key point detection model, and multiple key points are detected in the fish body image; wherein, the fish body key point detection model replaces all Basic Blocks in the HRNet model with BM Blocks, and the BM Blocks are composed of Basic Blocks and Mamba Blocks in parallel; after the fish body key point detection model extracts the basic features of the fish body image through the initial two convolution modules, it goes through multiple stages of multi-resolution feature processing modules, and the BM Blocks inside each multi-resolution feature processing module process feature maps of different scales in parallel and realize cross-resolution feature interaction through the fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image.
2. The fish body key point detection method based on smart fishery according to claim 1 is characterized in that: The Basic Block in the BMBlock uses convolution and normalization operations to enhance local feature expression. The Mamba Block leverages its long sequence modeling advantage to explore global dependencies between features. The Basic Block and Mamba Block operate in parallel and then fuse their outputs, allowing the features to have both local details and global associations.
3. The fish body key point detection method based on smart fishery according to claim 1 is characterized in that: The different branches in the multi-resolution feature processing module adjust the sizes of feature maps of different resolutions through upsampling and downsampling operations to make them adaptive and fused.
4. The fish body key point detection method based on smart fishery according to claim 1 is characterized in that: Each of the multi-resolution feature processing modules includes at least one BM Block.
5. The fish body key point detection method based on smart fishery according to claim 1 is characterized in that: The Basic Block sequentially includes a first convolutional layer, a first batch of normalization layers, a first activation function layer, a second convolutional layer, a second batch of normalization layers, and a second activation function layer; In the Basic Block, the input features are first extracted through the first convolution layer to extract local spatial features, normalized by the first batch of normalization layers, activated by the first activation function layer, and then adjusted through the second convolution layer to adjust the channels and refine the features, and normalized again by the second batch of normalization layers; finally, they are added to the original input features through a residual connection, activated by the second activation function layer, and the features processed by the branch are output.
6. The fish body key point detection method based on smart fishery according to claim 1, characterized in that: The Mamba Block backbone path is the root mean square normalization layer, the depth-separable 1D convolution layer, the first SiLU linear unit, the first linear layer, and the S6 module. In the Mamba Block, the input features are first normalized by the RMS normalization layer, and then enter the depthwise separable 1D convolutional layer to extract sequence or time-related features. They are activated by the first SiLU linear unit and linearly transformed by the first linear layer. The output of the first linear layer is multiplied by the output of the S6 module to enhance the features. At the same time, the Mamba Block trunk path branches out a branch path after the root mean square normalization layer, which is the connected second SiLU linear unit and the second linear layer. The trunk path and the branch path are finally integrated through the residual connection, and the features processed by the branch path are added to the features on the trunk path to output the integrated features.
7. The fish body key point detection method based on smart fishery according to claim 1 is characterized in that: The features after the Basic Block and Mamba Block operations are concatenated or fused at the end of the BM Block, and the fused features that combine the local spatial features and the long-sequence global features are output and passed to the subsequent network layers.
8. The fish body key point detection method based on smart fishery according to claim 1 is characterized in that: The loss function expression of the fish body key point detection model is: ; in, represents the regression loss of the key parts, It represents the structural proportional error term introduced, which is used to measure whether the predicted point satisfies the proportional stability between the segments of the tail fin; To adjust the hyperparameters, for balancing and The importance of two parts of loss.
9. The fish body key point detection method based on smart fishery according to claim 8, characterized in that: described The calculation formula is: ; in, represents the true length of the i-th key point pair, is the corresponding predicted length; 、 They refer to the true length and predicted length of another key point pair that is unique to each key point pair; is the structural importance weight of the key point pair.
10. A fish body key point detection system based on smart fishery, characterized in that: include: An acquisition module, used for acquiring fish body images; A detection module is provided, which is used to input the fish body image into a fish body key point detection model and detect multiple key points in the fish body image; wherein, the fish body key point detection model replaces all Basic Blocks in the HRNet model with BM Blocks, and the BM Blocks are composed of Basic Blocks and Mamba Blocks in parallel; after the fish body key point detection model extracts the basic features of the fish body image through the initial two convolution modules, the fish body key point detection model passes through multiple stages of multi-resolution feature processing modules, and the BM Blocks inside each multi-resolution feature processing module process feature maps of different scales in parallel and realize cross-resolution feature interaction through a fusion module. Finally, the multi-resolution feature processing module splices all feature maps, outputs coordinates containing multiple key points, and maps them to the fish body image.
Citation Information
Patent Citations
System for monitoring energy consumption of tail fin of robotic fish
CN106645933A
Complex underwater environment fish size estimation system and method based on key point detection
CN120452029A