Pseudosciaena crocea swimming speed monitoring method and system based on target detection

By using a dual-sensor visual acquisition system and an improved YOLOv7 network model, combined with biomechanics and flow velocity vector verification, the accuracy problem of monitoring the swimming speed of large yellow croaker in complex underwater environments was solved, and high-precision swimming speed calculation was achieved.

CN121053705AActive Publication Date: 2025-12-02GUANGDONG OCEAN UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511587595.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2025-12-02
Estimated Expiration
2045-11-03

AI Technical Summary

Technical Problem

Existing technologies cannot accurately monitor the swimming speed of large yellow croaker in complex underwater environments. The monitoring accuracy is low due to factors such as optical distortion and water turbidity.

Method used

A dual-sensor visual acquisition system using a narrow-beam laser rangefinder and an underwater wide-angle camera, combined with an improved YOLOv7 network model, is used for target detection. Through biomechanical constraints and flow vector verification, an attitude-fluid coupling model is constructed to optimize attitude characteristics and determine swimming speed.

Benefits of technology

It enables rapid and accurate acquisition of large yellow croaker coordinate data in complex underwater environments, reduces attitude noise, improves the scientific rigor and accuracy of swimming speed calculations, and supports research on ecological habits and behavioral patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053705A_ABST
    Figure CN121053705A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of aquatic product monitoring, and discloses a large yellow croaker swimming speed monitoring method and system based on target detection, and the method comprises the steps: obtaining a to-be-monitored image of a target water area, and real-time environment flow velocity data and real-time water turbidity data of the target water area; performing target detection on the to-be-monitored image to obtain coordinate data of the large yellow croaker, performing linkage verification on the coordinate data to obtain verification coordinate data, determining a real-time attitude feature of the large yellow croaker according to the verification coordinate data, and optimizing the real-time attitude feature through a flow velocity phase joint mechanism to obtain an optimized attitude feature; constructing an attitude fluid coupling model according to the optimized attitude features and the real-time environment flow velocity data to determine the swimming speed of the large yellow croaker; according to the method, the influence of the posture of the pseudosciaena crocea and the environment is fully considered, and the real swimming speed of the pseudosciaena crocea is accurately calculated in combination with biomechanical characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aquatic monitoring technology, and in particular to a method and system for monitoring the swimming speed of large yellow croaker based on target detection. Background Technology

[0002] In the research and aquaculture of large yellow croaker, understanding its swimming speed is crucial for grasping its physiological state, behavioral habits, and optimizing the aquaculture environment. Currently, the main methods for monitoring fish swimming speed include tag-and-recapture, sonar detection, and video image analysis. Among these, tag-and-recapture is cumbersome, can be harmful to fish, and has low monitoring accuracy. Sonar detection is greatly affected by underwater environmental interference and is not effective for monitoring small fish. While video image analysis is relatively simple to operate, it is difficult to accurately obtain the fish's position and posture information in complex underwater environments due to factors such as optical distortion and water turbidity, resulting in low accuracy in swimming speed monitoring.

[0003] Therefore, how to overcome the limitations of existing technologies in dealing with the complex underwater environment and thus accurately monitor the swimming speed of large yellow croaker has become a pressing technical problem for those skilled in the art. Summary of the Invention

[0004] This invention provides a method and system for monitoring the swimming speed of large yellow croaker based on target detection, solving the problem of how to overcome the influence of complex underwater environments and thus accurately monitor the swimming speed of large yellow croaker.

[0005] To address the aforementioned technical problems, this invention provides a method for monitoring the swimming speed of large yellow croaker based on target detection, comprising: The system acquires images of the target water area, real-time environmental flow velocity data, and real-time water turbidity data, respectively. Target detection is performed on the image to be monitored to obtain the coordinate data of the large yellow croaker; The coordinate data is subjected to linkage verification to obtain verified coordinate data. The linkage verification is selected by the camera based on the analysis results of the coordinate data, and includes at least: a first-level verification based on biomechanical constraints to verify the motion state reflected by the coordinate data; and / or a second-level verification based on the flow velocity vector confirmed by the real-time water turbidity data; and / or a third-level verification based on the motion law reflected by the coordinate data and the biomechanical constraints. The real-time attitude features of the large yellow croaker are determined based on the verified coordinate data, and the real-time attitude features are optimized through a flow velocity-phase joint mechanism to obtain optimized attitude features. An attitude-fluid coupling model is constructed based on the optimized attitude characteristics and the real-time environmental flow velocity data, and the swimming speed of the large yellow croaker is determined through the attitude-fluid coupling model.

[0006] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows: Target detection in monitored images can quickly and accurately obtain the coordinate data of large yellow croaker. Simultaneously, by performing linked verification on the coordinate data—specifically, a three-level biomechanical verification linking body length and movement limits—the accuracy of the coordinate data is further improved, providing a reliable foundation for subsequent analysis and reducing analytical biases caused by coordinate errors. Optimizing real-time attitude features through a flow-velocity-phase joint mechanism can more accurately reflect the true attitude of the large yellow croaker, removing potential noise or unreasonable fluctuations in the attitude data, making the attitude features more stable and reliable, and contributing to a more accurate understanding of the large yellow croaker's behavioral state. Combining the optimized attitude features with environmental flow velocity data from the target water area to construct an attitude-fluid coupling model to determine the swimming speed of the large yellow croaker fully considers the influence of the large yellow croaker's own attitude and the surrounding water flow environment on its swimming speed. This allows for a more scientific and accurate calculation of the actual swimming speed of the large yellow croaker in complex water flow environments, providing strong support for studying the ecological habits, behavioral patterns, and interactions with the environment of the large yellow croaker. Attached Figure Description

[0007] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a flowchart of a method for monitoring the swimming speed of large yellow croaker based on target detection, provided in a certain embodiment of the present invention; Figure 2 This is a structural diagram of a target detection-based swimming speed monitoring system for large yellow croaker provided in a certain embodiment of the present invention; Figure label: Among them, 11 is the data acquisition module; 12 is the target detection module; 13 is the linkage verification module; 14 is the attitude optimization module; and 15 is the velocity quantization module. Detailed Implementation

[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings and examples. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0010] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items. Those skilled in the art will be able to understand the specific meaning of the above terms in this application according to the specific circumstances.

[0011] In the description of this application, it should be noted that, unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is merely for describing specific embodiments and is not intended to limit the invention. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.

[0012] In one embodiment, such as Figure 1 As shown, the first aspect of the present invention provides a method for monitoring the swimming speed of large yellow croaker based on target detection, comprising: S1. Acquire the image to be monitored, real-time environmental flow velocity data, and real-time water turbidity data of the target water area respectively; Specifically, this invention uses a dual-sensor visual acquisition system, including a narrow-beam laser rangefinder and an underwater wide-angle camera, to acquire real-time images of the target water area. Eight 4K underwater wide-angle cameras are deployed in the target water area, with four arranged in a tetrahedral pattern to cover the vertical area to be monitored (at 2m depth intervals), and four arranged along a 10m radius to cover the horizontal area, forming a tetrahedral + circumferential heterogeneous camera array. This overcomes the limitations of traditional single-dimensional monitoring and achieves three-dimensional full-space coverage. The cameras and the narrow-beam laser rangefinder are integrated into a hardware system, and a synchronous triggering device ensures the synchronization of their data acquisition. This allows the cameras to acquire several sets of images of large yellow croakers at different angles in real time, and the laser beam of the narrow-beam laser rangefinder to cover the areas where large yellow croakers frequently swim in the target water area, thus enabling real-time acquisition of the distance to the large yellow croakers.

[0013] In addition, water flow monitoring employs three-dimensional grid sampling: within a 10m×10m×5m monitoring space, 25 Doppler current profilers (ADCPs) and environmental parameter sensors (temperature, salinity, dissolved oxygen, turbidity meter, etc.) are arranged at 2m×2m×1m intervals, with a sampling frequency of 10Hz. Real-time environmental flow velocity data and real-time water turbidity data of the target water area are output simultaneously, providing mathematical support for subsequent large yellow croaker aquaculture research. It should be noted that all equipment is equipped with starlight-level night vision sensors (minimum illumination 0.01 lux) and anti-biofouling coatings to ensure clear imaging even in turbid water. The target water area in this invention can be an aquaculture pond or a sea area; no specific limitation is made here. The above deployment method is commonly used in aquaculture ponds. For sea areas or other target scenarios, the equipment deployment interval and number can be adjusted according to the actual situation. This invention uses a dual-sensor visual acquisition system that integrates a wide-angle underwater visible light camera and a narrow-beam laser rangefinder to acquire images of the target water area, which can overcome underwater optical distortion and achieve real-time calibration of the three-dimensional coordinates of the fish.

[0014] S2. Perform target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker; S3. Perform linkage verification on the coordinate data to obtain verified coordinate data; the linkage verification is selected by the camera based on the analysis results of the coordinate data, and includes at least: a first-level verification based on biomechanical constraints to verify the motion state reflected by the coordinate data; and / or a second-level verification based on the flow velocity vector confirmed by the real-time water turbidity data; and / or a third-level verification based on the motion law reflected by the coordinate data and the biomechanical constraints. Specifically, in steps S2 and S3 above, due to the large changes in the posture and scale of the large yellow croaker in complex underwater scenes, as well as problems such as water turbidity and background interference, the present invention achieves target detection of the image to be monitored by an improved YOLOv7 network model. It is based on the YOLOv7 model as a framework, and the improved YOLOv7 network model formed by introducing multi-scale deformable convolution and channel space dual attention mechanism into the framework performs target detection on the image to be monitored.

[0015] In one embodiment, the improved YOLOv7 network model includes an input layer (i.e., the input layer of the YOLOv7 model), a backbone feature extraction network layer, a neck feature fusion network layer, and a head detection output layer; the backbone feature extraction network layer includes a first ELAN block, a second ELAN block, a third ELAN block, and a fourth ELAN block; the second ELAN block, the third ELAN block, and the fourth ELAN block all employ multi-scale deformable convolution; wherein, the step of performing target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker includes: The input layer converts the image to be monitored into a normalized tensor. Based on the first ELAN block, the normalized tensor is subjected to initial feature extraction to obtain the first feature used to characterize the edge of the fish body. The second ELAN block is used to perform secondary feature extraction on the first feature to obtain a second feature that characterizes the fish body outline. Based on the third ELAN block, the second feature is extracted three times to obtain the third feature used to characterize the fish body morphology; The third feature is extracted four times using the fourth ELAN block to obtain the fourth feature used to characterize the distant fish body.

[0016] Specifically, this invention uses the image to be monitored acquired by the dual-sensor visual acquisition system as input. The input image is an RGB three-channel color image with a fixed size (e.g., 640×640 pixels). By improving the input layer of the model, simple preprocessing operations such as normalization and data augmentation are performed on the image to be monitored to accelerate the convergence of the model and improve the stability of training. Then, a 640×640×3 normalized tensor is generated and sent to the backbone feature extraction network layer.

[0017] The backbone feature extraction network layer of this invention retains the original ELAN (Efficient Layer Aggregation Network) structure block of YOLOv7, used to extract multi-scale features from the input image (shallow features capture details such as edges and textures, while deep features capture the overall outline of the fish, its shape, and distant fish features). It replaces the three key convolutional layers in the backbone network (corresponding to small, medium, and large-scale feature extraction, respectively) with multi-scale deformable convolutions. The backbone feature extraction network layer consists of five stages. After normalizing the tensor input, in the first stage, it first performs downsampling and channel adjustment through basic convolutional units (Conv+BN+SiLU) to output a 320×320×32 image.

[0018] In the second stage, a 320×320×32 image is input into the first ELAN block. The first ELAN block consists of multiple convolutional layers and residual connections. These convolutional layers mainly employ standard convolution operations, performing weighted summation operations on local regions of the input image to extract preliminary shallow features such as fish body edges, resulting in a first feature of 320×320×64. This feature contains basic image information, laying the foundation for subsequent feature extraction. The residual connections help solve the gradient vanishing problem in deep neural networks, allowing information to be transmitted more smoothly within the network.

[0019] In stage 3, a 320×320×64 feature image is input to the second ELAN block. Some standard convolutions in the second ELAN block are replaced with multi-scale deformable convolutions, which introduce learnable offsets during the convolution process. For the multi-scale deformable convolution, an offset field is first generated through a 1×1 convolution operation with weights W. offset (These weights are automatically learned during the model's training process.) The input feature map is X(p), which is a 320×320×64 feature image, where p represents the spatial coordinates on the feature map. The formula for generating the offset field of the multi-scale deformable convolution is: Δp k =W offset *X(p) In the formula, Δp k W is the offset. offset is the 1×1 convolution weight; X(p) is the input feature map; The offset Δp at each position p is calculated using the offset field generation formula. k The offset Δp here k It is a two-dimensional vector indicating the offset of the convolution kernel's sampling position on the input feature map; subsequently, the shape and position of the convolution kernel are adjusted: traditional standard convolution kernels have a fixed shape and position, while multi-scale deformable convolution adjusts the kernel's shape and position based on the generated offset Δp. k The shape and position of the convolutional kernel are adaptively adjusted. Specifically, the sampling position of the convolutional kernel on the input feature map is no longer limited to regular grid points, but is offset within a preset range of the fish target according to the offset amount. This allows the convolutional kernel to better adapt to the medium-detail features of the large yellow croaker, such as the fish outline, because the outline of the large yellow croaker may not be a regular shape. This adaptive adjustment allows the convolutional kernel to capture the outline information more accurately. Then, the sampling process is performed: for each sampling point on the convolutional kernel, according to the corresponding offset amount Δp... k The process involves finding new sampling locations on the input feature map X(p) and extracting feature values ​​from these locations for subsequent convolution calculations. Next, the convolution calculation is performed: after sampling, the sampled feature values ​​are weighted and summed with the weights of the convolution kernel according to standard convolution calculation methods to obtain the output feature value at that location. Finally, feature extraction is performed: by performing the above convolution calculation on all locations across the entire feature map, the output feature of the second ELAN block is obtained, i.e., the second feature 160×160×128. Because multi-scale deformable convolution can adaptively adjust the convolution kernel, the extracted features have richer detailed information than the first feature, and can more accurately describe some features of the large yellow croaker, such as the fish's outline and local texture.

[0020] In stage 4, the 160×160×128 feature image is input to the third ELAN block. The third ELAN block also uses multi-scale deformable convolution to replace part of the standard convolution, with the offset calculation principle the same as the second feature. As the network depth increases, the third ELAN block can extract high-level semantic features such as the overall morphology of the fish, resulting in the third feature of 80×80×256. The introduction of multi-scale deformable convolution allows the convolution kernel to adapt to the feature changes of the large yellow croaker at different levels of abstraction, further enhancing the ability to extract features of the large yellow croaker. The third feature contains more semantic information about the overall morphology and category of the large yellow croaker, which is of great significance for the identification and classification of the large yellow croaker.

[0021] In stage 5, the 80×80×256 feature image is input to the fourth ELAN block. The fourth ELAN block also uses multi-scale deformable convolution for feature extraction, with the offset calculation principle the same as the second feature. At this level, the model further mines deeper features of the large yellow croaker, such as distant fish bodies. These features typically have stronger representational capabilities and can better distinguish the large yellow croaker from other objects, thus obtaining the fourth feature, 40×40×512. Multi-scale deformable convolution adaptively adjusts the convolution kernel, enabling the model to cope with various changes in the large yellow croaker in complex underwater scenes, such as pose and scale changes.

[0022] In the ELAN block of phases 3-5, the present invention replaces the original 3×3 standard convolution with multi-scale deformable convolution, so that it can dynamically adjust the sampling position of the convolution kernel through learning. For example, when the large yellow croaker bends, the convolution kernel automatically shifts in the bending direction, which better fits the shape of the fish.

[0023] This invention employs a hierarchical feature extraction method to comprehensively capture the feature information of large yellow croaker at different scales and different levels of abstraction, which helps the model to more accurately identify the target of large yellow croaker. Multi-scale deformable convolution introduces learnable offsets, enabling the convolution kernel to flexibly adapt to the different shapes and postures of large yellow croaker, thereby extracting more discriminative features. This helps the model to accurately distinguish large yellow croaker from other objects in complex underwater environments, such as when there is turbid water or background interference, thus improving the robustness of target detection.

[0024] In one embodiment, the neck feature fusion network layer includes a PAN block, a channel attention layer, and a spatial attention layer; the PAN block includes an upsampling path and a downsampling path; wherein, The step of performing target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker also includes: In the upsampling path, the fourth feature is upsampled and then concatenated with the third feature to obtain the fifth feature, and the fifth feature is upsampled and then concatenated with the second feature to obtain the sixth feature; In the downsampling path, the sixth feature is downsampled and then concatenated with the fifth feature to obtain the seventh feature, and the seventh feature is downsampled and then concatenated with the fourth feature to obtain the eighth feature; The sixth, seventh, and eighth features are all input into the channel attention layer and the spatial attention layer to perform weighted processing on the corresponding features in the channel dimension and the spatial dimension, respectively, to obtain the first enhanced feature, the second enhanced feature, and the third enhanced feature.

[0025] Specifically, the neck feature fusion network layer in this invention inherits the PAN (Path Aggregation Network) structure of YOLOv7. The PAN block achieves multi-scale feature fusion through a bidirectional path: top-down (high-level features transmit semantics to low-level features) + bottom-up (low-level features supplement details to high-level features). To address the problem of strong underwater background interference (such as bubbles) and the easy submersion of fish features, the improved YOLOv7 network model in this invention incorporates a channel-spatial dual attention mechanism into the PAN path. Channel attention layers and spatial attention layers are embedded after the original PAN block, enabling the model to focus more on features that significantly contribute to the detection of large yellow croaker when fusing features at different levels, thus improving the feature fusion effect and providing higher-quality feature representations for subsequent detection outputs. In the upsampling path, the fourth feature (40×40×512) is taken as input and upsampled by 2 times to enlarge the 40×40 feature map to 80×80. It is then concatenated with the third feature output from the backbone stage 4 and compressed to 256 channels through convolution to enhance the semantic information of small targets, resulting in the fifth feature (80×80×256). Next, the fifth feature is upsampled by 2 times to enlarge it to 160×160 and concatenated with the second feature output from the backbone stage 3 (160×160×128). It is then compressed to 128 channels through convolution to fuse detailed features, resulting in the sixth feature (160×160×128).

[0026] In the downsampling path (bottom-up), the sixth feature (160×160×128) is taken as input and downsampled by 2 times to transform the 160×160 feature map into 80×80. It is then concatenated with the fifth feature, the intermediate result of the upsampling path, to form 80×80×256. After convolution, the channel is compressed to 256 to supplement the mid-level localization information, resulting in the seventh feature (80×80×256). Next, the seventh feature is downsampled by 2 times to form 40×40 and concatenated with the fourth feature (40×40×512) output from the backbone stage 5. After convolution, the channel is compressed to 512 to optimize large target detection, resulting in the eighth feature (40×40×512).

[0027] The sixth feature is input into the channel attention layer and the spatial attention layer, and weighted in both the channel and spatial dimensions to obtain the first enhanced feature. This invention uses the sixth feature as an example to illustrate the feature processing procedures of the channel attention layer and the spatial attention layer: In the channel attention layer, the sixth feature 160×160×128 is processed by global average pooling, which averages the values ​​of all spatial locations in each channel to obtain a channel statistical vector. This vector contains global statistical information for each channel, and the feature map is compressed to 1×1×128. The channel statistical vector is then input into the fully connected layer for dimensionality reduction and expansion, outputting a 1×1×128 feature map. The output of the fully connected layer is then normalized using the sigmoid function to obtain a channel weight vector of 1×1×128. Finally, the channel weight vector is multiplied channel by channel with the eighth feature 160×160×128 to obtain the ninth feature 160×160×128.

[0028] In the spatial attention layer, the ninth feature 160×160×128 is globally averaged along both the width and height directions. This operation averages the values ​​of all spatial locations in each channel to obtain a vector of channel length, which contains global statistical information for each channel. The features in the two directions are then concatenated along the channel dimension to obtain a 160×160×256 feature. This feature is then processed by a 3×3 convolution to obtain a single-channel spatial weight map 160×160×1. The single-channel spatial weight map is then normalized using the sigmoid function to generate a spatial weight map (focusing on the area where the fish is located, with a value range of [0,1]). Values ​​close to 1 indicate key areas (such as the outline of the fish), and values ​​close to 0 indicate background areas (such as water). Finally, the spatial weight map is multiplied element-wise with the ninth feature 160×160×128 to obtain the first enhanced feature 160×160×128.

[0029] Similarly, after processing the seventh and eighth features as described above, we can obtain the second enhanced feature 80×80×256 and the third enhanced feature 40×40×512. Each enhanced feature map contains both high-level semantics and low-level details, and the background interference is suppressed through the attention mechanism.

[0030] This invention utilizes the upsampling and downsampling paths of the PAN block in the neck feature fusion network layer to capture features at different scales, and achieves multi-scale information complementarity through splicing operations, thereby improving the detection capability of small targets. It employs a channel attention mechanism to highlight important feature channels for large yellow croaker identification, and a spatial attention mechanism to focus on the key spatial region where the large yellow croaker is located, reducing the influence of irrelevant information such as water turbidity and background interference. It adopts a channel-first, then spatial processing approach, first optimizing feature semantics (channel dimension) and then focusing on key regions (spatial dimension), forming hierarchical feature enhancement, which significantly improves the detection accuracy of large yellow croaker. The introduction of a dual attention mechanism enhances the model's ability to extract key features of large yellow croaker, further improving the detection rate and tracking accuracy in complex underwater scenarios.

[0031] In one embodiment, the head detection output layer includes a first parallel detection branch, a second parallel detection branch, and a third parallel detection branch; the first parallel detection branch includes a first detection sub-branch and a first keypoint prediction sub-branch; the second parallel detection branch includes a second detection sub-branch and a second keypoint prediction sub-branch; the third parallel detection branch includes a third detection sub-branch and a third keypoint prediction sub-branch; wherein, The step of performing target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker also includes: The first enhanced feature is input into the first detection sub-branch and the first key point prediction sub-branch respectively for processing to obtain the first bounding box coordinate data and the first key point coordinate data. The second enhanced feature is input into the second detection sub-branch and the second key point prediction sub-branch respectively for processing to obtain the second bounding box coordinate data and the second key point coordinate data. The third enhanced feature is input into the third detection sub-branch and the third key point prediction sub-branch respectively for processing to obtain the third bounding box coordinate data and the third key point coordinate data; The coordinate data of the first bounding box, the second bounding box, the third bounding box, the first keypoint, the second keypoint, and the third keypoint are combined to generate the coordinate data of the large yellow croaker.

[0032] Specifically, the head detection output layer in this invention is based on a multi-branch detection head design using the YOLOv7 network model. This is achieved by adding a parallel keypoint prediction sub-branch after each multi-branch detection head in the YOLOv7 model. The original YOLOv7 head contains three parallel detection branches (corresponding to 160×160, 80×80, and 40×40 feature maps output from the neck, respectively). Each branch consists of a detection sub-branch and a keypoint prediction sub-branch, outputting bounding boxes, confidence scores, and categories. This invention improves the YOLOv7 network model to obtain biomechanical keypoints of the large yellow croaker (used to calculate swimming posture). A keypoint prediction sub-branch is added after each detection branch, forming a multi-task output of "detection + keypoint". Based on the core biomechanical features reflecting the swimming posture of the large yellow croaker, this invention selects the front end of the head (the foremost point of the mouth), the midpoint of the body axis (the midpoint of the line connecting the front end of the head and the root of the caudal fin), the root of the caudal fin (the starting point where the caudal fin connects to the body), and the tip of the caudal fin (the very end of the caudal fin) as keypoints.

[0033] The detection sub-branch (inherited from YOLOv7) consists of three 3×3 convolutional layers (the middle layers are activated with SiLU, and the output layer is unactivated). The output dimension is (4+1+1)×K for each feature map grid, where 4 is the bounding box offset (Δx, Δy, Δw, Δh). The offset is predicted based on "anchor boxes," with each anchor box corresponding to a grid on the feature map. The origin is the top-left corner of the image, and the center coordinates are: x = (top-left corner x + sigmoid(Δx)) × grid stride; y = (top-left corner y + sigmoid(Δy)) × grid stride. Step size: sigmoid ensures the offset is between 0 and 1, confined within the grid; Width and height: w = original anchor frame w × exp(Δw); h = original anchor frame h × exp(Δh), exp ensures width and height are positive); 1 is the confidence score (the probability that the grid contains the large yellow croaker, output by the sigmoid function); 1 is the class probability (the softmax function is used to predict the class probability. Since only "large yellow croaker" needs to be detected, it is a binary classification: "large yellow croaker" or "background", and the class with the highest probability is taken as the output label); K is the number of anchor frames for each grid (default 3, adapting to different shapes of fish).

[0034] The keypoint prediction sub-branch runs in parallel with the detection sub-branch and consists of two 3×3 convolutional layers (SiLU activation). The output dimension of this branch is as follows: for each feature map grid, the output keypoint coordinate data is 2×N×K, where: 2 is the x and y coordinate offset of each keypoint (which can also be considered as two-dimensional pixel coordinates). With the bounding box as a reference, the offset (dx, dy) of each keypoint relative to the upper left corner (x1, y1) of the bounding box is predicted. The final keypoint coordinates are u=x1+dx; v=y1+dy, where dx and dy are normalized to 0-1 by sigmoid and then multiplied by the width and height of the bounding box to ensure that the keypoint is within the box; N is the number of keypoints (default 4, i.e., the front of the head, the midpoint of the body axis, the root of the caudal fin, and the tip of the caudal fin); K is the number of anchor boxes shared with the detection sub-branch (3).

[0035] Based on the above principles, the first enhanced feature (160×160×128) is input into the first parallel detection branch, which is the small-scale detection branch (responsible for predicting small anchor boxes and matching distant targets). Through its first detection sub-branch and first keypoint prediction sub-branch, the first bounding box coordinate data and the first keypoint coordinate data are obtained. The second enhanced feature (80×80×256) is input into the second parallel detection branch, which is the medium-scale detection branch (responsible for predicting medium anchor boxes and matching medium-distance targets). Through its second detection sub-branch and second keypoint prediction sub-branch, the second bounding box coordinate data and the second keypoint coordinate data are obtained. The third enhanced feature (40×40×512) is input into the third parallel detection branch, which is the large-scale detection branch (responsible for predicting large anchor boxes and matching near targets). Through its third detection sub-branch and third keypoint prediction sub-branch, the third bounding box coordinate data and the third keypoint coordinate data are obtained. The bounding box coordinate data and keypoint coordinate data mentioned above are processed by non-maximum suppression (NMS). Specifically, for multiple overlapping candidate boxes of the same large yellow croaker, the bounding box with the highest confidence is retained, and duplicate boxes are removed. Then, keypoint filtering is performed on the processed bounding boxes, retaining only the keypoint coordinates corresponding to the bounding boxes after NMS, ensuring a one-to-one correspondence between the keypoints and bounding boxes for each large yellow croaker. Finally, the filtered keypoint coordinates are concatenated and organized, and the processed bounding box coordinates are also concatenated and organized to obtain the bounding box coordinate data and keypoint coordinate data of the target large yellow croaker. Furthermore, after feature processing at the head detection output layer, this improved model can also output a tracking ID for the large yellow croaker, assigning a unique ID to the same large yellow croaker in consecutive frames, thereby achieving cross-frame tracking.

[0036] This invention employs a joint design of a multi-branch detection head and a keypoint prediction sub-branch. Three parallel detection branches process enhanced features at different scales, covering the detection needs of large yellow croaker at various distances and improving compatibility with both small targets (far-distance fish) and large targets (close-range fish). Each detection branch adds an additional keypoint prediction sub-branch, directly regressing the coordinates of key points such as the fish head and tail, solving the problem of traditional bounding box detection's insensitivity to fish pose and improving the ability to discriminate overlapping or occluded scenes. The joint output of bounding box coordinate data and keypoint coordinate data retains the compatibility of traditional detection while providing refined pose information, facilitating subsequent behavior analysis.

[0037] In one embodiment, the improved YOLOv7 network model used in this invention is compared with other existing models, and the comparison results are shown in the table below: Table 1 Comparison of Results

[0038] As shown in Table 1 above, the detection accuracy of the improved YOLOv7 network model used in this invention is 8.6% higher than that of YOLOv5s and 5.4% higher than that of the original YOLOv7. The 63 FPS meets the real-time requirements (>30 FPS), which is slightly lower than that of the original YOLOv7 (71 FPS), but significantly higher than that of the two-stage model Faster R-CNN (22 FPS). The number of parameters is 28.4M, which is 23% less than that of the original YOLOv7 (36.9M→28.4M), and better than that of Faster R-CNN (41.5M). It maintains a real-time accuracy of 63 FPS with a high accuracy of 90.7% mAP and a 23% reduction in the number of parameters, achieving a balance between "high accuracy without sacrificing speed".

[0039] In one embodiment, the step of performing a three-level biomechanical verification of the coordinate data based on the real-time water turbidity data, linking body length with movement limits, to obtain verified coordinate data includes: The coordinate data is transformed to obtain three-dimensional coordinate data, and the body length of the large yellow croaker is determined based on the three-dimensional coordinate data; A biomechanical limit database of the large yellow croaker at multiple growth stages is constructed, and the biomechanical limit database is matched with the body length to achieve a first-level verification of the biomechanical constraints of the large yellow croaker. The three-dimensional coordinate data with a successful matching result is used as the verification coordinate data. When the matching result is determined to be a failure, a secondary compensation mechanism is triggered. The coordinate data is input into a pre-constructed grid velocity mapping table for spatial matching to obtain a velocity vector. This vector is then combined with the coordinate data of a preset frame to predict the theoretical coordinates of the current frame. Based on the real-time water turbidity data, residual correction and coordinate correction are performed on the theoretical coordinates to obtain corrected coordinates, which are then used as the verification coordinate data. When the corrected coordinates are detected to be invalid, a three-level completion mechanism is triggered to obtain the three-dimensional valid coordinate data of the large yellow croaker and fit it with a third-order Bézier curve to obtain the individual trajectory equation for extrapolation prediction, generating predicted coordinates. The predicted coordinates are then subjected to the first-level verification of the biomechanical constraints. If the match is successful, the predicted coordinates are used as the verification coordinate data of the large yellow croaker; otherwise, the control point weights of the Bézier curve are adjusted, and the prediction is repeated until the constraints are met.

[0040] Specifically, this invention converts the two-dimensional coordinates of key points output by the model into three dimensions based on the calibration parameters of the dual-sensor visual acquisition system (internal parameters of the wide-angle camera, spatial positional relationship between the laser rangefinder and the camera).

[0041] In this embodiment, the transformation of the coordinate data includes: The Z coordinate of the large yellow croaker is determined by the distance from the key point of the large yellow croaker to the narrow-beam laser rangefinder; Based on the Z coordinate, the coordinate data is distorted by using the calibration parameters of the underwater wide-angle camera to obtain the X and Y coordinates of the large yellow croaker, and then combined with the Z coordinate to generate the three-dimensional coordinate data of the large yellow croaker.

[0042] Specifically, this invention uses the straight-line distance from the narrow-beam laser rangefinder to a key point on the large yellow croaker as the Z-coordinate of that key point. Then, based on the Z-coordinate, and combined with the intrinsic parameters (focal length, principal point coordinates) of the underwater wide-angle camera, distortion correction is performed on the corresponding two-dimensional pixel coordinates, i.e., the coordinate data, to obtain the X and Y coordinates of the key point. The correction process is expressed by the following formula: X=(u-cx)×Z / f Y=(v-cy)×Z / f In the formula, X and Y are the X and Y coordinates of the key point, respectively; u and v are the two-dimensional pixel coordinates of the key point, respectively; cx and cy are the pixel positions of the camera optical axis in the image, i.e., the principal point coordinates, respectively; Z is the Z coordinate of the key point; and f is the focal length of the camera.

[0043] Finally, combining the X, Y, and Z coordinates of the key points yields the three-dimensional coordinate data of that point, which is the transformation result. Based on the above principle, the three-dimensional coordinate data of each key point within each large yellow croaker can be calculated.

[0044] This invention provides the distance (Z coordinate) from key points to the sensor directly through a narrow-beam laser rangefinder, avoiding the error accumulation of pure visual depth estimation and significantly improving the accuracy of 3D reconstruction. Although underwater wide-angle cameras have a wide field of view, they suffer from severe distortion. By using the absolute distance constraint of the laser rangefinder, the distortion can be corrected and the X and Y coordinates can be accurately recovered, balancing the field of view and measurement accuracy. The 3D reconstruction is decomposed into two steps: laser ranging (Z coordinate) and visual calculation (X and Y coordinates), reducing computational complexity and making it suitable for real-time processing.

[0045] Next, by improving the bounding box height H (pixel value) of the large yellow croaker detected in each frame output by the YOLOv7 model, and combining the calibration parameters of the dual-sensor system (1 pixel corresponds to the actual distance k, such as k=0.1cm / pixel), the actual body length L=H×k×1.2 (1.2 is the conversion factor of the offline calibration bounding box height - the actual body length) can be calculated, and the body length of the large yellow croaker in each frame can be obtained.

[0046] As a typical migratory fish, the large yellow croaker has physiological limits in its vertical (Z-axis) and horizontal (X / Y-axis) movement speeds (limited by muscle explosive power and body size). This invention constructs a database of the biological movement limits of the large yellow croaker at multiple growth stages to calibrate these limits offline, thereby determining in real time whether the current coordinates conform to the biological movement laws and eliminating obvious outliers. This invention selects 50 large yellow croakers at each of the three growth stages (juveniles: 10-15cm; subadults: 20-25cm; adults: 30-35cm). Swimming data for 30 minutes is recorded in a controlled aquatic environment (flow rate 0.2m / s, clear water) using a high-speed camera (200fps). The X / Y / Z coordinates of each frame are extracted, and the displacement between adjacent frames is calculated to obtain the movement speed (displacement divided by the time interval, preferably 0.005s). The Vx, Vy, and Vz values ​​for each growth stage are statistically analyzed, and the 99.9th percentile is used as the physiological limit threshold, stored in the system database to form a biological movement limit database. The physiological limit thresholds are shown in the table below. Table 2 Physiological Limit Thresholds

[0047] Next, the biomechanical limit database is matched with the body length to perform a first-level verification of the biomechanical constraints of the large yellow croaker, and the coordinate displacement of the large yellow croaker in the current frame and the previous frame is calculated to obtain Vx, Vy, and Vz. If it is within 1.2 times the limit threshold (leaving a 20% fluctuation space), it means that the matching is successful and is used as the verification coordinate data.

[0048] If any velocity exceeds 1.2 times the corresponding limit threshold, the matching is deemed a failure, and the current coordinates are treated as an outlier, triggering a secondary compensation mechanism: A grid-velocity mapping table for the target water area is pre-established, dividing the target water area into a three-dimensional grid according to the ADCP (Acoustic Doppler Current Profiler) deployment spacing (2m×2m×1m). Each grid corresponds to a unique ADCP number, storing the real-time velocity data of that ADCP (Ux: X-direction velocity, Uy: Y-direction velocity, Uz: Z-direction velocity). Then, any key point of the large yellow croaker in the current frame... The coordinate data, i.e., the two-dimensional pixel coordinates (u,v), is mapped to a predefined grid-flow velocity mapping table. The ADCP number corresponding to the grid cell is found, its flow velocity vector is read, and based on the fish body coordinates (X1,Y1,Z1), (X2,Y2,Z2), and (X3,Y3,Z3) of the previous three frames, the historical average fish propulsion velocity Vxavg=(X3-X1) / (2Δt) is calculated (the calculation process for Vyavg and Vzavg is similar to Vxavg). Combined with the current flow velocity vector, the current frame's velocity is predicted. Theoretical coordinates (Xpred = X³ + (Vxavg + Ux) × Δt, Ypred and Zpred are calculated similarly to Xpred); The theoretical coordinates are corrected for residuals based on real-time turbidity data of the target water area: If the current laser-measured Z-coordinate Zmeas has a deviation (i.e., Zmeas - Zpred > 0.1m), then the Z-coordinate is corrected using Zcorr = Zpred + α × (Zmeas - Zpred), where α is a dynamic compensation coefficient; when turbidity < 200 NTU, α... =0.8, prioritize trusting laser measurement; linear transition when 200≤NTU≤300, α=0.8-(0.5*(NTU-200) / 100), when turbidity>300NTU, α=0.3, prioritize trusting theoretical prediction; no correction is needed if there is no deviation; based on the corrected Zcorr, re-synchronize the X / Y coordinates through camera intrinsic parameters (i.e., distortion correction process during coordinate transformation) to ensure that the X / Y / Z coordinates are consistent with the correlation of the water flow environment, and generate corrected coordinates as verification coordinate data.

[0049] If the laser rangefinder returns no signal at all (i.e., Zcorr is NULL), it indicates that the corrected coordinates are invalid, triggering a three-level completion mechanism: For the large yellow croaker missing the Z coordinate, extract its three-dimensional coordinate data from the first 5 frames as control points for the Bézier curve. Using a third-order Bézier curve (4 control points), fit the horizontal (XY) and vertical (Zt) motion trajectories using the coordinates of the first 4 frames (P0, P1, P2, P3) to obtain the individual trajectory equation: X(t)=P0x×(1-t)³+3P1x×(1-t)²t+3P2x×(1-t)t²+P3x×t³ Y(t)=P0y×(1-t)³+3P1y×(1-t)²t+3P2y×(1-t)t²+P3y×t³ Z (t)=P0z×(1-t)³+3P1z×(1-t)²t+3P2z×(1-t)t²+P3z×t³ In the formula, t is a time parameter (t∈[0,1], corresponding to the time interval of the first 4 frames).

[0050] Let the missing frame be the 6th frame, and the time interval Δt = 1 / frame rate (e.g., Δt = 0.02s when the frame rate is 50fps). Then the t value corresponding to the 6th frame is t5 = 5 × Δt / (4 × Δt) = 1.25 (based on the time span of the first 4 frames, 4Δt, extrapolated to the 5th Δt). Substitute t5 into the individual trajectory equation to calculate the predicted coordinates of the 6th frame. Perform a first-level biomechanical constraint check on the predicted coordinates. If it meets the limit range, it is used as the coordinates of the missing frame, i.e., the check coordinate data. If it exceeds the limit, it means that the extrapolation is too aggressive. Then adjust the control point weights of the Bézier curve (using simple linear extrapolation or halving the speed) and re-predict until it meets the constraints.

[0051] Based on the principle that no fish's instantaneous movement speed can exceed its physiological limits, this invention introduces biomechanical constraints to quickly eliminate obvious outliers. By fusing water flow data, physical laws are used to constrain and compensate for multimodal residuals, correcting deviations caused by instantaneous errors of a single sensor. When the laser completely fails, the current position of the individual fish is predicted based on its historical movement trajectory. A three-level verification mechanism is adopted to ensure that the system can still provide continuous and reliable data output even in the event of a temporary failure of any single sensor, greatly improving the system's availability and reliability.

[0052] S4. Determine the real-time attitude characteristics of the large yellow croaker based on the verification coordinate data, and optimize the real-time attitude characteristics through the flow velocity phase joint mechanism to obtain optimized attitude characteristics; In one embodiment, the verification coordinate data includes the three-dimensional coordinates of the front end of the large yellow croaker's head, the three-dimensional coordinates of the midpoint of its body axis, the three-dimensional coordinates of the base of its caudal fin, and the three-dimensional coordinates of the tip of its caudal fin; wherein, determining the real-time posture characteristics of the large yellow croaker based on the verification coordinate data includes: The relative displacement of the three-dimensional coordinates of the tip of the caudal fin with respect to the three-dimensional coordinates of the root of the caudal fin in a single frame is calculated, so as to perform sliding window analysis on the relative displacement in consecutive frames, and the maximum value in the window is taken as the caudal fin sway amplitude for that consecutive frame period. Based on the three-dimensional coordinates of the front end of the head, the three-dimensional coordinates of the midpoint of the body axis, and the three-dimensional coordinates of the base of the tail fin, the body axis curve of the large yellow croaker is constructed, and the body axis curvature is determined based on the body axis curve, which, together with the tail fin swing amplitude, serves as the real-time posture feature.

[0053] Specifically, this invention calculates the core characteristic reflecting the swimming power output of the large yellow croaker—the tail fin amplitude—based on multiple verified three-dimensional coordinates. This amplitude is the maximum displacement distance between the tip of the tail fin and the base of the tail fin during continuous movement. The three-dimensional displacement vector of the tail fin tip relative to the base of the tail fin at a certain moment in a single frame image is calculated based on the three-dimensional coordinates of the tail fin tip and the base of the tail fin. The spatial distance of this three-dimensional displacement vector, also known as the Euclidean norm or Euclidean distance, is then calculated to obtain the relative displacement of the large yellow croaker in a single frame. Next, a sliding window analysis is performed on the displacement magnitude of N consecutive frames (e.g., 50 frames within 1 second), and the maximum value within the window is taken as the tail fin amplitude for that consecutive frame period. The window size is set according to the swimming frequency of the large yellow croaker, typically 0.2-0.5 seconds, to ensure coverage of a complete tail fin swing cycle.

[0054] The curvature of the body axis, reflecting the degree of bending of the large yellow croaker's trunk (the more pronounced the bending, the stronger the swimming power), is calculated using three-dimensional coordinates. This curvature is the curvature of the fish's body axis (the line connecting the head to the base of the caudal fin). Since the fish's movement is primarily in the horizontal plane (ignoring minor vertical fluctuations), a quadratic curve is fitted using XY plane coordinates (the fish's bending is approximated as a planar curve). The equation of the body axis curve is given by... By substituting the three-dimensional coordinates of the front of the head, the midpoint of the body axis, and the root of the caudal fin into the equation of the body axis curve, the coefficients of the equation can be obtained. After knowing the coefficients of the equation of the body axis curve, the curvature of the body axis curve can be calculated according to the curvature formula and any one of the three-dimensional coordinates of the front of the head, the midpoint of the body axis, and the root of the caudal fin. The caudal fin swing amplitude and body axis curvature of the same large yellow croaker can be used as the real-time posture characteristics of the large yellow croaker.

[0055] Based on the calibration parameters of a dual-sensor vision system, this invention converts two-dimensional pixel coordinates into three-dimensional world coordinates, eliminates the influence of perspective distortion, and accurately calculates the posture features of key points on the fish body. By analyzing the relative displacement of the tail fin in consecutive frames through a sliding window, the key motion feature of tail fin sway is extracted, reflecting the swimming intensity and rhythm of the fish. The three-dimensional coordinate transformation and sliding window statistics effectively suppress visual noise and instantaneous motion interference, improving feature stability.

[0056] In one embodiment, optimizing the real-time attitude features through a combined flow velocity and phase mechanism to obtain optimized attitude features includes: A fingerprint mapping table of tail wagging cycle and swimming state of various yellow croakers in the target water area is pre-constructed; the tail wagging cycle is determined based on the three-dimensional coordinates of the tail fin tip, and the swimming state is determined based on the body axis curvature. The swimming state of the large yellow croaker in the current frame is determined to match it with the fingerprint mapping table, and the current tail wagging period is obtained to dynamically adjust the frame number of the sliding window to obtain the target sliding window. Phase features within the target sliding window are extracted for phase segmentation to obtain segmentation results. Weights are dynamically assigned to the segmentation results according to the phase type for weighted calculation to obtain weighted swing amplitude. Based on the body axis curvature and the pre-built swing curvature correlation model, the theoretical swing of the current frame is determined, and the deviation between the weighted swing and the theoretical swing is judged to realize the attitude linkage verification of the real-time attitude features. When the condition is determined to be normal, the weighted swing amplitude and the body axis curvature are used as the optimized posture features; when the condition is determined to be abnormal, the body axis curvature is smoothed and corrected, and the correction result and the theoretical swing amplitude are used as the optimized posture features.

[0057] Specifically, the tail wagging cycle of different large yellow croakers varies from person to person, and the cycle of the same fish also changes under different states (calm, feeding, stress). This invention establishes a fingerprint mapping table between the tail wagging cycle and the swimming state, matches the current fish's tail wagging cycle in real time, and dynamically adjusts the sliding window to optimize the real-time posture features, avoiding missed or over-testing by the fixed window.

[0058] For large yellow croaker in the target waters, individual tracking IDs were marked using an improved OLOv7 tracking ID. Swimming data (including the three-dimensional coordinates of the caudal fin tip and base) was continuously collected for each ID over 30 minutes. For the caudal fin tip coordinates of each ID, the displacement sequence in the Z-direction (or Y-direction) was extracted and subjected to Fourier transform to obtain the frequency spectrum. The frequency with the largest amplitude in the frequency spectrum was taken as the main wagging frequency ω0, and the tail wagging period was T=1 / ω0. The state was determined by the mean curvature of the body axis (calm: mean curvature < 0.1cm). -1 Foraging: 0.1-0.2cm -1 Stress: >0.2cm -1 ), and establish a fingerprint mapping table between the tail cycle and the parade state for each ID.

[0059] For the large yellow croaker in the current frame, calculate the average body axis curvature Cavg of the previous 10 frames, match the corresponding state in the fingerprint database, and call the corresponding tail-wagging period Tcurr of the ID according to the current state to dynamically adjust the sliding window size to obtain the target sliding window: window frame number N = frame rate × Tcurr. Every 30 seconds, recalculate the main wagging frequency ω0 of the current ID and update Tcurr to ensure that the window size is synchronized with the real-time tail-wagging period, realize the flow velocity correction of the dynamic window, and solve the problem of missed window detection caused by the increased tail-wagging frequency of large yellow croaker under high flow velocity.

[0060] The tail-wagging process of large yellow croaker is divided into a power-generating phase (the tail fin swings to one side, with large displacement and significant propulsion) and a recovery phase (the tail fin returns to its original position, with small displacement and minimal contribution). This invention improves the accuracy of tail-wagging amplitude calculation by identifying the tail-wagging phase and assigning high weight to frames in the power-generating phase and low weight to frames in the recovery phase: For N frames of data within the target sliding window, the displacement Di of the tail fin tip relative to the tail fin root in each frame is calculated, and the first derivative of the displacement is calculated as a phase feature; when the phase feature is greater than 0 and greater than the average displacement Davg within the window, it is determined to be in the power-generating phase; when the phase feature is less than 0 and less than Davg, it is determined to be in the recovery phase; when the phase feature is equal to 0, it is determined to be at the boundary between the power-generating phase and the recovery phase; then, it is assigned a type based on the phase. Weights: Weight of the power-up frame: wi=1.2×e^(0.5×(Di / Dmax)), where Dmax is the maximum displacement within the window, with a weight range of 1.2~2.4; Weight of the swing-back frame: wi=0.8×e^(-0.3×(Davg / Di)) (Di≠0), with a weight range of 0.4~0.8; Weight of the phase inflection point: wi=1.0 (neutral weight); Then calculate the total weight within the window, normalize each wi, and after obtaining the normalized weight, perform a weighted summation of the displacements in each frame to highlight the contribution of the displacement during the power-up period and reduce the interference during the swing-back period, thereby obtaining the weighted swing amplitude. This achieves phase-weighted flow velocity constraints and ensures that the weighted swing amplitude can reflect the influence of water flow assistance / resistance on power-up efficiency.

[0061] The swimming postures of large yellow croaker exhibit a correlation: the greater the tail fin amplitude, the greater the body axis curvature (more pronounced trunk bending during exertion), and the two are positively correlated. This invention establishes a correlation model between amplitude and curvature to eliminate single posture anomalies caused by noise. For large yellow croaker at three growth stages, 1000 sets of "tail fin amplitude A - body axis curvature C" data were collected to ensure coverage of different swimming states. A linear regression model was used to fit the relationship between A and C: A = k × C + b, where k is the correlation coefficient (k = 5 for juveniles, k = 8 for subadults, and k = 12 for adults), b is the intercept (uniformly set to 0.05 cm), and the goodness of fit is... The correlation coefficient must be ≥0.9 (to ensure significant correlation), and the model parameters are stored in the system. For the current frame, based on the tail fin amplitude Acurr and the body axis curvature Ccurr, and combined with the correlation model, the theoretical amplitude Atheo = k × Ccurr + b (k is matched according to the growth stage) is calculated; the deviation rate δ = |Acurr - Atheo| / Atheo is calculated to realize the attitude linkage verification of real-time attitude features.

[0062] If δ ≤ 15% (normal fluctuation range), both Acurr and Ccurr are valid and used as optimized posture features; if δ > 15%, it is considered noise and a correction is triggered: if Acurr is abnormal (Acurr > Atheo × 1.15): replace Acurr with Atheo; if Ccurr is abnormal (Ccurr < Acurr / kb × 1.15): replace Ccurr with (Ccurr + (Acurr - b) / k) / 2 (smoothing correction). The corrected Acurr and Ccurr must again meet the biological movement limits (e.g., adult Acurr ≤ 15cm, Ccurr ≤ 0.3cm). -1 This ensures that the posture characteristics conform to physiological laws, and uses the correction results as optimized posture characteristics to achieve flow rate fallback for swing amplitude and curvature verification.

[0063] This invention establishes individual motion fingerprints, respecting the individual differences of organisms, making the feature extraction method more targeted and naturally more accurate; the phase weighting mechanism delves into the motion cycle, distinguishing between effective and ineffective motion, and the extracted features are more representative of the actual propulsion force generation process; the linkage verification mechanism links two originally independently calculated features, using the biological physical constraints for cross-validation, greatly enhancing the system's anti-interference ability and the reliability of the output results; the three-level mechanism is interlocked and ultimately forms a large closed loop with biomechanical constraints, constructing a highly intelligent attitude feature optimization algorithm.

[0064] S5. Construct an attitude-fluid coupling model based on the optimized attitude features and the real-time environmental flow velocity data, and determine the swimming speed of the large yellow croaker through the attitude-fluid coupling model; In one embodiment, step S5 includes: The real-time environmental flow velocity data is interpolated to obtain the water flow velocity vector of the target water area; Based on the optimized posture characteristics, the propulsion speed of the large yellow croaker relative to the water flow is determined; The attitude fluid coupling model is constructed based on the water flow velocity vector and the propulsion velocity, and the swimming speed of the large yellow croaker is determined based on the attitude fluid coupling model.

[0065] Specifically, this invention collects real-time environmental flow velocity data at different spatial points by deploying multiple flow sensors (such as Doppler current meters) in the target water area. The data includes the flow velocity (magnitude: m / s) and direction (angle: the angle with due north). The invention then performs interpolation processing on these environmental flow velocity data to generate a flow velocity vector for the target water area, i.e., a flow velocity vector at any location.

[0066] Based on the historical posture characteristics of the same historical period corresponding to the real-time posture characteristics of the large yellow croaker, the propulsion speed of the large yellow croaker relative to the water flow is obtained by fitting an empirical formula: v_fish_to_water = k1×A + k2×C + ε (A and C are the tail fin amplitude and body axis curvature, respectively; k1 and k2 are coefficients calibrated through experiments; ε is an error correction term, determined experimentally, preferably 10). -5 ).

[0067] Finally, the water flow velocity vector and propulsion velocity are added together to obtain the true swimming speed. The entire velocity summation formula is the attitude-fluid coupling model. Substituting the real-time attitude characteristics of the large yellow croaker into the attitude-fluid coupling model, the true swimming speed of the large yellow croaker can be output. The above embodiment calculates the true swimming speed of the large yellow croaker using a summation formula. Alternatively, it can be achieved using a corresponding AI (Artificial Intelligence) algorithm model: A sample dataset containing water flow velocity vectors and propulsion velocities is constructed, and the sample data is labeled with data result labels. These labels represent the true swimming speed corresponding to the water flow velocity vectors and propulsion velocities. Then, based on a learning algorithm, the AI ​​algorithm model is trained using the aforementioned sample dataset. During training, the convolutional neural network model can be retrained or fine-tuned using existing techniques to improve its generalization ability. Finally, the trained convolutional neural network model is obtained. In practical applications, the water flow velocity vectors and propulsion velocities are input into the convolutional neural network model, which analyzes and processes the data, then outputs the relevant data results (i.e., the corresponding true swimming speed). Through the processing of the above algorithm model, the convolutional neural network model can logically deduce and output the actual swimming speed from the input water flow velocity vector and propulsion speed. It should be noted that the above training method is merely an example; those skilled in the art can choose other suitable methods according to the scenario, such as reinforcement learning, federated learning, or transfer learning, and other general learning paradigms. No specific limitations are made in this embodiment of the invention. Furthermore, other data calculated using formulas can also be implemented using this AI algorithm model, and no specific limitations are made here.

[0068] In establishing the attitude-fluid coupling model, this invention establishes a mathematical model based on the extracted tail fin amplitude, body axis curvature, and the acquired environmental velocity vector field. This mathematical model comprehensively considers the swimming posture, power output, and water flow of the large yellow croaker, overcomes the influence of the complex underwater environment, and achieves accurate calculation of the actual swimming speed of the large yellow croaker.

[0069] Furthermore, the swimming speed of the large yellow croaker can be calculated by measuring the apparent velocity of the fish relative to the camera: for large yellow croakers with the same tracking ID in consecutive frames (0.02-second time intervals), the three-dimensional coordinates of the key points at the front of their heads are extracted, the displacement between the two coordinates is calculated, and the spatial distance between the two displacements is calculated. Dividing the calculated spatial distance by the time interval yields the apparent velocity of the large yellow croaker, with the direction being the displacement vector direction. By superimposing the apparent velocity with the water flow velocity vector in X, Y, and Z dimensions, the three-dimensional vector of the true velocity is obtained. Calculating the spatial distance between these three-dimensional vectors yields the true swimming speed of the large yellow croaker. This invention decouples fish movement from the water flow environment, enabling precise analysis of how fish adapt to the flow field and providing quantitative evidence for ecological research. By eliminating water flow interference through vector superposition and analyzing the fish propulsion mechanism separately, the accuracy of motion analysis is improved. In aquaculture, water flow design can be optimized to reduce the energy consumption of fish swimming. The interpolated flow velocity data fills the gaps in the sparse distribution of sensors, adapting to the modeling needs under complex flow fields.

[0070] In one embodiment, the method used in this invention is compared with the traditional closed-lane swimming test method (artificial water flow environment), and the comparison results are shown in the table below: Table 3 Comparison of Optimization Effects

[0071] As shown in Table 3, compared with the traditional method, the method adopted in this invention improves the detection of mAP@0.5 index by 8.6% and the measurement error by 65%. It can be seen that the present invention is significantly better than the traditional closed lane test method in terms of both mAP@0.5 index detection and measurement error, and is more suitable for monitoring the swimming speed of large yellow croaker in complex underwater environments.

[0072] In addition, to address the need for dynamic adaptability of the improved YOLOv7 network model in underwater large yellow croaker monitoring scenarios (such as model performance degradation caused by environmental changes and fish growth), this invention also proposes a multimodal, dual-factor driven progressive adaptive training method. This method breaks through the limitations of traditional static training, adopts a three-stage training approach: basic capability foundation, dynamic scenario adaptation, and online performance iteration. It also combines a dual-factor triggering mechanism based on environmental and fish characteristics to achieve continuous optimization of the model in complex underwater scenarios.

[0073] 1. Multimodal foundational capability training (offline): This is used to build the model's ability to recognize the core features (morphology, posture) of the large yellow croaker and basic underwater scenes, laying the foundation for generalization. The specific training process includes: First, a multimodal hybrid dataset is constructed by integrating three types of data: underwater images captured by a wide-angle camera (including scenes with different lighting and turbidity), three-dimensional coordinate data synchronously recorded by a laser rangefinder (used to generate physical scale labels for fish), and manually labeled "fish bounding boxes + key points + environmental labels (such as light intensity and turbidity level)". At the same time, these datasets need to cover different growth stages of large yellow croaker (juvenile to adult), typical postures (straight swimming, turning, tail wagging), and five typical underwater environments (clear, slightly turbid, bubble interference, net cage shadow, and sudden changes in lighting).

[0074] Next, using the laser ranging data in the dataset, a "fish posture-convolution offset" mapping pair is generated (such as the offset direction of the convolution kernel corresponding to the curved fish posture, etc. The mapping pair can be transformed into changes in spatial coordinates through changes in fish posture, and this change and the corresponding convolution features are input into the deep learning model, allowing the network to learn the mapping between this coordinate change and the convolution feature offset). Then, through supervised learning, the multi-scale deformable convolution in the model learns the morphological change law of the underwater fish in advance, and then the parameters of this layer are frozen to participate in the subsequent overall training, accelerating the model convergence, so as to realize the separate pre-training of the multi-scale deformable convolution layer in the backbone network.

[0075] Then, in the training of the neck network, an environment-aware loss is introduced. This loss is obtained by applying a weighted channel attention loss (a penalty term that enhances the weights of fish body feature channels) to samples with different environment labels. This forces the model to learn to suppress interfering features of specific environments (such as the high-brightness channel of bubbles) during the training phase, thereby achieving joint optimization of the dual attention mechanism. The environment-aware loss is expressed by the following formula:

[0076] In the formula, L env Loss of environmental perception; e k For environmental labels (if the sample belongs to the first category) k Class environment e k =1); K This represents the total number of environmental labels. C Total number of channels; M [ k,c [For the model, the first] k The level of attention given to interfering channels in similar environments; w c For the first c The weights of the channels. The larger this loss value, the more the model focuses on interfering channels, and the weights of interfering channels need to be reduced through backpropagation.

[0077] 2. Dynamic Scene Adaptation Training (Semi-Offline): This method allows the model to adapt to personalized scenes in specific monitoring areas (such as the structure of specific aquaculture cages, and the optical characteristics of water bodies caused by local water temperature), reducing performance fluctuations in the initial deployment phase. The specific training process includes: First, the model obtained from the foundation training is used as the pre-training weights, the first 3 layers of the backbone network are frozen (to retain the ability to extract general features), and only the neck and head networks are fine-tuned.

[0078] Next, a small number of samples (approximately 500 images) of the target monitoring area are collected. "Region-specific distortion samples" (simulating the optical refraction patterns of the region) are generated using laser ranging data. The sample diversity is expanded through a GAN network (e.g., generating region-specific "net cage shadow + fish body" hybrid samples). Specifically, a large number of net cage shadow and fish body images are collected, preprocessed, and labeled. Then, during CGAN training, the features of the net cage shadow and fish body images are input as conditional information along with a noise vector into the generator. The generator attempts to generate images that fuse net cage shadows and fish bodies, while the discriminator distinguishes between the generated hybrid images and the real hybrid images. Through adversarial training between the generator and the discriminator, the generator parameters are continuously optimized to enable the generation of realistic "net cage shadow + fish body" hybrid samples. For the generation of other hybrid samples, this invention only sets corresponding conditions and adjusts the model parameters according to the application scenario and requirements.

[0079] Finally, prior knowledge from the attitude-fluid coupling model is introduced: a "keypoint physical consistency constraint" is added to the loss function. When the deviation between the model's predicted tail fin amplitude and body axis curvature and the physical scale calculated by laser ranging exceeds a threshold, the keypoint loss weight is increased linearly or nonlinearly to ensure that the attitude characteristics output by the model conform to biomechanical laws. For example, a basic keypoint loss weight w0 is set, and when the physical scale deviation δ exceeds a threshold T, it is increased linearly. , where k is an amplification factor, the appropriate value of which can be determined experimentally, generally not exceeding 2, in order to better adapt to the influence of physical scale deviation on the loss weight of key points.

[0080] 3. Online incremental iterative training (real-time): This is used to learn about new scene changes (such as seasonal water quality changes and morphological changes in fish after growth) in real time after model deployment, avoiding performance degradation. The specific training process includes: First, the model output of "high confidence detection results + synchronous multimodal data" is stored in real time (such as automatically retaining samples with detection confidence > 0.85, along with laser ranging data and water flow sensor data) to form a dynamic sample pool (the window size is set to 5000 images, and the earliest sample is removed after exceeding the limit).

[0081] Next, incremental training is initiated every 1000 new samples: the currently deployed model is used as the teacher model, and the newly trained temporary model is used as the student model. When the student model learns new samples, it aligns with the teacher model's predicted distribution on historical samples through knowledge distillation loss (KL divergence) to avoid catastrophic forgetting (such as forgetting previously learned juvenile fish features). For new interferences appearing in the new samples (such as green water caused by increased algae in summer), the channel attention weight learning strategy is automatically adjusted (to enhance the channels that distinguish fish from the green background). For example, when new interferences appear, features are first extracted from the new samples, and then global features are extracted through operations such as global average pooling and global max pooling. Then, 1D convolution operations are used to enhance the interaction between different channels to capture the feature differences between fish and the background under new interferences. Finally, the sigmoid activation function is used to generate adaptive channel weights, and the channel weights are automatically adjusted according to the characteristics of the new interferences, so that the model can pay more attention to the channels that play an important role in distinguishing fish from new interferences (such as green water).

[0082] Simultaneously, the triggering of this training method requires combining performance degradation indicators with environmental and fish-related factors to achieve a combination of passive correction and active prevention: The training process is triggered if any of the following indicators exceed the threshold for 500 consecutive frames: detection accuracy is below 85% (normal state ≥ 90%); physical coordinate error of the head tip / tail fin tip > 5cm (normal state ≤ 3cm); the ID of the same fish body changes more than 3 times within 10 seconds (normal state ≤ 1 time). This automatically initiates online incremental training, calling the latest 1000 samples from the sliding window sample pool to fine-tune the head and neck networks. (Training rounds = 5 to avoid overfitting), seamlessly replace the original model after training is completed; while monitoring the water turbidity (NTU value) changes by more than 50% compared to the initial deployment, or the light intensity fluctuates by more than 30% (such as a sudden change from cloudy to sunny on a rainy day), and through laser ranging data statistics, the average body length of the fish increases by more than 20% compared to the initial deployment (growth leads to morphological changes), start dynamic scene adaptation training in advance, supplement and collect 200 new samples (specifically covering the current environment and fish status), and combine them with historical core samples (such as retaining 30% of the initial samples) for fine-tuning to prevent performance degradation.

[0083] This invention incorporates physical scale information from laser ranging into the training process to resolve pixel physical mapping deviations caused by underwater image distortion, making the model's output posture features more consistent with actual biomechanical laws. It breaks through the traditional passive mode of retraining after performance degradation, proactively adapting to dynamic changes through early warnings of the environment and fish growth, reducing monitoring interruptions. It retains historical knowledge during online training, preventing the model from forgetting previously learned features and achieving long-term stable performance iteration. This allows the improved YOLOv7 model to quickly adapt to specific scenarios in the initial deployment phase and continuously adapt to environmental and fish changes during long-term use, providing stable and reliable target detection and posture feature extraction capabilities for monitoring the swimming speed of large yellow croaker.

[0084] It should be noted that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.

[0085] In another embodiment, such as Figure 2 As shown, a second aspect of the present invention provides a large yellow croaker swimming speed monitoring system based on target detection, comprising: The data acquisition module 11 is used to acquire the image to be monitored, real-time environmental flow velocity data, and real-time water turbidity data of the target water area, respectively; the target detection module 12 is used to perform target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker; the linkage verification module 13 is used to perform linkage verification on the coordinate data to obtain verification coordinate data; the linkage verification is selected by the camera based on the analysis results of the coordinate data, and includes at least: a first-level verification based on biomechanical constraints to verify the motion state reflected by the coordinate data; and / or a second-level verification based on the flow velocity vector confirmed by the real-time water turbidity data; and / or a third-level verification based on the motion law reflected by the coordinate data and the biomechanical constraints; the attitude optimization module 14 is used to determine the real-time attitude characteristics of the large yellow croaker based on the verification coordinate data, and optimize the real-time attitude characteristics through a flow velocity phase joint mechanism to obtain optimized attitude characteristics; the velocity quantization module 15 is used to construct an attitude fluid coupling model based on the optimized attitude characteristics and the real-time environmental flow velocity data, and determine the swimming speed of the large yellow croaker through the attitude fluid coupling model.

[0086] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.

Claims

1. A method for monitoring the swimming speed of large yellow croaker based on target detection, characterized in that, include: The system acquires images of the target water area to be monitored, real-time environmental flow velocity data, and real-time water turbidity data, respectively. Target detection is performed on the image to be monitored to obtain the coordinate data of the large yellow croaker; The coordinate data is then subjected to linkage verification to obtain verified coordinate data; The linkage verification is selected by the camera based on the analysis results of the coordinate data, and it includes at least: a first-level verification based on biomechanical constraints to verify the motion state reflected by the coordinate data; And / or, a secondary verification based on the flow velocity vector confirmed by the real-time water turbidity data; and / or, a tertiary verification based on the motion patterns reflected by the coordinate data and the biomechanical constraints. The real-time attitude features of the large yellow croaker are determined based on the verified coordinate data, and the real-time attitude features are optimized through a flow velocity-phase joint mechanism to obtain optimized attitude features. An attitude-fluid coupling model is constructed based on the optimized attitude characteristics and the real-time environmental flow velocity data, and the swimming speed of the large yellow croaker is determined through the attitude-fluid coupling model.

2. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 1, characterized in that, The step of performing linkage verification on the coordinate data to obtain verification coordinate data includes: The coordinate data is transformed to obtain three-dimensional coordinate data, and the body length of the large yellow croaker is determined based on the three-dimensional coordinate data; A biomechanical limit database of the large yellow croaker at multiple growth stages is constructed, and the biomechanical limit database is matched with the body length to achieve a first-level verification of the biomechanical constraints of the large yellow croaker. The three-dimensional coordinate data with a successful matching result is used as the verification coordinate data. When the matching result is determined to be a failure, a secondary compensation mechanism is triggered. The coordinate data is input into a pre-constructed grid velocity mapping table for spatial matching to obtain a velocity vector. This vector is then combined with the coordinate data of a preset frame to predict the theoretical coordinates of the current frame. Based on the real-time water turbidity data, residual correction and coordinate correction are performed on the theoretical coordinates to obtain corrected coordinates, which are then used as the verification coordinate data. When the corrected coordinates are detected to be invalid, a three-level completion mechanism is triggered to obtain the three-dimensional valid coordinate data of the large yellow croaker and fit it with a third-order Bézier curve to obtain the individual trajectory equation for extrapolation prediction, generating predicted coordinates. The predicted coordinates are then subjected to the first-level verification of the biomechanical constraints. If the match is successful, the predicted coordinates are used as the verification coordinate data of the large yellow croaker; otherwise, the control point weights of the Bézier curve are adjusted, and the prediction is repeated until the constraints are met.

3. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 2, characterized in that, The process of transforming the coordinate data to obtain three-dimensional coordinate data includes: The Z coordinate of the large yellow croaker is determined by the distance from the key point of the large yellow croaker to the narrow-beam laser rangefinder; Based on the Z coordinate, the coordinate data is distorted by using the calibration parameters of the underwater wide-angle camera to obtain the X and Y coordinates of the large yellow croaker, and then combined with the Z coordinate to generate the three-dimensional coordinate data of the large yellow croaker.

4. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 1, characterized in that, The verification coordinate data includes the three-dimensional coordinates of the front end of the head, the three-dimensional coordinates of the midpoint of the body axis, the three-dimensional coordinates of the base of the caudal fin, and the three-dimensional coordinates of the tip of the caudal fin; wherein... Determining the real-time pose characteristics of the large yellow croaker based on the verified coordinate data includes: The relative displacement of the three-dimensional coordinates of the tip of the caudal fin with respect to the three-dimensional coordinates of the root of the caudal fin in a single frame is calculated, so as to perform sliding window analysis on the relative displacement in consecutive frames, and the maximum value in the window is taken as the caudal fin sway amplitude for that consecutive frame period. Based on the three-dimensional coordinates of the front end of the head, the three-dimensional coordinates of the midpoint of the body axis, and the three-dimensional coordinates of the base of the tail fin, the body axis curve of the large yellow croaker is constructed, and the body axis curvature is determined based on the body axis curve, which, together with the tail fin swing amplitude, serves as the real-time posture feature.

5. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 4, characterized in that, The optimization of the real-time attitude features through the combined flow velocity and phase mechanism to obtain optimized attitude features includes: A fingerprint mapping table of tail wagging cycle and swimming state of various yellow croakers in the target water area is pre-constructed; the tail wagging cycle is determined based on the three-dimensional coordinates of the tail fin tip, and the swimming state is determined based on the body axis curvature. The swimming state of the large yellow croaker in the current frame is determined to match it with the fingerprint mapping table, and the current tail wagging period is obtained to dynamically adjust the frame number of the sliding window to obtain the target sliding window. Phase features within the target sliding window are extracted for phase segmentation to obtain segmentation results. Weights are dynamically assigned to the segmentation results according to the phase type for weighted calculation to obtain weighted swing amplitude. Based on the body axis curvature and the pre-built swing curvature correlation model, the theoretical swing of the current frame is determined, and the deviation between the weighted swing and the theoretical swing is judged to realize the attitude linkage verification of the real-time attitude features. When the condition is determined to be normal, the weighted swing amplitude and the body axis curvature are used as the optimized posture features; when the condition is determined to be abnormal, the body axis curvature is smoothed and corrected, and the correction result and the theoretical swing amplitude are used as the optimized posture features.

6. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 1, characterized in that, Target detection in the image to be monitored is achieved using an improved YOLOv7 network model. This improved YOLOv7 network model includes an input layer, a backbone feature extraction network layer, a neck feature fusion network layer, and a head detection output layer. The backbone feature extraction network layer includes a first ELAN block, a second ELAN block, a third ELAN block, and a fourth ELAN block. The second, third, and fourth ELAN blocks all employ multi-scale deformable convolutions. The step of performing target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker includes: The input layer converts the image to be monitored into a normalized tensor. Based on the first ELAN block, the normalized tensor is subjected to initial feature extraction to obtain the first feature used to characterize the edge of the fish body. The second ELAN block is used to perform secondary feature extraction on the first feature to obtain a second feature that characterizes the fish body outline. Based on the third ELAN block, the second feature is extracted three times to obtain the third feature used to characterize the fish body morphology; The third feature is extracted four times using the fourth ELAN block to obtain the fourth feature used to characterize the distant fish body.

7. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 6, characterized in that, The neck feature fusion network layer includes a PAN block, a channel attention layer, and a spatial attention layer; the PAN block includes an upsampling path and a downsampling path; wherein... The step of performing target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker also includes: In the upsampling path, the fourth feature is upsampled and then concatenated with the third feature to obtain the fifth feature, and the fifth feature is upsampled and then concatenated with the second feature to obtain the sixth feature; In the downsampling path, the sixth feature is downsampled and then concatenated with the fifth feature to obtain the seventh feature, and the seventh feature is downsampled and then concatenated with the fourth feature to obtain the eighth feature; The sixth, seventh, and eighth features are all input into the channel attention layer and the spatial attention layer to perform weighted processing on the corresponding features in the channel dimension and the spatial dimension, respectively, to obtain the first enhanced feature, the second enhanced feature, and the third enhanced feature.

8. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 7, characterized in that, The head detection output layer includes a first parallel detection branch, a second parallel detection branch, and a third parallel detection branch; the first parallel detection branch includes a first detection sub-branch and a first keypoint prediction sub-branch. The second parallel detection branch includes a second detection sub-branch and a second keypoint prediction sub-branch; The third parallel detection branch includes a third detection sub-branch and a third keypoint prediction sub-branch; wherein... The step of performing target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker also includes: The first enhanced feature is input into the first detection sub-branch and the first key point prediction sub-branch respectively for processing to obtain the first bounding box coordinate data and the first key point coordinate data. The second enhanced feature is input into the second detection sub-branch and the second key point prediction sub-branch respectively for processing to obtain the second bounding box coordinate data and the second key point coordinate data. The third enhanced feature is input into the third detection sub-branch and the third key point prediction sub-branch respectively for processing to obtain the third bounding box coordinate data and the third key point coordinate data; The coordinate data of the first bounding box, the second bounding box, the third bounding box, the first keypoint, the second keypoint, and the third keypoint are combined to generate the coordinate data of the large yellow croaker.

9. The method for monitoring the swimming speed of large yellow croaker based on target detection according to claim 1, characterized in that, The process of constructing an attitude-fluid coupling model based on the optimized attitude features and the real-time environmental flow velocity data, and determining the swimming speed of the large yellow croaker using the attitude-fluid coupling model, includes: The real-time environmental flow velocity data is interpolated to obtain the water flow velocity vector of the target water area; Based on the optimized posture characteristics, the propulsion speed of the large yellow croaker relative to the water flow is determined; The attitude fluid coupling model is constructed based on the water flow velocity vector and the propulsion velocity, and the swimming speed of the large yellow croaker is determined based on the attitude fluid coupling model.

10. A swimming speed monitoring system for large yellow croaker based on target detection, characterized in that, include: The data acquisition module is used to acquire the image of the target water area to be monitored, as well as the real-time environmental flow velocity data and real-time water turbidity data of the target water area; The target detection module is used to perform target detection on the image to be monitored and obtain the coordinate data of the large yellow croaker; The linkage verification module is used to perform linkage verification on the coordinate data to obtain verification coordinate data. The linkage verification is selected by the camera based on the analysis results of the coordinate data, and includes at least: a first-level verification based on biomechanical constraints to verify the motion state reflected by the coordinate data; and / or a second-level verification based on the flow velocity vector confirmed by the real-time water turbidity data; and / or a third-level verification based on the motion law reflected by the coordinate data and the biomechanical constraints. The attitude optimization module is used to determine the real-time attitude characteristics of the large yellow croaker based on the verification coordinate data, and to optimize the real-time attitude characteristics through a flow velocity phase joint mechanism to obtain optimized attitude characteristics. The velocity quantization module is used to construct an attitude-fluid coupling model based on the optimized attitude features and the real-time environmental flow velocity data, and to determine the swimming speed of the large yellow croaker through the attitude-fluid coupling model.

Citation Information

Patent Citations

  • Deep learning-based fish school toxicity behavior analysis method and system

    CN120148110A

  • Pelteobagrus vachelli population dynamic monitoring system and method

    CN120391363A

  • Real-time attitude estimation method of underwater inertial navigation system

    CN120863846A