A bighead carp swimming speed monitoring method and system based on target detection

By using a dual-sensor visual acquisition system and an improved YOLOv7 network model, combined with biomechanical and water turbidity data verification, the accuracy problem of monitoring the swimming speed of large yellow croaker in complex underwater environments was solved, and high-precision swimming speed calculation was achieved.

CN121053705BActive Publication Date: 2026-02-06GUANGDONG OCEAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511587595.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-02-06
Estimated Expiration
2045-11-03

AI Technical Summary

Technical Problem

Existing technologies cannot accurately monitor the swimming speed of large yellow croaker in complex underwater environments. The monitoring accuracy is low due to factors such as optical distortion and water turbidity.

Method used

A dual-sensor visual acquisition system using a narrow-beam laser rangefinder and an underwater wide-angle camera, combined with an improved YOLOv7 network model, is used for target detection. Through biomechanical constraints and real-time water turbidity data verification, an attitude-fluid coupling model is constructed to determine swimming speed.

Benefits of technology

This technology enables rapid and accurate acquisition of large yellow croaker coordinate data in complex underwater environments, reducing attitude noise, improving monitoring accuracy, and scientifically calculating swimming speed, thus providing support for research on ecological habits and behavioral patterns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053705B_ABST
    Figure CN121053705B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of aquatic product monitoring, and discloses a large yellow croaker swimming speed monitoring method and system based on target detection, which comprises the following steps: acquiring a to-be-monitored image of a target water area, real-time environmental flow rate data and real-time water body turbidity data of the target water area; performing target detection on the to-be-monitored image to obtain coordinate data of the large yellow croaker; performing linkage verification on the coordinate data to obtain verified coordinate data; determining real-time posture features of the large yellow croaker according to the verified coordinate data; optimizing the real-time posture features through a flow rate phase joint mechanism to obtain optimized posture features; and constructing a posture fluid coupling model according to the optimized posture features and the real-time environmental flow rate data to determine the swimming speed of the large yellow croaker. The application fully considers the influence of the posture of the large yellow croaker and the environment, and combines the biomechanical features to realize accurate calculation of the real swimming speed of the large yellow croaker.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of aquatic product monitoring, in particular to a large yellow croaker swimming speed monitoring method and system based on target detection. BACKGROUND

[0002] In the research and breeding process of large yellow croaker, understanding its swimming speed is of great significance for grasping the physiological state, behavior habit and optimizing the breeding environment of large yellow croaker. At present, the monitoring methods of fish swimming speed mainly include marking recapture method, sonar detection method, video image analysis method, etc.; among them, the marking recapture method is tedious to operate, has certain harm to fish, and has low monitoring precision; the sonar detection method is greatly disturbed by underwater environment, and the monitoring effect of small fish is not good; although the video image analysis method is relatively simple to operate, in the complex underwater environment, due to the influence of optical distortion, turbidity of water body and other factors, it is difficult to accurately obtain the position and posture information of fish body, resulting in low swimming speed monitoring precision.

[0003] Therefore, how to solve the problem that the prior art cannot overcome the influence of complex underwater environment and accurately monitor the swimming speed of large yellow croaker has become a technical problem to be solved by the technical personnel in the field. SUMMARY

[0004] The present application provides a large yellow croaker swimming speed monitoring method and system based on target detection, which solves the problem of how to overcome the influence of complex underwater environment and accurately monitor the swimming speed of large yellow croaker.

[0005] To solve the above technical problems, the present application provides a large yellow croaker swimming speed monitoring method based on target detection, comprising:

[0006] Respectively acquiring the to-be-monitored image of the target water area, real-time environmental flow rate data and real-time water body turbidity data;

[0007] Performing target detection on the to-be-monitored image to obtain coordinate data of large yellow croaker;

[0008] Performing linkage verification on the coordinate data to obtain verified coordinate data; the linkage verification is selected according to the analysis result of the coordinate data, which at least includes: a first-level verification of verifying the motion state reflected by the coordinate data based on biomechanical constraints; and / or, a second-level verification of correcting the flow rate vector confirmed based on the real-time water body turbidity data; and / or, a third-level verification of correcting based on the motion law reflected by the coordinate data and the biomechanical constraints;

[0009] Determining the real-time posture feature of the large yellow croaker according to the verified coordinate data, and optimizing the real-time posture feature through a flow rate phase joint mechanism to obtain an optimized posture feature;

[0010] construct a posture fluid coupling model based on the optimized posture feature and the real-time environmental flow rate data, and determine the swimming speed of the large yellow croaker through the posture fluid coupling model.

[0011] Compared with the prior art, the embodiment of the present application has the following beneficial effects:

[0012] The target detection on the to-be-monitored image can quickly and accurately obtain the coordinate data of the large yellow croaker; at the same time, through linkage verification of the coordinate data, specifically, three-level biological force school verification of linkage of body length and movement limit, the accuracy of the coordinate data is further improved, a reliable foundation is provided for subsequent analysis, and deviation of subsequent analysis caused by coordinate error is reduced; the real-time posture feature is optimized through the flow rate phase joint mechanism, which can more accurately reflect the real posture of the large yellow croaker, remove possible posture data noise or unreasonable fluctuations, make the posture feature more stable and reliable, and help to more accurately understand the behavior state of the large yellow croaker; the optimized posture feature is combined with the environmental flow rate data of the target water area to construct a posture fluid coupling model to determine the swimming speed of the large yellow croaker, which fully considers the influence of the posture of the large yellow croaker itself and the surrounding water flow environment on the swimming speed, and can more scientifically and accurately calculate the actual swimming speed of the large yellow croaker in a complex water flow environment, thereby providing strong support for research on the ecological habit, behavior mode and interaction with the environment of the large yellow croaker. BRIEF DESCRIPTION OF DRAWINGS

[0013] In order to more clearly illustrate the technical solutions of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings described in the following only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0014] Figure 1 is a flow chart of a large yellow croaker swimming speed monitoring method based on target detection provided by an embodiment of the present application;

[0015] Figure 2 is a structural diagram of a large yellow croaker swimming speed monitoring system based on target detection provided by an embodiment of the present application;

[0016] Reference signs:

[0017] Among them, 11, data acquisition module; 12, target detection module; 13, linkage verification module; 14, posture optimization module; 15, speed quantification module. DETAILED DESCRIPTION

[0018] With reference to the drawings and embodiments, the technical solutions in the embodiments of the present application are clearly and completely described. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. The purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0019] In the description of the present application, the terms "first", "second", "third" and the like are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second", "third" and the like can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more. The term "and / or" used herein includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0020] In the description of the present application, it should be noted that, unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as generally understood by those skilled in the art. The terms used in the specification of the present application are only used to describe specific embodiments, and are not intended to limit the present application. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0021] In an embodiment, as shown in Figure 1 The first aspect of the present application provides a bighead carp swimming speed monitoring method based on target detection, comprising:

[0022] S1, acquiring the to-be-monitored image of the target water area, real-time environmental flow rate data and real-time water body turbidity data respectively;

[0023] Specifically, the present application realizes real-time collection of the to-be-monitored images of the target water area through a dual-sensing visual collection system comprising a narrow-beam laser range finder and an underwater wide-angle camera: 8 4K underwater wide-angle cameras are arranged in the target water area, 4 of which are distributed in a regular tetrahedron to cover the vertical to-be-monitored water area (with a depth interval of 2 m), and the other 4 are arranged along a circular periphery with a radius of 10 m to cover the horizontal area, so as to form a regular tetrahedron + circular periphery heterogeneous camera array, thereby breaking through the limitation of traditional single-dimensional monitoring and realizing three-dimensional full-space coverage; the camera and the narrow-beam laser range finder are integrated in hardware, and the synchronism of data collection of the two is realized through a synchronous triggering device, so that the camera can obtain image information of a plurality of groups of underwater large yellow croakers at different angles in real time, and the laser beam of the narrow-beam laser range finder can cover the area where the large yellow croakers often swim in the target water area to collect the distance from the large yellow croakers in real time.

[0024] In addition, the water flow monitoring adopts three-dimensional grid sampling: in a 10m*10m*5m monitoring space, 25 Doppler current profilers (ADCP) and environmental parameter sensors (temperature, salinity, dissolved oxygen, turbidity meter) are arranged at intervals of 2m*2m*1m, the sampling frequency reaches 10Hz, and the real-time environmental flow rate data and real-time water turbidity data of the target water area are synchronously output, thereby providing mathematical support for subsequent large yellow croaker breeding research. It should be noted that all the devices are equipped with starlight-level night vision sensors (minimum illumination 0.01 lux) and anti-bio-attachment coatings to ensure clear imaging in turbid water; the target water area in the present application can be a breeding pond or a certain sea area, which is not limited here; the above arrangement is commonly used in breeding ponds, and if it is a sea area or other target scene, the device arrangement interval and quantity can be adjusted according to the actual situation. The present application can overcome underwater optical distortion and realize real-time calibration of fish three-dimensional coordinates by fusing the dual-sensing visual collection system of the wide-angle underwater visible light camera and the narrow-beam laser range finder to collect the to-be-monitored images of the target water area.

[0025] S2, target detection is performed on the to-be-monitored images to obtain coordinate data of the large yellow croakers;

[0026] S3, linkage verification is performed on the coordinate data to obtain verified coordinate data; the linkage verification is selected according to the analysis result of the coordinate data, and at least includes: a first-level verification of verifying the motion state reflected by the coordinate data based on biomechanical constraints; and / or, a second-level verification of correcting the flow rate vector confirmed based on the real-time water turbidity data; and / or, a third-level verification of correcting based on the motion law reflected by the coordinate data and the biomechanical constraints;

[0027] Specifically, in the above steps S2 and S3, because in the underwater complex scene, the posture and scale of the large yellow croaker change greatly, there are problems such as turbidity of water body and background interference, the target detection of the to-be-monitored image is realized by improving the YOLOv7 network model, and the improved YOLOv7 network model formed after introducing the multi-scale deformation convolution and the channel-space dual attention mechanism into the framework based on the YOLOv7 model is used for target detection of the to-be-monitored image.

[0028] In an embodiment, the improved YOLOv7 network model comprises an input layer (i.e. the input layer of the YOLOv7 model), a backbone feature extraction network layer, a neck feature fusion network layer and a head detection output layer; the backbone feature extraction network layer comprises a first ELAN block, a second ELAN block, a third ELAN block and a fourth ELAN block; the second ELAN block, the third ELAN block and the fourth ELAN block all adopt multi-scale deformation convolution; wherein the target detection of the to-be-monitored image to obtain the coordinate data of the large yellow croaker comprises:

[0029] the to-be-monitored image is converted into a standardized tensor through the input layer;

[0030] the first ELAN block is used for primary feature extraction of the standardized tensor to obtain first features for representing the edges of the fish body;

[0031] the second ELAN block is used for secondary feature extraction of the first features to obtain second features for representing the contour of the fish body;

[0032] the third ELAN block is used for tertiary feature extraction of the second features to obtain third features for representing the morphology of the fish body;

[0033] the fourth ELAN block is used for quaternary feature extraction of the third features to obtain fourth features for representing the distant fish body.

[0034] Specifically, the to-be-monitored image obtained by the dual-sensing visual acquisition system is taken as input, the input image is an RGB three-channel color image with a fixed size (such as 640*640 pixels), simple preprocessing operations such as normalization and data enhancement are performed on the to-be-monitored image through the input layer of the improved model to accelerate the convergence of the model and improve the stability of the training, and then a 640*640*3 standardized tensor is generated and sent to the backbone feature extraction network layer.

[0035] The backbone feature extraction network layer of the application retains the original ELAN (Efficient Layer Aggregation Network) structure block of YOLOv7, which is used to extract multi-scale features (shallow features capture edges, textures and other details, and deep features capture the overall contour of the fish body, fish shape, and distant fish) from the input image. The three key convolution layers in the backbone network (corresponding to small, medium and large scale feature extraction respectively) are replaced by multi-scale deformation convolution to form. The backbone feature extraction network layer includes 5 stages. After the standardization of the tensor input, the 320x320x32 image is first down-sampled and the channel is adjusted by the basic convolution unit (Conv+BN+SiLU) in the first stage, and the output is 320x320x32 image.

[0036] In the second stage, the 320x320x32 image is input into the first ELAN block, and the first ELAN block is composed of multiple convolution layers and residual connections. These convolution layers mainly use standard convolution operations, which perform weighted summation operations on local regions of the input image to extract preliminary shallow features such as fish edges, and obtain the first feature 320x320x64. This feature contains basic information of the image and lays a foundation for subsequent feature extraction. The residual connection helps to solve the gradient vanishing problem in deep neural networks, making information pass more smoothly in the network.

[0037] In the third stage, the 320x320x64 feature image is input into the second ELAN block, and some standard convolutions in the second ELAN block are replaced by multi-scale deformation convolutions. The multi-scale deformation convolution introduces a learnable offset in the convolution process. For multi-scale deformation convolution, first, a 1x1 convolution operation is used to generate an offset field, and the weight of the 1x1 convolution is W offset (the weight is automatically learned through the training process of the model), and the input feature map is X(p), that is, the 320x320x64 feature image, where p represents the spatial position coordinates on the feature map. The offset field generation formula of the multi-scale deformation convolution is:

[0038] Δp k =W offset *X(p)

[0039] In the formula, Δp k is the offset; W offset is the 1x1 convolution weight; X(p) is the input feature map;

[0040] The offset Δp k at each position p is calculated by the offset field generation formula. The offset Δp kis a two-dimensional vector, which is used to indicate the offset of the sampling position of the convolution kernel on the input feature map; then the shape and position of the convolution kernel are adjusted: the traditional standard convolution kernel is fixed in shape and position, while the multi-scale deformation convolution adjusts the shape and position of the convolution kernel according to the generated offset Δp k The shape and position of the convolution kernel are adaptively adjusted, specifically, the sampling position of the convolution kernel on the input feature map is no longer limited to regular grid points, but is offset within the preset range of the fish target according to the offset, which enables the convolution kernel to better adapt to the medium detail features of the large yellow croaker, such as the fish contour, etc., because the contour of the large yellow croaker may not be regular in shape, and this adaptive adjustment can enable the convolution kernel to more accurately capture the contour information; then the sampling process is performed: for each sampling point on the convolution kernel, the corresponding offset Δp k The new sampling positions are found on the input feature map X(p), and the feature values are obtained from these new sampling positions for subsequent convolution calculation; then the convolution calculation is performed: after the sampling is completed, the feature values obtained by sampling are weighted and summed with the weights of the convolution kernel according to the standard convolution calculation method to obtain the output feature value of the position; finally, the feature extraction is performed: by performing the above convolution calculation on all positions on the entire feature map, the output feature of the second ELAN block, i.e., the second feature 160x160x128, is obtained. Since the multi-scale deformation convolution can adaptively adjust the convolution kernel, the extracted feature has more rich detail information than the first feature, and can more accurately describe part of the features of the large yellow croaker, such as the contour of the fish body and the local texture, etc.

[0041] In the fourth stage, the 160x160x128 feature image is input to the third ELAN block, and the multi-scale deformation convolution is also used in the third ELAN block to replace part of the standard convolution. The calculation principle of the offset is the same as that of the second feature, and as the network level deepens, the third ELAN block can extract high-level semantic features of the overall shape of the fish body, and obtain the third feature 80x80x256. The introduction of the multi-scale deformation convolution enables the convolution kernel to adapt to the feature changes of the large yellow croaker at different abstraction levels, further enhancing the extraction ability of the features of the large yellow croaker, and the third feature contains more semantic information about the overall shape and category of the large yellow croaker, which is of great significance for the recognition and classification of the large yellow croaker.

[0042] In the fifth stage, the 80x80x256 feature map is input to the fourth ELAN block, which also uses multi-scale deformable convolution for feature extraction. The calculation principle of the offset is the same as the second feature. In this level, the model further excavates the deep features of the large yellow croaker, such as the distant fish body, etc. These features usually have stronger representation ability and can better distinguish large yellow croaker from other objects, thereby obtaining the fourth feature 40x40x512. Multi-scale deformable convolution adjusts the convolution kernel adaptively, enabling the model to cope with various changes of the large yellow croaker in the complex underwater scene, such as posture change, size change, etc.

[0043] In the ELAN blocks of stages 3-5, the original 3x3 standard convolution is replaced by multi-scale deformable convolution, so that it dynamically adjusts the sampling position of the convolution kernel through learning. For example, when the large yellow croaker is bent, the convolution kernel automatically shifts to the bending direction, better fitting the fish body shape.

[0044] The hierarchical feature extraction method adopted by the present application comprehensively captures the feature information of the large yellow croaker at different scales and different abstraction levels, which helps the model to more accurately identify the large yellow croaker target. Multi-scale deformable convolution introduces a learnable offset, enabling the convolution kernel to flexibly adapt to different shapes and postures of the large yellow croaker, thereby extracting more discriminative features. This helps the model to accurately distinguish the large yellow croaker from other objects in a complex underwater environment, such as turbid water, background interference, etc., and improves the robustness of target detection.

[0045] In an embodiment, the neck feature fusion network layer comprises a PAN block, a channel attention layer and a spatial attention layer; the PAN block comprises an up-sampling path and a down-sampling path; wherein,

[0046] The target detection on the to-be-monitored image to obtain the coordinate data of the large yellow croaker further comprises:

[0047] In the up-sampling path, the fourth feature is up-sampled and spliced with the third feature to obtain a fifth feature, and the fifth feature is up-sampled and spliced with the second feature to obtain a sixth feature;

[0048] In the down-sampling path, the sixth feature is down-sampled and spliced with the fifth feature to obtain a seventh feature, and the seventh feature is down-sampled and spliced with the fourth feature to obtain an eighth feature;

[0049] The sixth feature, the seventh feature and the eighth feature are input into the channel attention layer and the spatial attention layer to perform weighted processing on the corresponding features in the channel dimension and the spatial dimension, respectively, to obtain a first enhanced feature, a second enhanced feature and a third enhanced feature.

[0050] Specifically, the neck feature fusion network layer in the present application inherits the PAN (Path Aggregation Network) structure of YOLOv7, and the PAN block realizes the fusion of multi-scale features through bidirectional paths of top-down (high-level features pass semantic to low-level) + bottom-up (low-level features supplement details to high-level); and the improved YOLOv7 network model in the present application adds a channel-space dual attention mechanism in the PAN path to solve the problem that the fish body features are easily submerged by strong underwater background interference (such as bubbles), and embeds a channel attention layer and a spatial attention layer after the original PAN block, so that the model pays more attention to the features that contribute to large yellow croaker detection when fusing features at different levels, improves the effect of feature fusion, and provides better feature representation for subsequent detection output:

[0051] In the up-sampling path, the fourth feature 40x40x512 is input, which is up-sampled by 2 times to enlarge the 40x40 feature map to 80x80, and is concatenated with the output third feature of the backbone stage 4, and is compressed to 256 channels through convolution to enhance the semantic information of small targets, to obtain the fifth feature 80x80x256; then the fifth feature is enlarged to 160x160 through 2 times up-sampling, and is concatenated with the output second feature 160x160x128 of the backbone stage 3, and is compressed to 128 channels through convolution to fuse detailed features, to obtain the sixth feature 160x160x128.

[0052] In the down-sampling path (bottom-up), the sixth feature 160x160x128 is input, which is down-sampled by 2 times to change the 160x160 feature map to 80x80, and is concatenated with the fifth feature 80x80x256 in the intermediate result of the up-sampling path, and is compressed to 256 channels through convolution to supplement the middle layer positioning information, to obtain the seventh feature 80x80x256; then the seventh feature is changed to 40x40 through 2 times down-sampling, and is concatenated with the output fourth feature 40x40x512 of the backbone stage 5, and is compressed to 512 channels through convolution to optimize large target detection, to obtain the eighth feature 40x40x512.

[0053] The sixth feature is input into the channel attention layer and the spatial attention layer to perform weighted processing on the sixth feature in the channel dimension and the spatial dimension, respectively, to obtain the first enhanced feature; the present application takes the sixth feature as an example to illustrate the processing process of the channel attention layer and the spatial attention layer on the feature:

[0054] In the channel attention layer, the sixth feature 160x160x128 is processed by a global average pooling operation, that is, the values of all spatial positions of each channel are averaged to obtain a channel statistic vector, which contains global statistical information of each channel, and the feature map is compressed to 1x1x128, and then the channel statistic vector is input into a fully connected layer for dimension reduction and dimension increase, and a 1x1x128 feature map is output, and then the output of the fully connected layer is normalized by using a sigmoid function to obtain a channel weight vector 1x1x128, and finally the channel weight vector is multiplied with the eighth feature 160x160x128 in each channel to obtain a ninth feature 160x160x128.

[0055] In the spatial attention layer, the ninth feature 160x160x128 is respectively processed by a global average pooling operation along the width direction and the height direction, which averages the values of all spatial positions of each channel to obtain a channel length vector, which contains global statistical information of each channel; the features in the channel dimension in the two directions are spliced to obtain a feature of 160x160x256; the feature is processed by a 3x3 convolution to obtain a single-channel spatial weight map 160x160x1; then the single-channel spatial weight map is normalized by a sigmoid function to generate a spatial weight map (focus on the region where the fish is located, value range [0, 1]), the value close to 1: key region (such as fish contour), the value close to 0: background region (such as water); finally, the spatial weight map is multiplied with the ninth feature 160x160x128 element by element to obtain a first enhanced feature 160x160x128.

[0056] Similarly, after the seventh feature and the eighth feature are processed by the above-mentioned processing, a second enhanced feature 80x80x256 and a third enhanced feature 40x40x512 can be obtained, each enhanced feature map contains high-level semantics and low-level details, and the background interference is suppressed by the attention mechanism.

[0057] The present application uses the up-sampling and down-sampling paths of the PAN block in the neck feature fusion network layer to capture features of different scales respectively, and realizes the complementation of multi-scale information through splicing operation, which improves the detection capability of small targets; the channel attention mechanism is used to highlight the feature channels important for large yellow croaker recognition, and the spatial attention mechanism is used to focus on the key spatial region where the large yellow croaker is located, so as to reduce the influence of irrelevant information such as turbid water and background interference; the order of processing is first channel and then space, that is, the feature semantics (channel dimension) is optimized first, and then the key region (spatial dimension) is focused, so as to form hierarchical feature enhancement, which significantly improves the detection accuracy of large yellow croaker; through the introduction of the double attention mechanism, the extraction capability of the model for the key features of large yellow croaker is enhanced, and the detection rate and tracking accuracy of the model in the complex underwater scene are further improved.

[0058] In an embodiment, the head detection output layer comprises a first parallel detection branch, a second parallel detection branch and a third parallel detection branch; the first parallel detection branch comprises a first detection sub-branch and a first key point prediction sub-branch; the second parallel detection branch comprises a second detection sub-branch and a second key point prediction sub-branch; the third parallel detection branch comprises a third detection sub-branch and a third key point prediction sub-branch; wherein,

[0059] The target detection on the image to be monitored to obtain the coordinate data of the large yellow croaker further comprises:

[0060] The first enhanced feature is respectively input into the first detection sub-branch and the first key point prediction sub-branch for processing to obtain first bounding box coordinate data and first key point coordinate data;

[0061] The second enhanced feature is respectively input into the second detection sub-branch and the second key point prediction sub-branch for processing to obtain second bounding box coordinate data and second key point coordinate data;

[0062] The third enhanced feature is respectively input into the third detection sub-branch and the third key point prediction sub-branch for processing to obtain third bounding box coordinate data and third key point coordinate data;

[0063] The first bounding box coordinate data, the second bounding box coordinate data, the third bounding box coordinate data, the first key point coordinate data, the second key point coordinate data and the third key point coordinate data are combined to generate the coordinate data of the large yellow croaker.

[0064] Specifically, the head detection output layer in the application is based on the multi-branch detection head of the YOLOv7 network model, and a parallel key point prediction sub-branch is added after each multi-branch detection head of the YOLOv7 model to form, the original head of YOLOv7 contains 3 parallel detection branches (corresponding to 160x160, 80x80 and 40x40 feature maps output by the neck respectively), each branch is composed of a detection sub-branch and a key point prediction sub-branch, and outputs a bounding box, a confidence and a class. And the application improves the YOLOv7 network model to obtain the biomechanical key points (used for calculating the swimming posture) of the large yellow croaker, adds a key point prediction sub-branch after each detection branch to form a multi-task output of “detection + key point”. According to the core biomechanical characteristics reflecting the swimming posture of the large yellow croaker, the front end of the head of the large yellow croaker (the most front end point of the mouth), the midpoint of the body axis (the midpoint of the connecting line between the front end of the head and the root of the tail fin), the root of the tail fin (the starting point of the connection between the tail fin and the trunk) and the tip of the tail fin (the most end point of the tail fin) are selected as key points.

[0065] The detection sub-branch (inherited from YOLOv7) is composed of 3 layers of 3x3 convolution (the middle layer is activated by SiLU, and the output layer is not activated), and the output dimension is: for each feature map grid, the output boundary box coordinate data is (4+1+1) x K, wherein 4 is the boundary box offset (Dx, D y, D w, D h) (based on "anchor box" prediction offset, each anchor box corresponds to a grid on the feature map, with the image top left corner as the origin, the center coordinates: x=(grid top left x+sigmoid(Dx))x grid step; y=(grid top left y+sigmoid(D y))x grid step, sigmoid ensures that the offset is between 0-1 and is limited within the grid; width and height: w=anchor box original w x exp(D w); h=anchor box original h x exp(D h), exp ensures that the width and height are positive); 1 is the confidence (the probability that the grid contains a large yellow croaker, output by the sigmoid function); 1 is the class probability (the class probability is predicted by the softmax function, since only "large yellow croaker" needs to be detected, it is a binary classification: "large yellow croaker" or "background", and the probability of the highest class is taken as the output label); K is the number of anchor boxes for each grid (default 3, suitable for fish bodies of different shapes).

[0066] The key point prediction sub-branch is parallel to the detection sub-branch and is composed of 2 layers of 3x3 convolution (SiLU activation), and the output dimension of this branch is: for each feature map grid, the output key point coordinate data is 2x N x K, wherein: 2 is the x, y coordinate offset of each key point (which can also be regarded as two-dimensional pixel coordinates), which predicts the offset (Dx, D y) of each key point relative to the left top corner (x1, y1) of the boundary box, and the final key point coordinates are u=x1+Dx; v=y1+D y, wherein D x, D y are normalized to 0-1 by sigmoid, and then multiplied by the width and height of the boundary box to ensure that the key point is within the box; N is the number of key points (default 4, i.e. head front end, body axis midpoint, tail fin root, and tail fin tip); K is the number of anchor boxes shared with the detection sub-branch (3).

[0067] Based on the above principle, the first enhanced feature 160x160x128 is input into the first parallel detection branch, that is, the small-scale detection branch (responsible for predicting small anchor boxes and matching long-distance targets), and is processed through the first detection sub-branch and the first key point prediction sub-branch to obtain first boundary box coordinate data and first key point coordinate data. The second enhanced feature 80x80x256 is input into the second parallel detection branch, that is, the medium-scale detection branch (responsible for predicting medium anchor boxes and matching medium-distance targets), and is processed through the second detection sub-branch and the second key point prediction sub-branch to obtain second boundary box coordinate data and second key point coordinate data. The third enhanced feature 40x40x512 is input into the third parallel detection branch, that is, the large-scale detection branch (responsible for predicting large anchor boxes and matching close-range targets), and is processed through the third detection sub-branch and the third key point prediction sub-branch to obtain third boundary box coordinate data and third key point coordinate data. The above boundary box coordinate data and key point coordinate data are subjected to non-maximum suppression processing, that is, for multiple overlapping candidate boxes of the same large yellow croaker, the boundary box with the highest confidence is retained and the repeated boxes are removed, and then the key points corresponding to the boundary box after NMS are retained to ensure that the key points of each large yellow croaker correspond to the boundary box one by one. Finally, the filtered key point coordinates are spliced and arranged, and the processed boundary box coordinates are spliced and arranged, so as to obtain the boundary box coordinate data and the key point coordinate data of the target large yellow croaker. In addition, after the feature is processed in the head detection output layer, the improved model can also output the tracking ID of the large yellow croaker, that is, the same large yellow croaker in the continuous frame is assigned a unique ID, so as to realize cross-frame tracking.

[0068] The present application improves the compatibility of small targets (long-distance fish body) and large targets (close-range fish body) by jointly designing the multi-branch detection head and the key point prediction sub-branch, and processing different scale enhanced features by three parallel detection branches to cover the detection needs of large yellow croaker at different distances. Each detection branch additionally increases the key point prediction sub-branch to directly regress the key point coordinates of the fish head, fish tail and the like, solves the problem that the traditional boundary box detection is not sensitive to the fish posture, and improves the discrimination ability in the overlapping or occlusion scene. The joint output of the boundary box coordinate data and the key point coordinate data not only retains the compatibility of the traditional detection, but also provides fine posture information, which is convenient for subsequent behavior analysis.

[0069] In an embodiment, the improved YOLOv7 network model used in the present application is compared with other existing models, and the comparison results are shown in the following table:

[0070] Table 1 result comparison table

[0071]

[0072] As can be seen from the above table 1, the detection accuracy of the improved YOLOv7 network model adopted in the application is improved by 8.6% compared with YOLOv5s and by 5.4% compared with the original YOLOv7; 63FPS meets the real-time requirement (> 30FPS), although it is slightly lower than the original YOLOv7 (71FPS), but it is significantly ahead of the two-stage model Faster R-CNN (22FPS); the parameter amount is 28.4M, which is reduced by 23% (36.9M→28.4M) compared with the original YOLOv7, which is better than Faster R-CNN (41.5M), and at the high precision of 90.7%mAP, the real-time performance of 63FPS is maintained, the parameter amount is reduced by 23%, and the balance of “high precision without sacrificing speed” is realized.

[0073] In an embodiment, based on the real-time water body turbidity data, a three-level biomechanics verification is performed on the coordinate data in combination with body length and motion limit, to obtain verified coordinate data, comprising:

[0074] Converting the coordinate data to obtain three-dimensional coordinate data, and determining the body length of the large yellow croaker according to the three-dimensional coordinate data;

[0075] A biological motion limit database of the large yellow croaker at multiple growth stages is constructed, and the biological motion limit database is matched with the body length to realize the first-level verification of the biomechanics constraint of the large yellow croaker, and the three-dimensional coordinate data with a matching result of success is taken as the verified coordinate data;

[0076] When it is determined that the matching result is a failure, a second-level compensation mechanism is triggered, the coordinate data is input into a pre-constructed grid flow rate mapping table for spatial matching to obtain a flow rate vector, the theoretical coordinates of the current frame are predicted in combination with the coordinate data of the preset frame, and the theoretical coordinates are corrected based on the real-time water body turbidity data to obtain corrected coordinates as the verified coordinate data;

[0077] When it is detected that the corrected coordinates are invalid, a third-level completion mechanism is triggered, three-dimensional valid coordinate data of the large yellow croaker are obtained and fitted using a three-order Bezier curve to obtain an individual trajectory equation for extrapolation prediction to generate predicted coordinates, the predicted coordinates are subjected to the first-level verification of the biomechanics constraint, if the matching is successful, the predicted coordinates are taken as the verified coordinate data of the large yellow croaker, otherwise, the control point weight of the Bezier curve is adjusted, and the prediction is re-performed until the constraint is met.

[0078] Specifically, the two-dimensional coordinates of the key points output by the model are converted into three-dimensional coordinates based on the calibration parameters (wide-angle camera internal parameters, spatial position relationship between the laser range finder and the camera) of the double-sensing visual acquisition system.

[0079] In the embodiment, the converting the coordinate data comprises:

[0080] The Z coordinate of the large yellow croaker is determined through the distance from the key point of the large yellow croaker to the narrow-beam laser range finder;

[0081] Based on the Z coordinate, the X and Y coordinates of the large yellow croaker are obtained through distortion correction of the coordinate data by the calibration parameters of the underwater wide-angle camera, and the three-dimensional coordinate data of the large yellow croaker is generated in combination with the Z coordinate.

[0082] Specifically, the present application takes the straight-line distance from the narrow-beam laser range finder to the key point of the large yellow croaker as the Z coordinate of the key point; then, based on the Z coordinate, the X and Y coordinates of the key point are obtained through distortion correction of the two-dimensional pixel coordinates, i.e. coordinate data, of the key point by the intrinsic parameters (focal length, principal point coordinates) of the underwater wide-angle camera, and the correction process is represented by the following formula:

[0083] X=(u-cx)×Z / f

[0084] Y=(v-cy)×Z / f

[0085] In the formula, X and Y are respectively the X and Y coordinates of the key point; u and v are respectively the two-dimensional pixel coordinates of the key point; cx and cy are respectively the pixel positions of the camera optical axis in the image, i.e. the principal point coordinates; Z is the Z coordinate of the key point; and f is the focal length of the camera.

[0086] Finally, the combination of the X, Y and Z coordinates of the key point is the three-dimensional coordinate data of the point, i.e. the conversion result. Based on the above principle, the three-dimensional coordinate data of each key point contained in each large yellow croaker can be calculated.

[0087] The present application directly provides the distance (Z coordinate) from the key point to the sensor through the narrow-beam laser range finder, avoids the error accumulation of pure visual depth estimation, and significantly improves the accuracy of three-dimensional reconstruction; the underwater wide-angle camera has a wide field of view but serious distortion, and through the absolute distance constraint of the laser range finder, the distortion can be corrected and the X and Y coordinates can be accurately restored, balancing the field of view range and measurement accuracy; the three-dimensional reconstruction is divided into two steps of laser ranging (Z coordinate) and visual calculation (X and Y coordinates), reducing the calculation complexity and being suitable for real-time processing.

[0088] Then, the height H (pixel value) of the detected large yellow croaker bounding box in each frame output by the improved YOLOv7 model is combined with the calibration parameters (1 pixel corresponds to an actual distance k, such as k=0.1 cm / pixel) of the dual-sensor system to calculate the actual body length L=H×k×1.2 (1.2 is the conversion coefficient of the height of the bounding box offline calibration-actual body length), so as to obtain the body length of the large yellow croaker in each frame.

[0089] As a typical migratory fish, there are physiological limits (restricted by muscle explosive force and body size) for the movement speed of large yellow croaker in vertical direction (Z axis) and horizontal direction (X / Y axis). The present application constructs a biological movement limit database of large yellow croaker in multiple growth stages to calibrate the limit value offline, and then determines whether the current coordinates conform to the biological movement rule in real time, and eliminates obvious outliers. The present application selects 3 growth stages of large yellow croaker (juvenile: body length 10-15 cm; subadult: 20-25 cm; adult: 30-35 cm), 50 tails each, records 30 minutes of swimming data in a controllable water environment (flow rate 0.2 m / s, clear water), extracts the X / Y / Z coordinates of each frame, calculates the displacement of adjacent frames, and then obtains the movement speed (displacement divided by time interval, and the time interval is preferably 0.005 s), and calculates the Vx, Vy, Vz of each growth stage, takes the 99.9% quantile as the physiological limit threshold, and stores it in the system database to form a biological movement limit database; wherein the physiological limit threshold is shown in the following table:

[0090] Table 2 physiological limit threshold table

[0091]

[0092] The biological movement limit database is matched with the body length to perform a primary verification of the biomechanics constraint of large yellow croaker, and the coordinate displacement of large yellow croaker in the current frame and the previous frame is calculated to obtain Vx, Vy, Vz, and if it is within 1.2 times of the limit threshold (20% fluctuation space is reserved), it means that the matching is successful, and it is used as the verification coordinate data.

[0093] If any speed exceeds 1.2 times the corresponding limit threshold, it is determined that the matching fails, the current coordinates are taken as abnormal values, and a secondary compensation mechanism is triggered: a grid-flow mapping table of the target water area is established in advance, that is, the target water area is divided into a three-dimensional grid according to the ADCP (Acoustic Doppler Current Profiler) deployment interval (2m*2m*1m), each grid corresponds to a unique ADCP number, and the real-time flow rate data (Ux: X-direction flow rate, Uy: Y-direction flow rate, Uz: Z-direction flow rate) of the ADCP are stored; then the coordinate data of any key point of large yellow croaker in the current frame, that is, the two-dimensional pixel coordinates (u, v), are mapped into the pre-defined grid-flow mapping table, the ADCP number corresponding to the grid unit is found, the flow rate vector is read, and the historical average value of the fish body propulsion speed Vxavg=(X3-X1) / (2Δt) is calculated based on the fish body coordinates (X1, Y1, Z1), (X2, Y2, Z2), (X3, Y3, Z3) of the previous 3 frames, and the calculation process of Vyavg and Vzavg is the same as that of Vxavg), combined with the current flow rate vector, the theoretical coordinates of the current frame (Xpred=X3+(Vxavg+Ux)×Δt, Ypred, Zpred) are predicted (the calculation process of Ypred and Zpred is the same as that of Xpred); the residual error of the theoretical coordinates is corrected according to the real-time water turbidity data of the target water area: if the Z coordinate Zmeas measured by the laser exists deviation (i.e. Zmeas-Zpred>0.1m), the Z coordinate is corrected by Zcorr=Zpred+α×(Zmeas-Zpred), wherein α is a dynamic compensation coefficient, α=0.8 when the turbidity is less than 200 NTU, the laser measurement is preferred; when 200≤NTU≤300, the linear transition is α=0.8-(0.5*(NTU-200) / 100), and when the turbidity is greater than 300 NTU, α=0.3, the theoretical prediction is preferred; if there is no deviation, no correction is needed; based on the corrected Zcorr, the X / Y coordinates are corrected again through the camera intrinsic parameters (i.e. the distortion correction process during coordinate conversion), to ensure the consistency of the association between the X / Y / Z coordinates and the water flow environment, and the corrected coordinates are generated as the verification coordinate data.

[0094] If the laser range finder has no return signal at all (i.e. Zcorr is NULL), it means that the corrected coordinates are invalid, and a tertiary compensation mechanism is triggered: for large yellow croaker with missing Z coordinates, the three-dimensional coordinate data of the previous 5 frames are extracted as the control points of the Bezier curve, a third-order Bezier curve (4 control points) is used to fit the horizontal direction (X-Y) and vertical direction (Z-t) motion trajectories with the previous 4 frame coordinates (P0, P1, P2, P3), and the individual trajectory equation is obtained:

[0095] X(t)=P0x×(1-t)³+3P1x×(1-t)²t+3P2x×(1-t)t²+P3x×t³

[0096] Y(t)=P0y×(1-t)³+3P1y×(1-t)²t+3P2y×(1-t)t²+P3y×t³

[0097] Z (t)=P0z×(1-t)³+3P1z×(1-t)²t+3P2z×(1-t)t²+P3z×t³

[0098] In the formula, t is a time parameter (t∈[0, 1], corresponding to the time interval of the first 4 frames).

[0099] Suppose that the missing frame is the 6th frame, the time interval Δt=1 / frame rate (for example, Δt=0.02 s when the frame rate is 50 fps), and the t value corresponding to the 6th frame is t5=5×Δt / (4×Δt)=1.25 (based on the time span of the first 4 frames 4Δt, extrapolated to the 5th Δt), the t5 is substituted into the individual trajectory equation to calculate the predicted coordinates of the 6th frame; the predicted coordinates are subjected to the first level of biomechanics constraint verification, if it is within the limit range, it is taken as the coordinate of the missing frame, that is, the verification coordinate data; if it is beyond, it means that the extrapolation is too aggressive, then the control point weight of the Bezier curve is adjusted (simple linear extrapolation or halving the speed), and the prediction is re-performed until the constraint is met.

[0100] The application introduces biomechanics constraints according to the principle that the instantaneous speed of any fish cannot exceed its physiological limit, to quickly eliminate obvious outliers; through fusion of flow data, physical laws are used for constraint to compensate for multi-modal residual errors, and deviations caused by instantaneous errors of a single sensor are corrected; when the laser completely fails, the current position of the fish is predicted based on the historical motion trajectory of the fish; a three-level verification mechanism is adopted to ensure that the system can still provide continuous and reliable data output in the case of temporary failure of any single sensor, greatly improving the usability and reliability of the system.

[0101] S4, determining a real-time posture feature of the large yellow croaker according to the verification coordinate data, and optimizing the real-time posture feature through a flow velocity phase joint mechanism to obtain an optimized posture feature;

[0102] In an embodiment, the verification coordinate data includes a head front end three-dimensional coordinate, a body axis midpoint three-dimensional coordinate, a tail fin root three-dimensional coordinate, and a tail fin tip three-dimensional coordinate of the large yellow croaker; wherein the real-time posture feature of the large yellow croaker is determined according to the verification coordinate data, including:

[0103] The relative displacement of the tail fin tip three-dimensional coordinate relative to the tail fin root three-dimensional coordinate in a single frame is calculated, a sliding window analysis is performed on the relative displacement in consecutive frames, and the maximum value in the window is taken as the tail fin swing amplitude in the consecutive frame period;

[0104] Based on the head front end three-dimensional coordinates, the body axis midpoint three-dimensional coordinates and the tail fin root three-dimensional coordinates, the body axis curve of the large yellow croaker is constructed, and the body axis curvature is determined based on the body axis curve to serve as the real-time posture feature with the tail fin swing.

[0105] Specifically, the present application calculates the tail fin swing, which is the maximum displacement distance of the tail fin tip relative to the tail fin root in the continuous motion process, as the core feature reflecting the swimming power output of the large yellow croaker based on the multiple three-dimensional coordinates after inspection. According to the three-dimensional coordinates of the tail fin tip and the three-dimensional coordinates of the tail fin root, the three-dimensional displacement vector of the three-dimensional coordinates of the tail fin tip relative to the three-dimensional coordinates of the tail fin root at a certain time in a single frame image is calculated, and the spatial distance of the three-dimensional displacement vector, also known as the Euclidean norm or Euclidean distance, is calculated, so that the relative displacement of the large yellow croaker in a single frame is obtained. Then, the displacement size of continuous N frames (such as 50 frames in 1 second) is analyzed by sliding window, and the maximum value in the window is taken as the tail fin swing in the continuous frame period. The window size is usually set to 0.2-0.5 seconds according to the swimming frequency of the large yellow croaker to ensure that a complete tail swing period is covered.

[0106] The body axis curvature, which is the bending curvature of the body axis (the connecting line from the head to the tail fin root), is calculated by three-dimensional coordinates to reflect the bending degree of the body of the large yellow croaker (the more severe the bending, the stronger the swimming power). Since the fish body motion is mainly in the horizontal plane (ignoring the slight fluctuation in the vertical direction), the X-Y plane coordinates are fitted to a quadratic curve (the fish body bending is approximately a plane curve), and the body axis curve equation is set as The head front end three-dimensional coordinates, the body axis midpoint three-dimensional coordinates and the tail fin root three-dimensional coordinates are substituted into the body axis curve equation, so that the equation coefficients are obtained. After the coefficients of the body axis curve equation are known, the body axis curvature of the body axis curve can be calculated according to the curvature formula and any one of the head front end three-dimensional coordinates, the body axis midpoint three-dimensional coordinates and the tail fin root three-dimensional coordinates, and the tail fin swing and the body axis curvature of the same large yellow croaker are taken as the real-time posture feature of the large yellow croaker.

[0107] Based on the calibration parameters of the double-sensing vision system, the present application converts two-dimensional pixel coordinates into three-dimensional world coordinates, eliminates the perspective distortion effect, and accurately calculates the posture features of the key points of the fish body. By analyzing the relative displacement of the tail fin in continuous frames through sliding window, the key motion feature of the tail fin swing is extracted, which reflects the swimming intensity and rhythm of the fish. The three-dimensional coordinate conversion and sliding window statistics effectively suppress the visual noise and instantaneous motion interference, and improve the feature stability.

[0108] In an embodiment, the real-time posture feature is optimized by the flow rate phase joint mechanism to obtain an optimized posture feature, which comprises:

[0109] pre-constructing a mapping table of tail-swing periods of each large yellow croaker in the target water area and swimming states, the tail-swing period being determined according to three-dimensional coordinates of a tail fin tip, and the swimming state being determined according to the body axis curvature;

[0110] determining a swimming state of the large yellow croaker in the current frame to match the mapping table, obtaining a current tail-swing period to dynamically adjust the number of frames of the sliding window, and obtaining a target sliding window;

[0111] extracting phase features in the target sliding window to perform phase division, obtaining a division result, and dynamically assigning weights to the division result according to phase types to perform weighted calculation, and obtaining a weighted amplitude;

[0112] based on the body axis curvature and a pre-constructed amplitude curvature correlation model, determining a theoretical amplitude of the current frame, performing deviation determination on the weighted amplitude and the theoretical amplitude to realize posture linkage verification on the real-time posture features;

[0113] when it is determined to be normal, the weighted amplitude and the body axis curvature are taken as the optimized posture features; when it is determined to be abnormal, the body axis curvature is smoothed and corrected, and the correction result and the theoretical amplitude are taken as the optimized posture features.

[0114] Specifically, tail-swing periods of different large yellow croakers have individual differences, and the same fish changes in periods under different states (calm, foraging, and stress). The present application establishes a mapping table of tail-swing periods and swimming states, matches the tail-swing period of the current fish in real time, dynamically adjusts the sliding window, and optimizes the real-time posture features to avoid missing or over-measuring of the fixed window.

[0115] For large yellow croakers in a target water area, an improved OLOv7 tracking ID is used to mark individuals, 30 minutes of swimming data (including three-dimensional coordinates of the tail fin tip and the tail fin root) of each ID is continuously collected, the displacement sequence of the Z direction (or the Y direction) of the tail fin tip coordinates of each ID is extracted, and Fourier transform is performed thereon to obtain a frequency spectrum, the frequency with the largest amplitude in the frequency spectrum is taken as the main swing frequency ω0, and the tail-swing period is T=1 / ω0; the state is determined by the average body axis curvature (calm: average curvature <0.1 cm -1 ; foraging: 0.1-0.2 cm -1 ; stress: >0.2 cm -1 ), and a mapping table of tail-swing periods and swimming states is established for each ID.

[0116] For the current frame of large yellow croaker, the average body axis curvature Cavg of the previous 10 frames is calculated, the corresponding state in the fingerprint library is matched, and according to the current state, the corresponding tail wagging period Tcurr of the ID is called to dynamically adjust the size of the sliding window, and the target sliding window is obtained: window frame number N = frame rate x Tcurr, and the main wagging frequency ω0 of the current ID is recalculated every 30 seconds, and Tcurr is updated to ensure that the window size is synchronized with the real-time tail wagging period, the flow velocity correction of the dynamic window is realized, and the window missing problem caused by the accelerated tail wagging frequency of large yellow croaker under high flow velocity is solved.

[0117] The tail wagging process of large yellow croaker is divided into a force period (the tail fin swings to one side, the displacement is large, and the contribution to the propulsive force is more) and a return swing period (the tail fin resets, the displacement is small, and the contribution is less), the invention improves the accuracy of the swing amplitude calculation by identifying the phase of the tail wagging, assigning high weight to the force period frame and low weight to the return swing period, improving the accuracy of the swing amplitude calculation: for the N frame data in the target sliding window, the displacement Di of the tail fin tip relative to the tail fin root of each frame is calculated, and the first derivative of the displacement is calculated as a phase feature; when the phase feature is greater than 0 and greater than the displacement average Davg in the window, it is determined to be in the force period; when the phase feature is less than 0 and less than Davg, it is determined to be in the return swing period; when the phase feature is equal to 0, it is determined to be at the boundary between the force period and the return swing period; then the phase type is assigned a weight: the force period frame weight wi=1.2x e^(0.5x(Di / Dmax)), where Dmax is the maximum displacement in the window, the weight range is 1.2~2.4; the return swing period frame weight wi=0.8x e^(-0.3x(Davg / Di)) (Di≠0), the weight range is 0.4~0.8; the phase turning point weight wi=1.0 (neutral weight); then the total weight in the window is calculated, each wi is normalized to obtain the normalized weight, and then the weighted sum of the displacement of each frame is calculated, which highlights the contribution of the force period displacement and reduces the interference of the return swing period, to obtain the weighted swing amplitude, realize the phase weighted flow velocity constraint, and ensure that the weighted swing amplitude can reflect the influence of water flow assistance / resistance on the force efficiency.

[0118] The swimming posture of large yellow croaker has linkage: the larger the tail fin swing amplitude, the larger the body axis curvature (the trunk bends more obviously when force is applied), and the two are positively correlated, and the invention can eliminate abnormal single posture features caused by noise by establishing a swing amplitude-curvature correlation model: 1000 groups of "tail fin swing amplitude A-body axis curvature C" data of large yellow croaker at 3 growth stages are collected to ensure coverage of different swimming states, and a linear regression model is used to fit the relationship between A and C: A=kxC+b, wherein k is the correlation coefficient (k=5 for larvae, k=8 for sub-adults, and k=12 for adults), and b is the intercept (uniformly taken as 0.05 cm), the goodness of fit The correlation coefficient must be ≥0.9 (to ensure significant correlation), and the model parameters are stored in the system. For the current frame, based on the tail fin amplitude Acurr and the body axis curvature Ccurr, and combined with the correlation model, the theoretical amplitude Atheo = k × Ccurr + b (k is matched according to the growth stage) is calculated; the deviation rate δ = |Acurr - Atheo| / Atheo is calculated to realize the attitude linkage verification of real-time attitude features.

[0119] If δ ≤ 15% (normal fluctuation range), both Acurr and Ccurr are valid and used as optimized posture features; if δ > 15%, it is considered noise and a correction is triggered: if Acurr is abnormal (Acurr > Atheo × 1.15): replace Acurr with Atheo; if Ccurr is abnormal (Ccurr < Acurr / kb × 1.15): replace Ccurr with (Ccurr + (Acurr - b) / k) / 2 (smoothing correction). The corrected Acurr and Ccurr must again meet the biological movement limits (e.g., adult Acurr ≤ 15cm, Ccurr ≤ 0.3cm). -1 This ensures that the posture characteristics conform to physiological laws, and uses the correction results as optimized posture characteristics to achieve flow rate fallback for swing amplitude and curvature verification.

[0120] This invention establishes individual motion fingerprints, respecting the individual differences of organisms, making the feature extraction method more targeted and naturally more accurate; the phase weighting mechanism delves into the motion cycle, distinguishing between effective and ineffective motion, and the extracted features are more representative of the actual propulsion force generation process; the linkage verification mechanism links two originally independently calculated features, using the biological physical constraints for cross-validation, greatly enhancing the system's anti-interference ability and the reliability of the output results; the three-level mechanism is interlocked and ultimately forms a large closed loop with biomechanical constraints, constructing a highly intelligent attitude feature optimization algorithm.

[0121] S5. Construct an attitude-fluid coupling model based on the optimized attitude features and the real-time environmental flow velocity data, and determine the swimming speed of the large yellow croaker through the attitude-fluid coupling model;

[0122] In one embodiment, step S5 includes:

[0123] The real-time environmental flow velocity data is interpolated to obtain the water flow velocity vector of the target water area;

[0124] Based on the optimized posture characteristics, the propulsion speed of the large yellow croaker relative to the water flow is determined;

[0125] The attitude fluid coupling model is constructed based on the water flow velocity vector and the propulsion velocity, and the swimming speed of the large yellow croaker is determined based on the attitude fluid coupling model.

[0126] Specifically, the present application collects real-time environmental flow velocity data of different spatial points in real time by deploying a plurality of water flow sensors (such as Doppler current meters) in the target water area, which includes water flow speed (size: m / s) and direction (angle: included angle with the north direction), and interpolates the environmental flow velocity data to generate a flow velocity vector of the target water area, i.e. the water flow velocity vector of any position.

[0127] Based on the historical posture features of the corresponding historical period of the real-time posture features of the large yellow croaker, the propulsion speed of the large yellow croaker relative to the water flow is fitted by an empirical formula: v fish to water = k1 x A + k2 x C + ε (A and C are tail fin amplitude and body axis curvature, respectively; k1 and k2 are coefficients calibrated by experiments; ε is an error correction term, which is determined according to experiments and is preferably 10 -5 ).

[0128] Finally, the water flow velocity vector and the propulsion speed are added to obtain the real swimming speed, and the whole velocity summation formula is the posture fluid coupling model. Finally, the real-time posture features of the large yellow croaker are substituted into the posture fluid coupling model to output the real swimming speed of the large yellow croaker. The above embodiment is to obtain the real swimming speed of the large yellow croaker by summation formula calculation, in addition to which, the corresponding AI (Artificial Intelligence, artificial intelligence) algorithm model can also be used to realize: a sample data set containing the water flow velocity vector and the propulsion speed is constructed, and the sample data is labeled with data result labels, which are used to represent the real swimming speed corresponding to the water flow velocity vector and the propulsion speed; then based on the learning algorithm, the AI algorithm model is trained by neural network learning based on the above sample data set. In the training process, the convolutional neural network model can be retrained or fine-tuned according to the method of the prior art to improve the generalization ability of the model; finally, the trained convolutional neural network model is obtained. In actual application, the water flow velocity vector and the propulsion speed are input into the convolutional neural network model, which is analyzed and processed by the model, and then the relevant data results (i.e. the corresponding real swimming speed) are output. Through the processing process of the above algorithm model, the convolutional neural network model can have the ability to input the water flow velocity vector and the propulsion speed, logically deduce and output the real swimming speed. It should be noted that the above training method is only an example, and those skilled in the art can also select other suitable methods according to the scene, such as reinforcement learning, federated learning or transfer learning, etc. Other general learning paradigms are not specifically limited in the present embodiment. In addition, other formula calculation data can also be realized by using this AI algorithm model, which is not specifically limited herein.

[0129] The application establishes a mathematical model based on the extracted tail fin swing, body axis curvature and obtained environmental flow velocity vector field when creating a posture fluid coupling model, which comprehensively considers the swimming posture, power output of the large yellow croaker and the influence of water flow, overcomes the influence of the complex underwater environment, and realizes accurate calculation of the real swimming speed of the large yellow croaker.

[0130] In addition, the calculation of the swimming speed of the large yellow croaker can also be performed by calculating the apparent speed of the fish body relative to the camera: for the large yellow croaker with the same tracking ID in the continuous frames (time interval 0.02 seconds), the three-dimensional coordinates of the key points at the front end of the head are extracted, the displacement of the two is calculated, and the spatial distance of the displacement of the two is calculated. The apparent speed of the large yellow croaker can be obtained by dividing the calculated spatial distance by the time interval, and the direction is the displacement vector direction; the apparent speed and the water flow speed vector are superimposed in X, Y and Z three-dimensional vectors, and the three-dimensional vector of the real speed is obtained, and the spatial distance of the three-dimensional vector is calculated, and the real swimming speed of the large yellow croaker is obtained. The fish body motion and the water flow environment are decoupled, the fish can accurately analyze how to adapt to the flow field, and provide quantitative basis for ecological research; the water flow interference is eliminated through vector superposition, the fish body propulsion mechanism is analyzed alone, and the accuracy of motion analysis is improved; in aquaculture, the water flow design can be optimized to reduce the energy consumption of fish swimming; the flow rate data after interpolation processing fills the blank of the sparse distribution of the sensor, and meets the modeling demand in the complex flow field.

[0131] In an embodiment, the method used in the application is compared with the traditional closed swimming pool test method (artificial water flow environment), and the comparison results are shown in the following table:

[0132] Table 3 comparison table of optimization effect

[0133]

[0134] As shown in Table 3, compared with the traditional scheme, the method used in the application improves the mAP@0.5 index by 8.6% and the measurement error by 65%; it can be seen that the method is significantly better than the traditional closed swimming pool test method in detecting the mAP@0.5 index and the measurement error, and is more suitable for monitoring the swimming speed of the large yellow croaker in the complex underwater environment.

[0135] In addition, in view of the dynamic adaptability requirement of the improved YOLOv7 network model in the underwater large yellow croaker monitoring scene (such as model performance degradation caused by environmental changes and fish growth), the application also proposes a multi-modal double-factor driven gradual adaptive training method. The method breaks through the limitations of traditional static training, adopts three-stage training of basic ability foundation, dynamic scene adaptation and online performance iteration, and combines the double-factor triggering mechanism of environment and fish body characteristics to realize continuous optimization of the model in the complex underwater scene.

[0136] 1. Multimodal foundational capability training (offline): This is used to build the model's ability to recognize the core features (morphology, posture) of the large yellow croaker and basic underwater scenes, laying the foundation for generalization. The specific training process includes:

[0137] First, a multimodal hybrid dataset is constructed by integrating three types of data: underwater images captured by a wide-angle camera (including scenes with different lighting and turbidity), three-dimensional coordinate data synchronously recorded by a laser rangefinder (used to generate physical scale labels for fish), and manually labeled "fish bounding boxes + key points + environmental labels (such as light intensity and turbidity level)". At the same time, these datasets need to cover different growth stages of large yellow croaker (juvenile to adult), typical postures (straight swimming, turning, tail wagging), and five typical underwater environments (clear, slightly turbid, bubble interference, net cage shadow, and sudden changes in lighting).

[0138] Next, using the laser ranging data in the dataset, a "fish posture-convolution offset" mapping pair is generated (such as the offset direction of the convolution kernel corresponding to the curved fish posture, etc. The mapping pair can be transformed into changes in spatial coordinates through changes in fish posture, and this change and the corresponding convolution features are input into the deep learning model, allowing the network to learn the mapping between this coordinate change and the convolution feature offset). Then, through supervised learning, the multi-scale deformable convolution in the model learns the morphological change law of the underwater fish in advance, and then the parameters of this layer are frozen to participate in the subsequent overall training, accelerating the model convergence, so as to realize the separate pre-training of the multi-scale deformable convolution layer in the backbone network.

[0139] Then, in the training of the neck network, an environment-aware loss is introduced. This loss is obtained by applying a weighted channel attention loss (a penalty term that enhances the weights of fish body feature channels) to samples with different environment labels. This forces the model to learn to suppress interfering features of specific environments (such as the high-brightness channel of bubbles) during the training phase, thereby achieving joint optimization of the dual attention mechanism. The environment-aware loss is expressed by the following formula:

[0140]

[0141] In the formula, L env Loss of environmental perception; e k For environmental labels (if the sample belongs to the first category) k Class environment e k =1); K This represents the total number of environmental labels. C Total number of channels; M [ k,c [For the model, the first] k The level of attention given to interfering channels in similar environments; w c For the firstc the greater the loss value, the more the model focuses on the interference channel, and the weight of the interference channel needs to be reduced through back propagation.

[0142] 2. Dynamic scene adaptation training (semi-offline), used to adapt the model to the personalized scene of a specific monitoring area (such as the structure of a specific breeding net cage, the optical properties of the water body caused by the local water temperature), to reduce the performance fluctuation in the early deployment, and the specific training process includes:

[0143] First, the model obtained by the foundation training is used as the pre-training weight, the first 3 layers of the backbone network are frozen (to retain the general feature extraction capability), and only the neck and head networks are fine-tuned.

[0144] Then, a small amount of samples (about 500 images) of the target monitoring area are collected, and "region-specific distortion samples" are generated using laser ranging data (to simulate the optical refraction law of the region), and the sample diversity is expanded through a GAN network (such as generating "net cage shadow + fish" mixed samples specific to the region), specifically, a large number of net cage shadow images and fish images are collected, preprocessed and labeled; then, during the training of the CGAN, the features of the net cage shadow images and the features of the fish images are input into the generator as conditional information together with the noise vector, and the generator tries to generate images that fuse the net cage shadow and the fish, and the discriminator is responsible for distinguishing the generated mixed images from the real mixed images. Through the adversarial training of the generator and the discriminator, the parameters of the generator are constantly optimized, so that it can generate realistic "net cage shadow + fish" mixed samples. For the generation of other mixed samples, the application only sets the corresponding conditions according to the application scenarios and requirements and adjusts the parameters of the model.

[0145] Finally, the prior knowledge of the pose-fluid coupling model is introduced: a "key point physical consistency constraint" is added to the loss function - when the tail fin swing and body axis curvature predicted by the model deviate from the physical scale calculated by the laser ranging by more than a threshold value, the key point loss weight is increased in a linear or nonlinear manner to ensure that the pose features output by the model conform to the biomechanical laws. For example, set a basic key point loss weight w0, when the physical scale deviation δ exceeds the threshold T, increase the weight in a linear manner where k is the increase coefficient, which can be determined by experiments to determine a suitable value, generally not more than 2, to better adapt to the influence of the physical scale deviation on the key point loss weight.

[0146] 3. Online incremental iterative training (real-time), used to learn new scene changes (such as changes in water quality caused by seasons, changes in morphology after fish growth) in real time after the model is deployed, to avoid performance degradation, and the specific training process includes:

[0147] First, store the "high-confidence detection results + synchronous multi-modal data" output by the model in real time (e.g., automatically retain samples with a detection confidence greater than 0.85, along with laser ranging data and water flow sensor data), forming a dynamic sample pool (window size set to 5000, discard the earliest samples when exceeded).

[0148] Next, every 1000 new samples are accumulated, and incremental training is started: the currently deployed model is used as the teacher model, and the newly trained temporary model is used as the student model; when the student model learns new samples, the prediction distribution on historical samples is aligned with the teacher model through knowledge distillation loss (KL divergence) to avoid catastrophic forgetting (e.g., forgetting the characteristics of juvenile fish previously learned); for new disturbances that appear in new samples (e.g., increased algae in summer leading to green water), automatically adjust the weight learning strategy of channel attention (enhance the differentiation channel between fish body and green background), such as when new disturbances appear, first extract features from new samples, then extract global features through global average pooling and global maximum pooling operations; then, use 1D convolution operations to enhance the interaction between different channels to capture the feature differences between fish body and background under new disturbances; use the sigmoid activation function to generate adaptive channel weights, automatically adjust the channel weights according to the characteristics of new disturbances, so that the model can pay more attention to channels that are important for distinguishing fish body and new disturbances (e.g., green water).

[0149] At the same time, the triggering of this training method needs to be combined with the performance degradation index and the environmental fish body double factors to realize the combination of passive correction and active prevention: monitor the following indicators in real time, and trigger the training process if any of the indicators exceeds the threshold for 500 consecutive frames: detection accuracy is less than 85% (normal state ≥ 90%); physical coordinate error of head front / tail fin tip > 5 cm (normal state ≤ 3 cm); ID of the same fish body switches > 3 times within 10 seconds (normal state ≤ 1 time), automatically start online incremental training, call the latest 1000 samples in the sliding window sample pool, fine-tune the head and neck network (training rounds = 5, avoid overfitting), and seamlessly replace the original model after training; at the same time, monitor the water turbidity (NTU value) in real time, which changes by more than 50% from the initial deployment, or the light intensity fluctuates by more than 30% (e.g., sudden sunny weather after rainy weather), and through laser ranging data statistics, the average body length of fish body grows by more than 20% from the initial deployment (growth leads to morphological changes), start dynamic scene adaptation training in advance, supplement 200 new samples (targeted to cover the current environment and fish body state), fine-tune with historical core samples (e.g., retain 30% of the initial samples), and prevent performance degradation.

[0150] The application introduces the physical scale information of laser ranging into training when training the model, solves the pixel physical mapping deviation caused by underwater image distortion, makes the attitude features output by the model more in line with the actual biomechanical laws, breaks through the passive mode of retraining after the traditional performance degradation, actively adapts to dynamic changes through early warning of the environment and fish growth, reduces monitoring interruption, retains historical knowledge in online training, avoids the model forgetting the features learned in the early stage, realizes long-term stable performance iteration, enables the improved YOLOv7 model to quickly adapt to specific scenarios in the early stage of deployment and continuously adapt to environmental and fish changes in long-term use, and provides stable and reliable target detection and attitude feature extraction capability for large yellow croaker swimming speed monitoring.

[0151] It should be noted that although each step in the above flowchart is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps has no strict order limitation, and these steps can be executed in other orders.

[0152] In another embodiment, as shown in Figure 2 The second aspect of the application provides a large yellow croaker swimming speed monitoring system based on target detection, comprising:

[0153] A data acquisition module 11 is configured to acquire a to-be-monitored image of a target water area, real-time environmental flow rate data, and real-time water turbidity data; a target detection module 12 is configured to perform target detection on the to-be-monitored image to obtain coordinate data of a large yellow croaker; a linkage verification module 13 is configured to perform linkage verification on the coordinate data to obtain verified coordinate data; the linkage verification is selected according to an analysis result of the coordinate data, and at least includes: a first-level verification of verifying a motion state reflected by the coordinate data based on biomechanical constraints; and / or, a second-level verification of correcting a flow rate vector confirmed based on the real-time water turbidity data; and / or, a third-level verification of correcting based on a motion law reflected by the coordinate data and the biomechanical constraints; an attitude optimization module 14 is configured to determine real-time attitude features of the large yellow croaker according to the verified coordinate data, and optimize the real-time attitude features through a flow rate phase joint mechanism to obtain optimized attitude features; and a speed quantification module 15 is configured to construct an attitude fluid coupling model based on the optimized attitude features and the real-time environmental flow rate data, and determine a swimming speed of the large yellow croaker through the attitude fluid coupling model.

[0154] The above embodiments only express several preferred embodiments of the present application, which are described in a more specific and detailed manner, but cannot be understood as a limitation to the patent scope of the application. It should be pointed out that, for ordinary skilled in the art, several improvements and replacements can be made without departing from the technical principles of the present application, and these improvements and replacements should also be considered as the protection scope of the present application. Therefore, the protection scope of the present application patent should be subject to the protection scope of the claims.

Claims

1. A method for monitoring the swimming speed of Pseudosciaena crocea based on target detection, characterized in that, The method comprises the following steps: acquiring a to-be-monitored image, real-time environmental flow rate data and real-time water turbidity data of a target water area respectively; performing target detection on the to-be-monitored image to obtain coordinate data of Pseudosciaena crocea; performing linkage verification on the coordinate data to obtain verified coordinate data; the linkage verification is selected according to an analysis result of the coordinate data, and at least includes: a first-level verification of verifying a motion state reflected by the coordinate data based on biomechanical constraints; and / or, a second-level verification of correcting a flow rate vector confirmed based on the real-time water turbidity data; and / or, a third-level verification of correcting based on a motion law reflected by the coordinate data and the biomechanical constraints; determining real-time posture features of the Pseudosciaena crocea according to the verified coordinate data, and optimizing the real-time posture features through a flow rate phase linkage mechanism to obtain optimized posture features; constructing a posture fluid coupling model based on the optimized posture features and the real-time environmental flow rate data, and determining a swimming speed of the Pseudosciaena crocea through the posture fluid coupling model; the verified coordinate data includes three-dimensional coordinates of a head front end, three-dimensional coordinates of a body axis midpoint, three-dimensional coordinates of a tail fin root and three-dimensional coordinates of a tail fin tip of the Pseudosciaena crocea; wherein the determination of the real-time posture features of the Pseudosciaena crocea according to the verified coordinate data comprises: calculating a relative displacement of the three-dimensional coordinates of the tail fin tip relative to the three-dimensional coordinates of the tail fin root in a single frame, performing sliding window analysis on the relative displacements in continuous frames, and taking a maximum value in the window as a tail fin swing amplitude of the continuous frame period; constructing a body axis curve of the Pseudosciaena crocea based on the three-dimensional coordinates of the head front end, the three-dimensional coordinates of the body axis midpoint and the three-dimensional coordinates of the tail fin root, and determining a body axis curvature based on the body axis curve, so as to take the tail fin swing amplitude as the real-time posture features; the optimization of the real-time posture features through the flow rate phase linkage mechanism to obtain the optimized posture features comprises: previously constructing a mapping table of tail swing periods and swimming states of each Pseudosciaena crocea in the target water area; the tail swing period is determined according to the three-dimensional coordinates of the tail fin tip, and the swimming state is determined according to the body axis curvature; determining a swimming state of the Pseudosciaena crocea in a current frame to match the mapping table to obtain a current tail swing period to dynamically adjust the number of frames of the sliding window to obtain a target sliding window; extracting phase features in the target sliding window to perform phase division to obtain a division result, and dynamically assigning weights to the division result according to phase types to perform weighted calculation to obtain a weighted swing amplitude; determining a theoretical swing amplitude of the current frame based on the body axis curvature and a previously constructed swing amplitude-curvature correlation model, and performing deviation judgment on the weighted swing amplitude and the theoretical swing amplitude to realize posture linkage verification of the real-time posture features; when it is determined to be normal, taking the weighted swing amplitude and the body axis curvature as the optimized posture features; when it is determined to be abnormal, performing smoothing correction on the body axis curvature, and taking a correction result and the theoretical swing amplitude as the optimized posture features; The posture fluid coupling model is constructed based on the optimized posture feature and the real-time environmental flow velocity data, and the swimming speed of the large yellow croaker is determined through the posture fluid coupling model, including: The real-time environmental flow velocity data is interpolated to obtain a water flow velocity vector of the target water area; Based on the optimized posture feature, a propulsion speed of the large yellow croaker relative to the water flow is determined; The posture fluid coupling model is constructed according to the water flow velocity vector and the propulsion speed, and the swimming speed of the large yellow croaker is determined according to the posture fluid coupling model.

2. The target detection-based large yellow croaker swimming speed monitoring method according to claim 1, characterized in that, The linkage verification is performed on the coordinate data to obtain verified coordinate data, including: The coordinate data is converted to obtain three-dimensional coordinate data, and the body length of the large yellow croaker is determined according to the three-dimensional coordinate data; A biological motion limit database of the large yellow croaker at multiple growth stages is constructed, and the biological motion limit database is matched with the body length to realize the first-level biomechanics constraint verification of the large yellow croaker, and the three-dimensional coordinate data with a matching result of success is taken as the verified coordinate data; When it is determined that the matching result fails, a second-level compensation mechanism is triggered, the coordinate data is input into a pre-constructed grid flow velocity mapping table for spatial matching to obtain a flow velocity vector, the theoretical coordinates of the current frame are predicted by combining the coordinate data of the preset frame, and the theoretical coordinates are residual corrected and coordinate corrected based on the real-time water body turbidity data to obtain corrected coordinates as the verified coordinate data; When it is detected that the corrected coordinates are invalid, a third-level completion mechanism is triggered, three-dimensional valid coordinate data of the large yellow croaker are obtained and fitted using a three-order Bezier curve to obtain an individual trajectory equation for extrapolation prediction to generate predicted coordinates, the predicted coordinates are subjected to the first-level biomechanics constraint verification, if the matching is successful, the predicted coordinates are taken as the verified coordinate data of the large yellow croaker, otherwise, the control point weight of the Bezier curve is adjusted for re-prediction until the constraint is met.

3. The method according to claim 2, wherein, The coordinate data is converted to obtain three-dimensional coordinate data, including: The Z coordinate of the large yellow croaker is determined through the distance from the key point of the large yellow croaker to the narrow-beam laser range finder; Based on the Z coordinate, the coordinate data is distortion corrected through the calibration parameters of the underwater wide-angle camera to obtain the X and Y coordinates of the large yellow croaker, and the three-dimensional coordinate data of the large yellow croaker are generated in combination with the Z coordinate.

4. The target detection-based large yellow croaker swimming speed monitoring method according to claim 1, characterized in that, The target detection of the to-be-monitored image is realized by improving a YOLOv7 network model; wherein, the improved YOLOv7 network model comprises an input layer, a backbone feature extraction network layer, a neck feature fusion network layer and a head detection output layer; the backbone feature extraction network layer comprises a first ELAN block, a second ELAN block, a third ELAN block and a fourth ELAN block; the second ELAN block, the third ELAN block and the fourth ELAN block all adopt multi-scale deformation convolution; wherein, The target detection of the to-be-monitored image is realized by improving a YOLOv7 network model; wherein, the improved YOLOv7 network model comprises an input layer, a backbone feature extraction network layer, a neck feature fusion network layer and a head detection output layer; the backbone feature extraction network layer comprises a first ELAN block, a second ELAN block, a third ELAN block and a fourth ELAN block; the second ELAN block, the third ELAN block and the fourth ELAN block all adopt multi-scale deformation convolution; wherein, The to-be-monitored image is converted into a standardized tensor through the input layer; According to the first ELAN block, the standardized tensor is subjected to primary feature extraction to obtain first features for representing fish body edges; According to the second ELAN block, the first features are subjected to secondary feature extraction to obtain second features for representing fish body contours; According to the third ELAN block, the second features are subjected to tertiary feature extraction to obtain third features for representing fish body shapes; According to the fourth ELAN block, the third features are subjected to quaternary feature extraction to obtain fourth features for representing distant fish bodies.

5. The target detection-based bighead carp swimming speed monitoring method according to claim 4, characterized in that, The neck feature fusion network layer comprises a PAN block, a channel attention layer and a spatial attention layer; the PAN block comprises an up-sampling path and a down-sampling path; wherein, The target detection on the to-be-monitored image to obtain coordinate data of the large yellow croaker further comprises: In the up-sampling path, the fourth features are subjected to up-sampling processing and then spliced with the third features to obtain fifth features, and the fifth features are subjected to up-sampling processing and then spliced with the second features to obtain sixth features; In the down-sampling path, the sixth features are subjected to down-sampling processing and then spliced with the fifth features to obtain seventh features, and the seventh features are subjected to down-sampling processing and then spliced with the fourth features to obtain eighth features; The sixth features, the seventh features and the eighth features are all input into the channel attention layer and the spatial attention layer to respectively perform weighted processing on corresponding features in channel dimensions and spatial dimensions to obtain first enhanced features, second enhanced features and third enhanced features.

6. The bighead carp swimming speed monitoring method based on target detection according to claim 5, characterized in that, The head detection output layer comprises a first parallel detection branch, a second parallel detection branch and a third parallel detection branch; the first parallel detection branch comprises a first detection sub-branch and a first key point prediction sub-branch; The second parallel detection branch comprises a second detection sub-branch and a second key point prediction sub-branch; The third parallel detection branch comprises a third detection sub-branch and a third key point prediction sub-branch; wherein, The target detection on the to-be-monitored image to obtain coordinate data of the large yellow croaker further comprises: The first enhanced features are respectively input into the first detection sub-branch and the first key point prediction sub-branch for processing to obtain first bounding box coordinate data and first key point coordinate data; The second enhanced features are respectively input into the second detection sub-branch and the second key point prediction sub-branch for processing to obtain second bounding box coordinate data and second key point coordinate data; The third enhanced features are respectively input into the third detection sub-branch and the third key point prediction sub-branch for processing to obtain third bounding box coordinate data and third key point coordinate data; The first bounding box coordinate data, the second bounding box coordinate data, the third bounding box coordinate data, the first key point coordinate data, the second key point coordinate data and the third key point coordinate data are combined to generate the coordinate data of the large yellow croaker.

7. A target detection-based monitoring system for swimming speed of Pseudosciaena crocea, characterized in that, Comprise: The data acquisition module is used for acquiring a to-be-monitored image of a target water area, and real-time environmental flow rate data and real-time water turbidity data of the target water area; The target detection module is used for performing target detection on the to-be-monitored image to obtain coordinate data of the large yellow croaker; The linkage verification module is used for performing linkage verification on the coordinate data to obtain verified coordinate data; The linkage verification is selected according to an analysis result of the coordinate data, and at least includes: a first-level verification of verifying a motion state reflected by the coordinate data based on a biomechanics constraint; and / or a second-level verification of correcting a flow rate vector confirmed based on the real-time water turbidity data; and / or a third-level verification of correcting based on a motion law reflected by the coordinate data and the biomechanics constraint; The posture optimization module is used for determining real-time posture features of the large yellow croaker according to the verified coordinate data, and optimizing the real-time posture features through a flow rate phase joint mechanism to obtain optimized posture features; The speed quantification module is used for constructing a posture fluid coupling model based on the optimized posture features and the real-time environmental flow rate data, and determining a swimming speed of the large yellow croaker through the posture fluid coupling model; The verified coordinate data includes a head front end three-dimensional coordinate, a body axis midpoint three-dimensional coordinate, a tail fin root three-dimensional coordinate and a tail fin tip three-dimensional coordinate of the large yellow croaker; wherein The real-time posture features of the large yellow croaker are determined according to the verified coordinate data, including: A relative displacement of the tail fin tip three-dimensional coordinate relative to the tail fin root three-dimensional coordinate in a single frame is calculated to perform sliding window analysis on the relative displacement in continuous frames, and a maximum value in the window is taken as a tail fin amplitude of the continuous frame period; A body axis curve of the large yellow croaker is constructed based on the head front end three-dimensional coordinate, the body axis midpoint three-dimensional coordinate and the tail fin root three-dimensional coordinate, and a body axis curvature is determined based on the body axis curve to take the tail fin amplitude as the real-time posture features; The real-time posture features are optimized through the flow rate phase joint mechanism to obtain optimized posture features, including: A mapping table of tail beating periods and swimming states of each large yellow croaker in the target water area is constructed in advance; the tail beating period is determined according to the tail fin tip three-dimensional coordinate, and the swimming state is determined according to the body axis curvature; A swimming state of the large yellow croaker in a current frame is determined to match the mapping table to obtain a current tail beating period to dynamically adjust a frame number of the sliding window to obtain a target sliding window; Phase features in the target sliding window are extracted to perform phase division to obtain a division result, and weights are dynamically assigned to the division result according to phase types to perform weighted calculation to obtain a weighted amplitude; The body axis curvature and a pre-constructed amplitude curvature correlation model are used to determine a theoretical amplitude of the current frame, and deviation determination is performed on the weighted amplitude and the theoretical amplitude to realize posture linkage verification on the real-time posture features. When the determination result is normal, the weighted swing and the body axis curvature are taken as the optimized posture features; when the determination result is abnormal, the body axis curvature is modified by smoothing, and the modified result and the theoretical swing are taken as the optimized posture features; The posture fluid coupling model is constructed based on the optimized posture features and the real-time environmental flow velocity data, and the swimming speed of the large yellow croaker is determined through the posture fluid coupling model, including: The real-time environmental flow velocity data is processed by interpolation to obtain a water flow velocity vector of the target water area; Based on the optimized posture features, a propulsion speed of the large yellow croaker relative to the water flow is determined; The posture fluid coupling model is constructed according to the water flow velocity vector and the propulsion speed, and the swimming speed of the large yellow croaker is determined according to the posture fluid coupling model.

Citation Information

Patent Citations

  • Deep learning-based fish school toxicity behavior analysis method and system

    CN120148110A

  • Real-time attitude estimation method of underwater inertial navigation system

    CN120863846A