Underwater fish length measurement method based on binocular vision and pose estimation

By combining binocular vision and attitude estimation methods with an underwater contextual attention enhancement module and geometric statistical filtering, the problem of feature localization difficulties and measurement distortion caused by non-rigid swimming in underwater fish length measurement is solved, achieving high-precision and stable fish length measurement.

CN122391329APending Publication Date: 2026-07-14DALIAN OCEAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN OCEAN UNIV
Filing Date
2026-04-30
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing technologies for measuring the length of underwater fish suffer from difficulties in feature localization due to complex underwater environments and measurement distortion caused by non-rigid swimming. In particular, under the influence of factors such as light refraction, water turbidity, fish body bending, and tilting posture, it is difficult to guarantee measurement accuracy and stability.

Method used

A binocular vision and attitude estimation-based approach is adopted, which combines an underwater contextual attention enhancement module (UCA) and a geometric statistical filtering engine into a cascaded perception and measurement framework. Through topological perception of visual features and spatiotemporal joint filtering, the accurate capture of key points of the fish body and the suppression of non-rigid swimming deformation are achieved.

Benefits of technology

It significantly improves the measurement anti-distortion capability and overall output steady-state performance in underwater dynamic environments, and achieves high-precision and stable measurement of fish length.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391329A_ABST
    Figure CN122391329A_ABST
Patent Text Reader

Abstract

The application provides a kind of underwater fish length measurement method based on binocular vision and attitude estimation, belongs to image recognition and underwater measurement technical field.The method constructs a new type of cascaded perception measurement framework combined with underwater context attention enhancement module and geometric statistical filtering engine;The framework is based on a top-down target detection and key point positioning cascaded network, introduces an underwater context attention enhancement module at the key point feature fusion end, and introduces a spatial orthogonal circular arc compensation and double-stage time signal filtering mechanism at the three-dimensional object understanding algorithm end, through the collaborative optimization of visual feature topology perception and space-time joint filtering, realizes the accurate capture of fish key points under light and shadow interference, and effectively suppresses the non-rigid swimming deformation and depth noise jump, effectively solves the distortion problem of underwater dynamic length measurement, significantly improves the anti-distortion ability and overall output stability of measurement in underwater dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image recognition and underwater measurement technology, and relates to an underwater fish length measurement method based on binocular vision and attitude estimation. Background Technology

[0002] Fish farming is a crucial component of aquaculture, with high-density and intensive farming methods increasingly becoming the primary means of improving production efficiency. However, traditional contact-based manual measurement methods easily induce strong stress responses in fish, leading to physical damage and even death, which has become a key issue hindering the sustainable and healthy development of the industry. Growth data monitoring is an effective means of evaluating aquaculture benefits, and automatic fish length measurement, as a fundamental technology of smart fisheries, plays a pivotal role. Achieving non-contact and accurate extraction of underwater fish dimensions using machine vision-based methods provides a reliable basis for subsequent scientific feeding and growth assessment, and is an important guarantee for the sustainable development of aquaculture.

[0003] However, achieving this goal faces two major challenges: the difficulty of feature localization caused by the complex underwater environment, and the measurement distortion caused by non-rigid swimming and physical distortion. In the field of aquaculture, the low contrast and blurriness of images due to light refraction and water turbidity limit the accuracy of key point localization on the fish body. At the same time, the bending of the fish's body and the tilting of its posture during swimming, combined with the radial distortion at the edge of the underwater camera and the phenomenon of deep reflection penetration, can easily induce violent fluctuations in the measurement values, increasing the difficulty of accurate steady-state measurement.

[0004] Visual feature extraction technology is fundamental to improving measurement accuracy. Early studies, such as the underwater binocular video measurement method proposed by Harvey et al., while achieving high measurement accuracy in controlled environments, suffer from inefficiencies due to their semi-automated operation logic heavily reliant on manual calibration. To reduce hardware deployment costs, Monkman et al. proposed a measurement method based on a monocular camera and a physical reference (fiducial marker), correcting parallax errors using a reference scale. While this method effectively lowers the equipment barrier, its limitation lies in the extreme difficulty of arranging physical references in real large-scale aquaculture cages, and the severe perspective distortion caused by changes in the relative depth between the fish and the reference.

[0005] Given the high degree of interplay between non-rigid deformation of fish and underwater optical noise, a stable measurement model with superior anti-interference performance is crucial. With the development of deep learning, more and more studies are attempting to calculate fish length through automatic segmentation and keypoint detection. Zhou Minggang et al. proposed an underwater measurement system based on a depth camera, using the GrabCut algorithm to segment the fish body and employing a skeleton extraction algorithm to address fish body curvature. Although this strategy can reconstruct the fish outline under specific conditions, in actual production, traditional segmentation algorithms such as GrabCut are highly susceptible to interference from water speckle and local target occlusion features, leading to skeleton breakage or geometric truncation. Furthermore, most of the aforementioned studies focus on independent geometric calculation tasks for single-frame images, and have not yet explored the non-biological, drastic jumps in measurement values ​​caused by edge distortion and high-frequency depth-penetration noise in continuous video streams. To address this deficiency, it is urgent to introduce spatial distortion isolation and temporal dynamic filtering mechanisms to achieve industrial-grade stable measurement output.

[0006] In summary, while computer vision technology has made some progress in underwater fish size measurement, it still faces the following challenges:

[0007] (1) Existing automated extraction methods are too dependent on environmental contrast and segmentation accuracy, and are prone to positioning drift in complex backgrounds, which affects the accuracy of key point extraction.

[0008] (2) Existing studies are mostly limited to linear distance or single-frame independent calculation, which fails to effectively compensate for geometric distortion caused by non-rigid body bending, and lacks a spatiotemporal joint filtering mechanism for numerical jumps caused by depth noise. Summary of the Invention

[0009] To address the aforementioned problems in existing technologies, this invention proposes a method for accurate underwater fish length measurement based on binocular vision and attitude estimation. This method constructs a novel cascaded perception and measurement framework combining an Underwater Contextual Attention (UCA) module and a geometric statistical filtering engine. Based on a top-down cascaded network for target detection and keypoint localization, the framework introduces an Underwater Contextual Attention (UCA) module at the keypoint feature fusion stage and spatial orthogonal circular arc compensation and a two-order temporal signal filtering mechanism at the 3D physics computation stage. Through the synergistic optimization of visual feature topology perception and spatiotemporal joint filtering, it achieves accurate capture of key points on the fish body under light and shadow interference, and effective suppression of non-rigid swimming deformation and depth noise jumps, significantly improving the anti-distortion capability and overall output stability of measurements in dynamic underwater environments.

[0010] The technical solution of the present invention is as follows:

[0011] A method for measuring the length of underwater fish based on binocular vision and pose estimation includes the following steps:

[0012] S1: Acquire several images of the fish to be tested in real underwater scenes and the corresponding 3D point cloud depth map matrix;

[0013] S2: Input all underwater fish images to be tested into the keypoint perception network model to obtain the two-dimensional pixel coordinates of three core physical key points of all target individuals in each underwater fish image: the front end of the fish mouth, the physical midpoint of the fish body lateral line, and the end of the fish tail. The keypoint perception network model is a two-level hierarchical architecture consisting of a front-end target detection network and a back-end keypoint localization network. The front-end target detection network extracts the target two-dimensional bounding box of the underwater fish image and obtains a local image containing a single target individual based on the target two-dimensional bounding box, realizing macroscopic target locking and redundant background removal. The back-end keypoint localization network obtains the two-dimensional pixel coordinates of the three core physical key points of the fish mouth, the physical midpoint of the fish body lateral line, and the end of the fish tail based on the local image.

[0014] S3: Based on the two-dimensional pixel coordinates of three physical key points of all target individuals in the underwater fish images and the corresponding three-dimensional point cloud depth map matrix, the true physical length of the fish to be tested is calculated using spatial orthogonal circular arc compensation and two-order temporal signal filtering mechanism.

[0015] Furthermore, step S1 specifically includes:

[0016] S1.1: Use an underwater binocular camera to record videos of the fish activity to be tested, and obtain a continuous video material sequence containing thousands of original images; execute an offline frame extraction strategy to filter and extract hundreds of key image frames from the video material sequence as samples for subsequent calculation; use the camera's distortion and rotation translation matrix to perform epipolar correction on the sample images, and project the left and right images onto a coplanar and row-aligned plane.

[0017] S1.2: Perform stereo matching on the corrected sample images. For the left eye image, the pixel coordinates are... The feature points are in the same row of the right eye image. Search for corresponding coordinates And calculate pixel disparity ;

[0018] S1.3: Combined with the factory-calibrated intrinsic focal length of the binocular camera. Distance between physical extrinsic baseline and the optical center of the left and right lenses The principle of triangulation is used to convert pixel parallax into the true physical depth of the target distance from the camera. ;

[0019] S1.4: The true physical depth values ​​of all pixels in the obtained sample image. The two-dimensional pixel coordinates of the left eye image are arranged in an array to generate a three-dimensional point cloud depth map matrix that is completely consistent with the resolution of the original two-dimensional color image and is spatially physically aligned.

[0020] Furthermore, the front-end object detection network acquires local images containing a single object based on the following method:

[0021] The front-end target detection network takes an underwater image of the fish to be tested as input and extracts multi-scale feature maps from shallow to deep through layer-by-layer downsampling. The shallow features retain spatial details such as the fish's edges and scale textures, while the deep features expand the receptive field to extract global semantic information. Subsequently, the multi-scale feature maps are fused across scales and output corresponding refined feature maps to adapt to the drastic scale fluctuations of the underwater target. Finally, the target's two-dimensional bounding box is predicted based on the refined feature map, and the original underwater image of the fish to be tested is precisely spatially cropped according to the target's two-dimensional bounding box to remove complex water backgrounds and output a local image containing only a single target individual.

[0022] Furthermore, the process of the front-end target detection network acquiring a local image containing a single target individual is implemented using the RTMDet model.

[0023] Furthermore, the backend key point localization network obtains the two-dimensional pixel coordinates of three core physical key points—the front end of the fish mouth, the physical midpoint of the fish's lateral line, and the end of the fish tail—based on the following method:

[0024] The backend keypoint localization network consists of a feature extraction network, a neck network embedded with an underwater contextual attention enhancement (UCA) module, and a sub-pixel keypoint prediction head connected in series. The feature extraction network, as the backbone network, performs deep convolutional extraction on the local image of a single target individual to obtain a multi-scale feature map containing the target's physical topology. The neck network receives this multi-scale feature map and performs anti-spectral feature decoupling and 2D topology locking processing through the UCA module. After effectively suppressing feature drift caused by water surface reflection and target local occlusion, a refined fused feature map with spatial crosstalk eliminated is obtained. The sub-pixel keypoint prediction head uses the refined fused feature map as input and employs a SimCC-based decoding network to perform sub-pixel-level spatial probability distribution regression on it, accurately outputting the 2D pixel coordinates of three core physical keypoints: the front of the fish mouth, the physical midpoint of the fish's side line, and the end of the fish tail.

[0025] Furthermore, the UCA module aims to solve the feature drift problem caused by underwater specular interference and target local overlap through feature decoupling and spatiotemporal multi-sensing mechanisms; the UCA module includes a sequentially connected pre-feature decoupling submodule, a channel attention submodule, and a two-dimensional spatial context submodule;

[0026] The pre-feature decoupling submodule receives the multi-scale feature map output by the feature extraction network, filters it through the channel-level random mask component, and outputs a preliminary screening feature map that is resistant to specular light.

[0027] The channel attention submodule takes the initial screening feature map as input and uses a global average pooling component and a global max pooling component to compress the features in the spatial dimension, outputting two spatial aggregation vectors. The two sets of vectors are then fed into a shared parameter MLP to learn the nonlinear dependencies between channels. The MLP uses a SiLU nonlinear smooth activation layer instead of the traditional hard truncation activation. The first fusion unit adds the outputs of the two MLPs element by element and then processes them through an activation function to generate a channel weight matrix. The first multiplier multiplies the initial screening feature map with the channel weight matrix channel by channel, outputting a channel-enhanced transition feature map.

[0028] The two-dimensional spatial context submodule performs average pooling and max pooling operations on the transition feature map along the channel dimension, respectively; the first stitcher stitches the outputs of the two along the channel dimension and feeds them into the first large receptive field convolutional component to forcibly lock the macroscopic two-dimensional topology of a single fish; then it passes through the batch normalization component and the SiLU smoothing activation layer in sequence, and then through the second standard convolutional component and activation function to extract the two-dimensional spatial weight matrix; the second multiplier multiplies the transition feature map and the two-dimensional spatial weight matrix element by element, and finally outputs a refined fusion feature map that eliminates spatial crosstalk and is robust to slight overlap.

[0029] Furthermore, the feature extraction network includes CSPNEXt network, ResNet50, or HRNet-w32, etc.

[0030] Furthermore, the construction process of the key point perception network model includes:

[0031] Continuous video recording of the target koi in a real water tank was performed to obtain an SVO format video stream, and the video stream was then subjected to equal-interval frame extraction.

[0032] Labelme was used for two-level customized annotation. The first level was for target localization annotation, using rectangular bounding boxes to accurately locate the koi carp's body and assigning it the category label "fish". The second level was for physical feature keypoint annotation, strictly marking the coordinates of three key points within the bounding box of each fish: the front of the mouth, the physical midpoint of the side line, and the end of the tail, labeled "head", "midbp", and "tail", respectively. After annotation, the system initially generated a JSON file containing the bounding box and keypoint coordinates. An automated script then converted the JSON file into the standard COCO dataset format (COCO-format JSON) to obtain the underwater multimodal visual dataset.

[0033] The keypoint perception network model is trained using the underwater multimodal visual dataset to obtain a trained keypoint perception network model.

[0034] Furthermore, in step S3, the true physical length of the fish under test is calculated using spatial orthogonal circular arc compensation and a two-order temporal signal filtering mechanism, specifically including:

[0035] S3.1: The two-dimensional pixel coordinates of the three physical key points of each target individual in the underwater fish image are fused with the corresponding three-dimensional point cloud depth map matrix. A 5×5 local neighborhood pixel block is extracted centered on each key point. After filtering out invalid depth values, the remaining valid depth values ​​are arranged in ascending order to construct a local depth sample set for the corresponding key point. Only the first half of the depth value samples in the local depth sample set are extracted, and the median is taken as the robust depth value of the corresponding key point. , ;

[0036] Subsequently, combined with the camera intrinsic parameter matrix Using the reverse perspective principle of a pinhole camera, the two-dimensional pixel coordinates of three physical key points are obtained. Mapped to three-dimensional physical coordinates respectively Its mapping formula is:

[0037]

[0038]

[0039]

[0040] in, The camera intrinsic parameter matrix is ​​expressed as follows:

[0041]

[0042] In the formula, This is the equivalent focal length of the camera. The coordinates of the camera's optical center;

[0043] S3.2: Establish a global field of view and cold start gating mechanism to perform pre-legality screening on the continuously acquired video image data stream of the target fish. Only when the center of the local image containing the target individual output by the front-end target detection network is within the safe field of view and the average confidence of the key points is greater than a preset threshold, is it allowed to extract the three-dimensional spatial physical coordinates of the three key points corresponding to the target individual in the underwater image of the target fish. The entry requirements for the cold start phase are:

[0044]

[0045] The average confidence level of the key points The calculation formula is:

[0046]

[0047] in, The confidence level for the three key points;

[0048] For a valid admission frame, calculate the spatial straight-line distance between adjacent keypoints to construct a spatial triangle, and define the basic polyline distance. The area and radius of curvature of the circumcircle of the triangle are calculated based on Heron's formula. and combined with the central angle Calculation of theoretical arc unfolding length ;

[0049] To prevent feature point coordinate drift from incorrectly fitting an overly curved closed loop, and to avoid non-physical stretching exceeding the biomechanical limits of the fish's spine, based on extensive underwater measurement statistical priors, the following settings are established. The polyline distance is a hard, empirical upper limit threshold for dynamic deformation. When the calculated theoretical arc crosses this threshold, the system triggers a singularity degradation mechanism, stripping away unreliable curvature expansion and forcing a return to the absolutely safe baseline polyline distance; this is used to extract safe original measurement values. Its piecewise mathematical constraint expression is:

[0050]

[0051] S3.3: Raw measurement values ​​for continuous valid admission frames output The sequence executes a two-stage temporal signal filtering mechanism to eliminate spatial high-frequency jitter; the first stage involves real-time calculation of the current single-frame raw measurement value. Sliding window with system memory history All The absolute deviation of the median, if the deviation is abnormal If this happens, the extreme value circuit breaker will be triggered, intercepting deep breakdown noise and forcibly maintaining the output of the previous frame's measurement value.

[0052]

[0053] in, The sample size of the historical sliding window is the deviation rate. Select 90 frames per second;

[0054] The second level is for circuit breakers that have not been triggered. The queue is sorted, and the top and bottom 25% of extreme outliers are removed. Data, and 50% of the measurements are retained at the center. Calculate the arithmetic mean of the sample set and output the physical total length to be corrected. :

[0055]

[0056] in, This represents the total number of samples in the set.

[0057] Furthermore, a dynamic curvature gain coefficient based on the central angle is introduced. Compensation coefficient for static refraction of water body right System-level corrections are performed, and the actual physical measurement values ​​of each target fish in a single frame of underwater fish image are continuously calculated and output. .

[0058] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, the electronic device performs the underwater fish length measurement method based on binocular vision and attitude estimation.

[0059] A storage medium comprising a computer program that, when run on an electronic device, causes the electronic device to perform the underwater fish length measurement method based on binocular vision and attitude estimation.

[0060] The beneficial effects of this invention are as follows: By synergistically optimizing visual feature topology perception and spatiotemporal joint filtering, this invention significantly improves the accuracy and output stability of length measurement in complex underwater dynamic environments. Addressing the challenge of keypoint localization drift caused by light and shadow interference and local target occlusion, this invention utilizes a keypoint perception network embedded with an underwater contextual attention enhancement module (UCA). Through channel-level random feature discarding, it effectively suppresses feature dependence caused by specular reflection. Combined with large receptive field convolution to forcibly lock the two-dimensional geometric topology of a single target, it achieves sub-pixel-level accurate capture of keypoints, providing a high-precision coordinate reference for subsequent three-dimensional physical calculations. To address the non-rigid deformation of the fish body during swimming and lens edge distortion, the system utilizes core field of view (ROI) physical constraints to avoid edge imaging distortion. Furthermore, it employs an orthogonal circular arc compensation algorithm to restore the straight-line distance between keypoints to physical arc length, correcting measurement truncation errors caused by attitude shifts and torso bending from a geometric perspective. Based on this, the present invention constructs a two-order time-series signal filtering architecture composed of extreme value fusion and tail-cutting mean fusion. It monitors and intercepts numerical collapse caused by deep breakdown in real time by historical median, and then uses a sliding window to dynamically remove transient interference generated by high-frequency swimming of fish, thus ensuring the physical consistency of the measurement sequence at the time domain level. Finally, it realizes fish length calculation output with industrial-grade steady state. Attached Figure Description

[0061] Figure 1 The dataset contains images; where (a) is the reflection, (b) is the bending of the fish body, and (c) is the reflection on the water surface.

[0062] Figure 2 Annotate the data.

[0063] Figure 3 This is an overall flowchart of the underwater fish length measurement method in an embodiment of the present invention.

[0064] Figure 4 This is a diagram of the cascaded perception logic and network architecture of the key point perception network model in this embodiment of the invention.

[0065] Figure 5 This is a channel attention topology diagram of the UCA module in an embodiment of the present invention.

[0066] Figure 6 This is a spatial context topology diagram of the UCA module in an embodiment of the present invention.

[0067] Figure 7 This is a geometric model diagram of spatial orthogonal circular arc compensation in an embodiment of the present invention.

[0068] Figure 8 This is a diagram of a two-order time-series signal filtering mechanism based on field-of-view constraints in an embodiment of the present invention.

[0069] Figure 9The following are some of the actual length measurements of fish in the embodiments of the present invention; wherein (a) fish is 15.5cm, (b) fish is 16cm, and (c) fish is 15cm.

[0070] Figure 10 The following is a diagram showing the actual effect of the system in an embodiment of the present invention; wherein, (a) is the measurement of the length of the fish reflected on the water surface in its normal state, (b) is the measurement of the length of the fish reflecting light in its normal state, (c) and (e) are the measurement of the length of the fish in its bent state, and (d) is the measurement of the length of the fish swimming close to the camera. Detailed Implementation

[0071] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0072] like Figure 3 As shown, this embodiment of the invention provides a method for measuring the length of underwater fish based on binocular vision and attitude estimation, including the following steps:

[0073] Step S10: Construct an underwater multimodal visual dataset.

[0074] This embodiment's underwater fish dataset encompasses swimming koi individuals in real-world aquaculture environments, fully demonstrating the diversity and complexity of non-rigid swimming postures. This helps evaluate the model's localization stability and generalization ability for dynamic bending deformation in continuous spatiotemporal sequences. The underwater multimodal visual dataset in this embodiment is designed around realistic clear water scenes, focusing on specific optical interference conditions such as specular reflections on the left and right sidewalls of the tank, and water surface reflections that are easily generated when koi swim close to the surface. This data distribution closely matches the optical characteristics that easily induce visual feature confusion and false detection in actual clear aquaculture scenarios.

[0075] A high-precision underwater binocular ZED camera was used to continuously record video of the target koi carp in a real water tank. High-quality SVO format video streams were acquired through the camera's native data interface, followed by objective, equally spaced frame extraction. To avoid high physical redundancy between adjacent frames and improve network learning efficiency, a basic data cleaning mechanism was implemented, ultimately selecting hundreds of valid, non-repeating left-eye RGB images and a strictly physically aligned 3D point cloud depth map matrix from thousands of basic images. Since the data originates from natural slices of the real environment, the dataset not only realistically reproduces the various dynamic tail-wagging postures of the koi carp but also naturally preserves the objectively existing optical interference conditions such as tank sidewall reflections and water surface mirror reflections, thus constructing a high-precision, robust visual measurement dataset. The collected underwater multimodal aligned data is shown below. Figure 1 As shown.

[0076] Two-level customized annotation was performed using Labelme. First, the first level involved target localization annotation, using rectangular bounding boxes to precisely locate the koi carp's body and assigning it the category label "fish." Second, the second level involved annotation of key physical features, meticulously marking the coordinates of three key points within the bounding box for each fish: the front of the mouth, the physical midpoint of the side line, and the end of the tail, namely head, midbp, and tail. The completed annotation looks like this. Figure 2 As shown. After annotation is completed, the system will automatically generate a JSON file containing bounding boxes and keypoint coordinates. Finally, to make the results more reliable, the dataset is automatically divided into training and validation sets in an 8:2 ratio using an automated computer script.

[0077] Step S20: Construct a key point awareness network model based on an RTM cascade architecture.

[0078] Because fish undergo non-rigid bending deformation when swimming in water, and the goal of this invention is to perform absolute measurement of their physical length, traditional single-object detection models can only output macroscopic rectangular bounding boxes, failing to obtain the precise anchor points needed to calculate the actual bending length. Furthermore, to meet the engineering requirements of real-time non-contact measurement, this embodiment adopts a top-down, two-tiered cascaded architecture. Specifically, it uses RTMDet as the front-end object detection model and the similarly derived RTMPose-UCA as the back-end keypoint localization model. Its cascaded sensing logic and network architecture are as follows: Figure 4 As shown.

[0079] Specifically, the cascaded model in this embodiment mainly consists of two stages: target detection and keypoint estimation. First, the original underwater image of the fish to be tested is fed into the front-end RTMDet network, whose backbone network (CSPNeXt) extracts multi-scale feature maps from shallow to deep through layer-by-layer downsampling. shallow features Preserving spatial details such as the fish's body edges and scale textures, deep features The receptive field is then expanded to extract global semantic information; subsequently, the features are fused across scales by the neck network of RTMDet and a refined feature map is output. It adapts to the scale of underwater targets, causing drastic fluctuations; ultimately, the head network of RTMDet is based on The system predicts the target's 2D bounding box and uses this to precisely crop the original underwater image of the fish, removing complex water backgrounds and outputting a local image containing only the individual target. This cropped local image is then fed into the backend RTMPose-UCA network, combined with... Figure 4 It can be seen that the RTMPose-UCA network includes a CSPNEXt backbone network, a neck network with embedded UCA modules, and a sub-pixel keypoint prediction head. The RTMPose-UCA uses CSPNEXt as the backbone network for deep feature extraction, performs multi-scale feature fusion through the neck network, and finally outputs the coordinates of three core physical keypoints: the fish mouth, the midpoint of the fish body side line, and the fish tail by the prediction head.

[0080] Furthermore, in actual intensive aquaculture water bodies, not only are there localized high-light interferences caused by mirror reflections from tank sidewalls, water surface reflections, and dynamic ripples, but also severe physical feature confusion due to multi-target crowding and partial occlusion. These high-frequency optical noises and spatial feature interferences easily cause feature extraction to deviate from the true structure, leading to severe coordinate space crosstalk and keypoint localization drift when using traditional attention mechanisms or one-dimensional coordinate pooling. To address the localization drift problem caused by the complex underwater environment, this invention reconstructs the underlying structure of the keypoint perception network. In the neck multi-scale feature fusion stage before the backbone network features of the RTMPose-UCA enter the prediction head, an underwater contextual attention enhancement (UCA) module is introduced.

[0081] The UCA module abandons the conventional dimensionality reduction pooling mechanism that is prone to spatial crosstalk, and achieves high-precision positioning through the collaborative optimization of the following three dimensions:

[0082] (1) Anti-highlight and feature decoupling mechanism: The channel-level random feature dropout mechanism (Dropout2d) is introduced in the front end of the UCA module to effectively break the model’s excessive dependence on water surface reflection or local high-brightness channels, and force the network to mine deeper fish body physical topology features.

[0083] (2) Two-dimensional topology locking and anti-overlap mechanism: The complete two-dimensional geometric topology is preserved in the spatial dimension, and a local spatial convolutional network with a 7×7 large receptive field is used to forcibly block the Cartesian coordinate crosstalk caused by the interference of the target local occlusion features, so as to achieve accurate physical locking of the geometric contour of a single fish.

[0084] (3) Continuous gradient flow and sub-pixel adaptation mechanism: The full-link nonlinear smooth activation function (SiLU) is used in the channel and spatial attention branches to eliminate the hard truncation effect in the traditional attention module and provide a high-quality continuous gradient flow for the sub-pixel classification regression prediction head at the end of the network.

[0085] Through the UCA module, the network constructed in this invention can adaptively suppress optical noise and completely eliminate feature space crosstalk in complex underwater lighting and target occlusion environments, achieving a leap in sub-pixel-level high-precision positioning performance for key points and providing a high-precision coordinate reference for subsequent three-dimensional geometric and physical measurements. The UCA module consists of a pre-feature decoupling submodule, a channel attention submodule, and a two-dimensional spatial context submodule sequentially connected in series. Let the multi-scale feature map received from the CSPNEXt backbone network be... ,in , , These represent the number of channels, height, and width of the feature map, respectively.

[0086] like Figure 5 As shown, the pre-feature decoupling submodule receives multi-scale feature maps. The model is filtered using a channel-level random masking (Dropout2d) component to break its over-reliance on local highlight channels on the water surface; this channel-level random masking component operates according to a set probability. Randomly set some channels of the feature map to zero to output a preliminary feature map with anti-speculative properties. The process is represented as:

[0087]

[0088]

[0089] in, For channel mask vectors that follow a Bernoulli distribution, This indicates an element-wise multiplication operation.

[0090] The channel attention submodule uses the initial screening feature map As input, the Global Average Pooling (GAP) component and the Global Max Pooling (GMP) component are used respectively in the spatial dimension ( Compress features on the channel to generate two channel-level aggregated vectors. :

[0091]

[0092]

[0093] in, These represent the row and column coordinates of the feature map in the spatial dimension, respectively;

[0094] The two sets of vectors are fed into a multilayer perceptron (MLP) with shared parameters to learn the nonlinear dependencies between channels. To maintain a continuous, high-quality gradient flow during backpropagation, the MLP uses a SiLU nonlinear smoothing activation layer instead of the traditional hard-truncation activation. The first fusion unit adds the outputs of the two MLPs element-wise and then processes them through a Sigmoid activation function to generate the channel weight matrix. :

[0095]

[0096] in, The activation function is Sigmoid; the first multiplier will initially screen the feature maps. With channel weight matrix Channel-by-channel multiplication produces an enhanced transition feature map. .

[0097] like Figure 6 As shown, the two-dimensional spatial context submodule affects the transition feature map. Perform average pooling and max pooling operations along the channel dimension to extract two two-dimensional spatial descriptors that retain the complete physical topology. :

[0098]

[0099]

[0100] The first concatenator concatenates the outputs of both along the channel dimension and then feeds them into the first large receptive field convolutional component. The convolutional kernel is used to forcibly lock the macroscopic two-dimensional topology of a single fish; then it passes through a batch normalization (BatchNorm) component and a SiLU smoothing activation layer, and then through a second standard convolutional component ( The convolution kernel and sigmoid activation function are used to process the data, and finally, a two-dimensional spatial weight matrix is ​​extracted. :

[0101]

[0102] in, This indicates a splicing operation along the channel dimension;

[0103] The second multiplier will transition the feature map. With the two-dimensional spatial weight matrix Element-wise multiplication yields a refined, fused feature map that eliminates spatial crosstalk and is robust to mild overlap. .

[0104] Finally, the key point prediction head receives the refined fused feature map, performs sub-pixel level spatial probability distribution regression through the SimCC-based decoding network, and finally outputs high-precision two-dimensional pixel coordinates of the three core anchor points: the front pole of the fish mouth, the physical midpoint of the side line of the fish body, and the end pole of the fish tail.

[0105] Step S30: Training and evaluation of the key point perception network model.

[0106] S31: Loss Function for Keypoint Awareness Network: In real, clear aquaculture environments, surface ripples and sidewall reflections are easily confused with fish edge features, generating a large amount of indistinguishable background noise. This imbalance in the distribution of easy and difficult samples poses a severe challenge to network training, causing the model to pay insufficient attention to difficult samples affected by light and shadow interference during training. Since this invention employs a two-class connected architecture, loss functions need to be constructed separately for the front-end object detection network and the back-end keypoint localization network. In the front-end object detection network, Quality Focal Loss (QFL) is introduced to address the imbalance between easy and difficult samples. Its calculation formula is as follows:

[0107]

[0108] in, Predict the probability score for the front-end object detection network that the candidate region contains the object; A soft label for the intersection-over-union (IoU) ratio between the predicted bounding box and the ground truth bounding box, used to distinguish localization quality; This is a smoothing adjustment constant used to control the rate of weight reduction for simple water body background samples.

[0109] Furthermore, to address the feature point prediction drift problem caused by water reflection, the backend keypoint localization network of this invention employs the KL divergence loss function based on SimCC. This method transforms two-dimensional coordinate regression into a one-dimensional classification task in the horizontal and vertical directions, and its calculation formula is as follows:

[0110]

[0111] in, and These represent one-dimensional Gaussian probability distribution labels generated from real keypoints in the horizontal and vertical dimensions, respectively. and These represent the probability distributions of the network's predicted output in the corresponding dimensions.

[0112] The combined application of the aforementioned loss functions addresses the problem of underwater sample imbalance by using QFL (Quick Light Function) to assign high weights to difficult samples, making the model focus more on fish areas affected by reflective interference during training. Furthermore, combining KL divergence with one-dimensional coordinate classification smooths out differences in feature point probability distributions and avoids the quantization errors of traditional heatmaps. This prevents drastic fluctuations in loss values ​​caused by individual extreme reflective points, thus helping the cascaded model converge more stably.

[0113] S32: Training and Evaluation of Key Point Awareness Network Model

[0114] To ensure the accuracy and reliability of the experimental data, the training and testing of the underwater key point sensing network in this embodiment were both conducted under the same standard hardware and software environment. The experimental environment is shown in Table 1.

[0115] Table 1 Experimental Environment

[0116]

[0117] Experimental hyperparameter settings: Using an underwater image dataset, the batch size for both the front-end and back-end networks was set to 4. To accommodate the characteristics of the two-class connected architecture, a phased independent training strategy was adopted: the front-end object detection network had 150 epochs with an initial learning rate of 0.00003125; the back-end keypoint localization network had 400 epochs with an initial learning rate of 0.0002. AdamW was used as the optimizer for both networks.

[0118] In terms of the recognition accuracy of the front-end object detection network, precision rate (P), recall rate (R), and mean average precision (mAP@0.5) were selected as evaluation indicators.

[0119] Regarding the accuracy of backend keypoint localization, since traditional detection metrics cannot measure the pixel deviation of spatial points, this embodiment additionally introduces the Percentage of Correct Keypoints (PCK) as a core evaluation metric, and sets the distance tolerance threshold to 0.2. The calculation formula is as follows:

[0120]

[0121] in, Represents the total number of key points; Representing the The Euclidean distance between each predicted keypoint and the actual manually labeled keypoint; This represents the set distance tolerance threshold; This is an indicator function that returns 1 if the condition within the parentheses is met, and 0 otherwise. The higher the value, the more accurate the sub-pixel coordinates output by the backend key point localization model in complex water bodies, and the stronger its anti-drift capability.

[0122] S33: Experimental Results and Analysis

[0123] S331: Ablation Experiment of Underwater Contextual Attention Enhancement Module (UCA)

[0124] To verify the effectiveness of the various sub-mechanisms within the Underwater Contextual Attention (UCA) module in complex underwater environments, this experiment used the standard RTMPose network as a baseline. Under uniform conditions of 400 rounds of maximum computational budget and strong occlusion (CoarseDropout) data augmentation, ablation tests were conducted on the core components. Specific comparison configurations included: a network model incorporating traditional 1D coordinate pooling (CoordAtt), a UCA model stripped of the high-spectrum channel-level random mask (Dropout2d), and a UCA model replacing the nonlinear smooth activation function (SiLU) with hard truncation (ReLU). Experimental results are shown in Table 2.

[0125] Table 2 Results of the internal ablation experiment of the UCA attention module

[0126]

[0127] Due to the crowded nature of underwater multi-target environments and the susceptibility to local feature interference, the baseline model achieved only 59.1% accuracy on the stringent high-precision localization metric (AP@0.75). While introducing traditional 1D coordinate dimensionality reduction pooling (CoordAtt) improved the overall localization capability on a more lenient metric, spatial crosstalk risks due to X / Y axis feature decoupling still existed. In contrast, the complete UCA module of this invention, by introducing a 7×7 large receptive field 2D convolution to forcibly lock the 2D geometric topology of a single fish, successfully blocked crosstalk misalignment of Cartesian spatial features, achieving a significant leap in high-precision localization capability.

[0128] Meanwhile, ablation comparison data of the internal sub-mechanisms further confirms the scientific nature of the module design: stripping the Dropout2d mechanism leads to an over-reliance on local water surface reflections and specular features, weakening the model's generalization robustness; while the model that replaces the SiLU activation function with ReLU hard truncation disrupts the continuous gradient flow in backpropagation. The complete UCA module of this invention, through the synergistic coupling of anti-spectral masking and the full-link SiLU activation function, provides high-quality smooth gradient features for the backend SimCC-based sub-pixel classification head, ultimately enabling the model of this invention to achieve optimal performance on comprehensive evaluation metrics such as PCK and mAP.

[0129] S332: Cross-sectional Comparison Experiment of Backend Key Point Localization Backbone Network

[0130] To verify the performance advantages and practical deployment feasibility of the model of this invention in underwater non-rigid target feature extraction tasks, this experiment fixed the front-end target detection network as RTMDet. In the back-end keypoint localization network of this invention, CSPNEXt is used as the feature extraction backbone network, and an underwater contextual attention enhancement (UCA) module is embedded in its neck network to construct the RTMPose-UCA back-end keypoint localization network.

[0131] To conduct a cross-sectional performance evaluation, the classic serial network ResNet-50 and the high-resolution parallel network HRNet-w32 were selected as the baselines for comparison. The evaluation system takes into account both the prediction accuracy and computational cost of the models, and the specific evaluation metrics include PCK, mAP, mAR, as well as the number of model parameters (params) and computational cost (FLOPs). The comparative experimental results are shown in Table 3.

[0132] Table 3 Comparison Experiment of Key Point Perception Models

[0133]

[0134] Existing conventional networks (such as HRNet-w32) typically rely on a multi-branch parallel architecture to maintain high-resolution feature representations, resulting in high computational overhead (up to 15.368G) and limited performance on the stringent sub-pixel localization metric (AP@0.75), achieving only 0.633. In contrast, the model proposed in this invention achieves a better balance between accuracy and model complexity. With the number of parameters controlled at 24.51M and the computational cost reduced to 5.983G (only 38.9% of HRNet-w32), this model achieves improvements in all accuracy metrics. Specifically, the model's PCK metric reaches 0.998, and the overall localization accuracy (mAP) and recall (mAR) are improved to 0.594 and 0.650, respectively. These results demonstrate that this invention, utilizing the topology awareness and locking mechanism of the UCA module, effectively improves target localization accuracy in complex underwater scenarios while significantly reducing model computational complexity, and can better meet the engineering deployment requirements for real-time, high-precision computation of underwater edge devices.

[0135] Step S40: Fish body length calculation based on spatial geometry and temporal filtering

[0136] After obtaining high-precision pixel coordinates of two-dimensional key points of a non-rigid target, this invention combines the spatial intrinsic parameters of a binocular vision / depth camera with three-dimensional orthogonal mapping, geometric arc compensation, and two-order temporal digital signal filtering to calculate the true physical length of the target individual. Specific steps include:

[0137] S41: Local noise-resistant sampling and 3D spatial mapping based on depth map

[0138] To overcome underwater speckle noise and the "depth penetration" phenomenon caused by highly permeable water (i.e., the depth camera misinterpreting the background wall as the fish surface), this invention does not employ direct single-point depth mapping, but instead constructs a robust depth extraction mechanism. Specifically, it uses the two-dimensional pixel coordinates of three key points: the fish mouth, the midpoint of the lateral line, and the tail peduncle. Centered on, extract respectively The local neighborhood pixels (ROI) are analyzed, and invalid depth values ​​(non-positive or infinite) are filtered out. All remaining valid depth values ​​are aggregated into a local depth sample set corresponding to the keypoint. This local depth sample set is then arranged in ascending order, and only the depth value samples in the first half of the sorted set (i.e., the physical surface closest to the lens) are extracted. The median of these median values ​​is then used as the final robust depth value of the target in that frame. Subsequently, the camera intrinsic parameter matrix was combined. Using the following pinhole camera inverse perspective formula, the two-dimensional pixel coordinates of three physical key points are obtained. Mapped to three-dimensional Cartesian coordinates :

[0139]

[0140]

[0141]

[0142] The camera intrinsic parameter matrix is ​​expressed as follows:

[0143]

[0144] In the formula, This is the equivalent focal length of the camera. These are the coordinates of the camera's optical center.

[0145] Thus, the fish mouth was obtained. Midpoint of the side line Fish tail The three-dimensional physical coordinates.

[0146] S42: Establish a global field of view and cold start gating mechanism to perform pre-legality screening on the continuously acquired video image data stream of the target fish. Only when the center of the local image containing the target individual output by the front-end target detection network is within the safe field of view and the average confidence of the key points is greater than a preset threshold, is it allowed to extract the three-dimensional spatial physical coordinates of the three key points corresponding to the target individual in the underwater image of the target fish. The average confidence level of the key points The calculation formula is:

[0147]

[0148] in, The confidence level for the three key points;

[0149] The admission criteria for the cold start phase can then be expressed as:

[0150]

[0151] in, This is the preset high sensitivity threshold.

[0152] When a target object undergoes swimming deformation in water, simple Euclidean linear distance will produce a severe physical truncation error. Therefore, for valid admission frames, this invention establishes a three-dimensional spatial geometric compensation engine, whose spatial orthogonal circular arc compensation geometric model is as follows: Figure 7 As shown.

[0153] First, calculate the spatial linear Euclidean distance between adjacent key points: the line segment from the fish mouth to the midpoint. The line segment from the midpoint to the fish tail The straight section from the fish's mouth to its tail .

[0154] Traditional broken line approximation for length measurement is To accurately represent the physical arc length, this invention utilizes Heron's formula to calculate the area of ​​a spatial triangle formed by three points. :

[0155]

[0156] Among them, half perimeter The calculation formula is:

[0157]

[0158] The radius of curvature of the circumcircle of this spatial triangle is further derived. :

[0159]

[0160] If the radius of curvature If the fish is within a reasonable threshold range (determined to be curved rather than perfectly straight), then calculate the corresponding central angle. Compared with the theoretical arc length :

[0161]

[0162]

[0163] To prevent feature point coordinate drift from incorrectly fitting an overly curved closed loop, and to avoid non-physical stretching exceeding the biomechanical limits of the fish's spine, based on extensive underwater measurement statistical priors, the following settings are established. The polyline distance is a hard, empirical upper limit threshold for dynamic deformation. When the calculated theoretical arc crosses this threshold, the system triggers a singularity degradation mechanism, stripping away unreliable curvature expansion and forcing a return to the absolutely safe baseline polyline distance; this is used to extract safe original measurement values. Its piecewise mathematical constraint expression is:

[0164]

[0165] S43: Two-order timing signal filtering mechanism

[0166] In acquiring the raw measurement values ​​of consecutive valid admission frames Following the sequence, to eliminate optical distortion at the lens edges and underwater dynamic high-frequency noise, this invention constructs a two-order temporal signal filtering mechanism based on field-of-view constraints. The specific algorithm processing logic is as follows: Figure 8As shown. Specifically, a first-order extreme value fusing mechanism is implemented for the continuous data stream. This mechanism maintains a sliding window historical data queue in memory for a specific time span. When new measurement data is input for each frame, the system calculates the current median value of the historical queue in real time. If the physical measurement value of the current frame experiences a drastic abnormal jump, i.e., exceeds a specific multiple threshold of the historical median, the system determines that the data is a "depth breakdown" distortion value caused by underwater speckle noise or complete occlusion of the target. At this time, the circuit breaker protection is triggered, the system rejects the abnormal data from entering the sliding window, and maintains the safe and valid value of the previous frame as the current output.

[0167] For valid measurement data that successfully penetrates the first-level circuit breaker defense, the system immediately triggers the second-level truncated mean fusion mechanism. In this processing phase, the system sorts the valid historical data sequence accumulated within the current sliding window in ascending order and, according to a preset truncation ratio, precisely removes the 25% of extreme distribution data at the highest and lowest points of the queue. Finally, the system calculates the arithmetic mean of the remaining central 50% core data pool after removing marginal extremes and outputs the physical length to be corrected. This two-stage serial filtering architecture not only completely filters out spatial coordinate collapse caused by hardware sensors, but also effectively smooths geometric fluctuations within a reasonable range caused by the high-frequency movement of the target individual, ultimately outputting a length measurement sequence with industrial-grade stability.

[0168] Finally, the curvature gain coefficient is introduced. And combined with the physical compensation coefficient of water refractive index (In this embodiment, 0.95 is used) to calculate the final physical measurement value for a single frame. :

[0169]

[0170]

[0171] S431: Dynamic Length Measurement Steady-State Analysis Based on Spatial Geometric Compensation and Two-Order Temporal Filtering

[0172] To further verify the accuracy and noise resistance of the present invention in three-dimensional physical measurement under continuous time sequence, an ablation comparison experiment was carried out on the dynamic continuous length measurement task of underwater non-rigid targets based on the key point coordinates extracted by the aforementioned RTMPose-UCA model.

[0173] Table 4 Comparison Experiment of Dynamic Length Measurement and Two-Order Temporal Filtering Mechanism Ablation

[0174]

[0175] Table 4 presents a comparison of the system noise immunity and measurement accuracy of different length measurement strategies under complex underwater environments (including multi-target spatial crosstalk and individual non-rigid deformation). This experiment used a target individual with an absolute physical truth value of 15.5 cm as the spatial tracking anchor point, accurately capturing its continuous dynamic swimming sequence from severe torso bending to complete flattening. The experiment employed the anomalous jump rate (AJR) and mean relative error (ARE) as core quantitative evaluation indicators.

[0176] As shown in Table 4, comparing the experimental data, this invention verifies the synergistic effect of two-order temporal filtering and spatial orthogonal circular arc compensation mechanism in complex underwater environments. In this high-frequency continuous sampling, traditional straight-line ranging inevitably produces severe geometric truncation errors when the target undergoes non-rigid deformation (torso bending), resulting in an average relative error (ARE) as high as 33.4%. Furthermore, when simply introducing the midpoint of the lateral line for single-frame curve ranging, the complex underwater lighting, water disturbances, and high-frequency target twisting easily induce local point cloud distortion in the 3D vision sensor (i.e., depth penetration), causing the measurement results to be extremely sensitive to noise. The abnormal jump rate (AJR) surges to 25.1%, and the ARE reaches 34.2%. This demonstrates that simply introducing spatial dimension geometric compensation lacks effective measurement robustness in real underwater high-noise environments.

[0177] In contrast, the dual-order filtering length measurement system constructed in this invention, while effectively restoring the true physical arc length of the target by utilizing ROI anti-distortion and orthogonal circular arc compensation at the front end, innovatively introduces a time-series digital signal processing mechanism at the back end. Specifically, the first-order "extreme value throttling" mechanism accurately intercepts and eliminates spatial coordinate collapse data caused by depth penetration; the second-order "sliding truncated mean" further smooths fluctuation noise within a reasonable range caused by high-frequency movement. Final test results show that this invention significantly suppresses the abnormal jump rate to 2.9%, and the average relative error significantly converges to 6.8%. In summary, the technical solution of this invention completely overcomes the problem of dynamic length measurement distortion caused by the superposition of continuous non-rigid deformation of the target and optical speckle noise in complex underwater environments, achieving high robustness and industrial-grade measurement accuracy.

[0178] This invention proposes a keypoint perception network model based on a two-stage concatenated architecture and a two-stage temporal filtering method for length measurement of non-rigid targets in complex underwater environments. This aims to address challenges such as crosstalk from multiple underwater targets, optical interference, and abrupt changes in depth data. By introducing an underwater contextual attention enhancement (UCA) module at the feature decoding front end, and leveraging the synergistic effect of two-dimensional large receptive field convolution and anti-spectral feature masks, the Cartesian space crosstalk problem caused by traditional one-dimensional dimensionality reduction is successfully blocked, significantly improving the model's ability to perceive the fish's topology. Combined with orthogonal circular arc geometric compensation and a two-stage temporal statistical filtering mechanism based on extreme value melting and truncated mean, dynamic measurement jitter caused by high-frequency swimming deformation and depth map speckle noise is further eliminated. In model ablation experiments, the complete UCA module enabled the network to achieve 68.6% accuracy on the stringent sub-pixel localization index AP@0.75, and the overall localization PCK accuracy was improved to 0.998. Compared with current mainstream sensing models (such as ResNet-50 and HRNet-w32), this invention achieves a better balance between computational efficiency and accuracy under lightweight conditions with only 24.51M parameters and a computational load reduced to 5.983G. In continuous spatiotemporal sequence length measurement verification, this invention significantly reduces the anomalous jump rate (AJR) to 2.9% and the average relative error (ARE) to 6.8%. These results demonstrate that this invention significantly lowers the computational power threshold of end-side devices while constructing a highly robust physical measurement benchmark, providing reliable algorithmic support for industrial-grade underwater high-frequency dynamic measurement.

[0179] Based on the above embodiments, this application also provides a computer program that, when run on a computer, causes the computer to execute the methods provided in the above embodiments.

[0180] Based on the above embodiments, this application also provides a computer storage medium storing a computer program, which, when executed by a computer, causes the computer to perform the methods provided in the above embodiments.

[0181] The storage medium can be any available medium that a computer can access. For example, but not limited to, a computer-readable medium can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer.

[0182] Finally, it should be noted that the above embodiments are intended to illustrate the technical solutions of the present invention and do not constitute any limitation on the present invention. Those skilled in the art should fully understand that modifications to the technical solutions described in the foregoing embodiments or equivalent substitutions for any part or all of the technical features are entirely feasible. Such modifications or substitutions, as long as they do not depart from the scope of protection defined by the claims of the present invention, should be considered reasonable extensions of the present invention.

Claims

1. A method for measuring the length of underwater fish based on binocular vision and attitude estimation, characterized in that, include: S1: Acquire several images of the fish to be tested in real underwater scenes and the corresponding 3D point cloud depth map matrix; S2: Input all underwater fish images to be tested into the key point perception network model to obtain the two-dimensional pixel coordinates of three core physical key points of all target individuals in each underwater fish image: the front end of the fish mouth, the physical midpoint of the fish body side line, and the end of the fish tail; the key point perception network model is a two-level hierarchical architecture composed of a front-end target detection network and a back-end key point localization network. The front-end target detection network extracts the target two-dimensional bounding box of the underwater fish image and obtains a local image containing a single target individual based on the target two-dimensional bounding box; the back-end key point localization network obtains the two-dimensional pixel coordinates of three core physical key points: the front end of the fish mouth, the physical midpoint of the fish body side line, and the end of the fish tail based on the local image. S3: Based on the two-dimensional pixel coordinates of three physical key points of all target individuals in the underwater fish images and the corresponding three-dimensional point cloud depth map matrix, the true physical length of the fish to be tested is calculated using spatial orthogonal circular arc compensation and two-order temporal signal filtering mechanism.

2. The underwater fish length measurement method based on binocular vision and attitude estimation according to claim 1, characterized in that, The front-end object detection network acquires local images containing a single object based on the following method: The front-end target detection network takes the underwater fish image as input and extracts multi-scale feature maps from shallow to deep through layer-by-layer downsampling. The multi-scale feature maps are fused across scales and output corresponding refined feature maps. Based on the refined feature maps, the target two-dimensional bounding box is predicted, and the original underwater fish image is precisely spatially cropped according to the target two-dimensional bounding box to remove complex water backgrounds and output a local image containing only a single target individual.

3. The underwater fish length measurement method based on binocular vision and attitude estimation according to claim 1 or 2, characterized in that, The process of acquiring a local image containing a single target individual by the front-end target detection network is implemented using the RTMDet model.

4. The underwater fish length measurement method based on binocular vision and attitude estimation according to claim 1, characterized in that, The backend key point localization network obtains the two-dimensional pixel coordinates of three core physical key points: the front end of the fish mouth, the physical midpoint of the fish body lateral line, and the end of the fish tail, based on the following method: The backend keypoint localization network consists of a feature extraction network, a neck network embedded with a UCA module, and a sub-pixel keypoint prediction head connected in series. The feature extraction network, as the backbone network, performs deep convolution extraction on the local image of the input single target individual to obtain a multi-scale feature map containing the target's physical topology. The neck network receives this multi-scale feature map and performs anti-spectral feature decoupling and 2D topology locking processing through the UCA module. After effectively suppressing feature drift caused by water surface reflection and target local occlusion, a refined fusion feature map with spatial crosstalk eliminated is obtained. The sub-pixel keypoint prediction head takes the refined fusion feature map as input and uses a SimCC-based decoding network to perform sub-pixel-level spatial probability distribution regression on it, accurately outputting the 2D pixel coordinates of three core physical key points: the front of the fish mouth, the physical midpoint of the fish's side line, and the end of the fish tail.

5. The underwater fish length measurement method based on binocular vision and attitude estimation according to claim 4, characterized in that, The UCA module includes a sequentially connected pre-feature decoupling submodule, a channel attention submodule, and a two-dimensional spatial context submodule; The pre-feature decoupling submodule receives the multi-scale feature map output by the feature extraction network, filters it through the channel-level random mask component, and outputs a preliminary screening feature map that is resistant to specular light. The channel attention submodule takes the initial screening feature map as input and uses a global average pooling component and a global max pooling component to compress the features in the spatial dimension, outputting two spatial aggregation vectors. The two sets of vectors are then fed into a shared parameter MLP to learn the nonlinear dependencies between channels. The MLP uses a SiLU nonlinear smooth activation layer instead of the traditional hard truncation activation. The first fusion unit adds the outputs of the two MLPs element by element and then processes them through an activation function to generate a channel weight matrix. The first multiplier multiplies the initial screening feature map with the channel weight matrix channel by channel, outputting a channel-enhanced transition feature map. The two-dimensional spatial context submodule performs average pooling and max pooling operations on the transition feature map along the channel dimension, respectively; the first stitcher stitches the outputs of the two along the channel dimension and feeds them into the first large receptive field convolutional component to forcibly lock the macroscopic two-dimensional topology of a single fish; then it passes through the batch normalization component and the SiLU smoothing activation layer in sequence, and then through the second standard convolutional component and activation function to extract the two-dimensional spatial weight matrix; the second multiplier multiplies the transition feature map and the two-dimensional spatial weight matrix element by element, and finally outputs a refined fusion feature map that eliminates spatial crosstalk and is robust to slight overlap.

6. The underwater fish length measurement method based on binocular vision and attitude estimation according to claim 4 or 5, characterized in that, The feature extraction network includes CSPNEXt network, ResNet50, or HRNet-w32.

7. The underwater fish length measurement method based on binocular vision and attitude estimation according to claim 1, characterized in that, In step S3, the calculation of the true physical total length of the fish under test using spatial orthogonal circular arc compensation and a two-order temporal signal filtering mechanism specifically includes: S3.1: The two-dimensional pixel coordinates of the three physical key points of each target individual in the underwater fish image are fused with the corresponding three-dimensional point cloud depth map matrix. A 5×5 local neighborhood pixel block is extracted centered on each key point. After filtering out invalid depth values, the remaining valid depth values ​​are arranged in ascending order to construct a local depth sample set for the corresponding key point. Only the first half of the depth value samples in the local depth sample set is extracted, and the median is calculated as the robust depth value for the corresponding key point. Subsequently, combined with the camera intrinsic parameter matrix, the two-dimensional pixel coordinates of the three physical key points are fused using the reverse perspective principle of a pinhole camera. Mapped to three-dimensional physical coordinates respectively , ; S3.2: Establish a global field of view and cold start gating mechanism to perform pre-legality screening on the continuously acquired video image data stream of the target fish. Only when the center of the local image containing the target individual output by the front-end target detection network is within the safe field of view and the average confidence of the key points is met, the system will ensure that the data is valid. Greater than the preset threshold Only then is it permitted to extract the three-dimensional spatial physical coordinates of the three key points corresponding to the target individual in the underwater image of the fish under test. The entry requirements for the cold start phase are: The average confidence level of the key points The calculation formula is: in, The confidence level for the three key points; For a valid admission frame, calculate the spatial straight-line distance between adjacent keypoints to construct a spatial triangle, and define the basic polyline distance. The area and radius of curvature of the circumcircle of the triangle are calculated based on Heron's formula. and combined with the central angle Calculation of theoretical arc unfolding length ; To prevent feature point coordinate drift from incorrectly fitting an overly curved closed loop, and to avoid non-physical stretching exceeding the biomechanical limits of the fish's spine, the following settings are adopted: The polyline distance is a hard, empirical upper limit threshold for dynamic deformation. When the calculated theoretical arc crosses this threshold, the system triggers a singularity degradation mechanism, stripping away unreliable curvature expansion and forcing a return to the absolutely safe baseline polyline distance; this is used to extract safe original measurement values. Its piecewise mathematical constraint expression is: S3.3: Raw measurement values ​​for continuous valid admission frames output The sequence executes a two-stage temporal signal filtering mechanism to eliminate spatial high-frequency jitter; the first stage involves real-time calculation of the current single-frame raw measurement value. Sliding window with system memory history All The absolute deviation of the median; if the deviation is abnormal, the extreme value circuit breaker is triggered to intercept the deep breakdown noise and force the output to maintain the measurement value of the previous frame. in, The sample size of the historical sliding window is the deviation rate. Select 90 frames per second; The second level is for circuit breakers that have not been triggered. The queue is sorted, and the top and bottom 25% of extreme outliers are removed. Data, and 50% of the measurements are retained at the center. Calculate the arithmetic mean of the sample set and output the physical total length to be corrected. : in, This represents the total number of samples in the set. Furthermore, a dynamic curvature gain coefficient based on the central angle is introduced. Compensation coefficient for static refraction of water body right System-level corrections are performed, and the actual physical measurement values ​​of each target fish in a single frame of underwater fish image are continuously calculated and output. .

8. The underwater fish length measurement method based on binocular vision and attitude estimation according to claim 1, characterized in that, The construction process of the key point perception network model includes: Continuous video recording of the target koi in a real water tank was performed to obtain an SVO format video stream, and the video stream was then subjected to equal-interval frame extraction. Labelme was used for two-level customized annotation. The first level was for target localization annotation, using rectangular bounding boxes to accurately locate the koi carp's body and assigning it the category label "fish". The second level was for physical feature keypoint annotation, strictly marking the coordinates of three key points within the bounding box of each fish: the front of the mouth, the physical midpoint of the side line of the body, and the end of the tail, labeled as head, midbp, and tail, respectively. After annotation, the system initially generated a JSON file containing the bounding box and keypoint coordinates. An automated script then converted the JSON file into the standard COCO dataset format to obtain the underwater multimodal visual dataset. The keypoint perception network model is trained using the underwater multimodal visual dataset to obtain a trained keypoint perception network model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor; characterized in that, When the processor executes the computer program, it causes the electronic device to perform the underwater fish length measurement method based on binocular vision and attitude estimation as described in any one of claims 1 to 8.

10. A storage medium comprising a computer program, characterized in that, When the computer program is run on an electronic device, the electronic device performs the underwater fish length measurement method based on binocular vision and attitude estimation as described in any one of claims 1 to 8.