Method, system and program product for measuring fish length

CN122820802APending Publication Date: 2026-09-25SOUTHERN MARINE SCIENCE & ENGINEERING GUANGDONG LABORATORY (ZHANJIANG)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611136294.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]本申请的主要目的在于提供一种鱼体长测量方法、系统及程序产品,旨在解决如何在完整计算帧内重新组织任务调度,使异构资源高效协同工作,以降低单帧处理延迟的技术问题

Benefits of technology

本申请的技术方案首先通过获取包含左目图像和右目图像的双目视频流,为后续的立体视觉处理提供了原始图像数据基础;然后通过对左目图像和右目图像进行校正处理,使左右图像满足极线约束条件,消除了镜头畸变和双目安装误差带来的影响,从而为后续的立体匹配和深度计算提供了空间对齐的图像数据,保证了视差计算的准确性;接着通过多线程执行步骤,在一完整计算帧内并行执行视差深度计算和鱼体检测与实例分割推理两个子任务,视差深度计算在第一线程中由第一类处理器执行,鱼体检测与实例分割推理在第二线程中由第二类处理器执行,两个线程并行运行,使得原本串行执行时依次占用计算资源的两个耗时环节转变为在同一时间段内重叠执行,从而将完整计算帧中两个主要耗时环节对整体处理时间的贡献从两者之和优化为两者中的较大值,显著降低了单帧处理的墙钟时间,提高了边缘计算平台上异构计算资源的协同利用效率;最后根据鱼体实例分割掩膜和深度图计算鱼体长度,通过将语义分割前景与双目深度信息相结合,基于鱼体实例分割掩膜约束深度采样范围,减少了背景区域和遮挡区域对视差采样的干扰,从而在三维空间中实现了鱼体长度的准确测量,克服了单目二维图像测量缺乏深度信息的缺陷。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820802A_ABST
    Figure CN122820802A_ABST
Patent Text Reader

Abstract

The application discloses a fish body length measurement method, system and program product, and relates to the technical field of biological measurement. The method comprises the following steps: acquiring a binocular video stream; the binocular video stream comprises a plurality of left-eye images and a plurality of right-eye images; performing correction processing on the left-eye images and the right-eye images to obtain left-eye corrected images and right-eye corrected images; in a complete calculation frame, performing parallax depth calculation on the left-eye corrected images and the right-eye corrected images in a first thread to generate a depth map, and performing fish body detection and instance segmentation inference on the left-eye corrected images in a second thread to obtain a fish body instance segmentation mask; the first thread and the second thread are executed in parallel, and the first thread and the second thread are executed by different types of processors; and the fish body length is calculated according to the fish body instance segmentation mask and the depth map. According to the application, the binocular parallax depth calculation and the instance segmentation inference are executed in parallel in a complete calculation frame, so that heterogeneous resources can be efficiently cooperated, and the single-frame processing delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biometrics, and in particular to methods, systems and programs for measuring fish body length. Background Technology

[0002] Binocular vision measurement technology is widely used in aquaculture, marine ranching, and other scenarios. It recovers target depth information through the parallax of left and right cameras and combines it with target detection, instance segmentation, or contour extraction to measure fish length. With the development of edge computing technology, related measurement algorithms are gradually migrating to embedded platforms in order to achieve low-cost, low-power field deployment.

[0003] Currently, a typical binocular fish length measurement system usually includes multiple stages such as image correction, neural network detection and segmentation, binocular disparity calculation, depth conversion, and length post-processing. In practical deployments, object detection and segmentation tasks mainly rely on accelerated resources such as graphics processing units (GPUs) or neural processing units (NPUs) for inference, while disparity calculation mainly consumes CPU and memory bandwidth. Current systems typically execute these stages sequentially, either by completing neural network inference first, followed by disparity calculation and depth generation, or by executing them in reverse order.

[0004] However, this serial execution method results in a large processing delay per frame, making it difficult to meet real-time requirements. At the same time, although domestic edge computing platforms, represented by the RK3588, possess heterogeneous computing resources such as CPU, NPU, and RGA (Raster Graphics Acceleration), the traditional serial process cannot fully utilize their collaborative processing capabilities, resulting in a waste of computing resources.

[0005] Therefore, how to reorganize task scheduling within a complete computing frame to enable heterogeneous resources to work together efficiently and reduce single-frame processing latency has become an urgent technical problem to be solved.

[0006] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0007] The main purpose of this application is to provide a method, system and program product for measuring fish body length, which aims to solve the technical problem of how to reorganize task scheduling within a complete computing frame, so as to enable heterogeneous resources to work together efficiently and reduce the processing latency of a single frame.

[0008] To achieve the above objectives, this application proposes a method for measuring the body length of a fish, the method comprising: Acquire a binocular video stream; the binocular video stream includes multiple frames of left-eye images and multiple frames of right-eye images; The left and right eye images are corrected to obtain a corrected left eye image and a corrected right eye image. The multi-threaded execution steps are as follows: within a complete calculation frame, in the first thread, disparity depth is calculated based on the left eye correction image and the right eye correction image to generate a depth map; and in the second thread, fish body detection and instance segmentation inference are performed on the left eye correction image to obtain a fish body instance segmentation mask. The first thread and the second thread are executed in parallel, with the first thread executed by the first type of processor and the second thread executed by the second type of processor. The length of the fish is calculated based on the segmentation mask of the fish instance and the depth map.

[0009] In one embodiment, the step of calculating disparity depth based on the left-eye corrected image and the right-eye corrected image in a first thread to generate a depth map includes: The left and right eye corrected images are subjected to grayscale conversion and texture enhancement processing to obtain a left eye grayscale image and a right eye grayscale image; The left-eye grayscale image and the right-eye grayscale image are scaled according to a preset scaling factor s to obtain a scaled left-eye grayscale image and a scaled right-eye grayscale image; where 0 <s≤1; Calculate a small-scale disparity map based on the zoomed grayscale image of the left eye and the zoomed grayscale image of the right eye; The small-scale disparity map is filtered, and the filtered small-scale disparity map is restored to the full resolution size to obtain the restored small-scale disparity map. Based on the scale factor s, the recovered small-scale disparity map is converted into a full-resolution equivalent disparity map; The depth map is calculated based on the full-resolution equivalent disparity map.

[0010] In one embodiment, the step of converting the recovered small-scale disparity map into a full-resolution equivalent disparity map based on the scale factor s includes: Based on the disparity values ​​of each pixel in the restored small-scale disparity map and the scale factor, calculate the full-resolution equivalent disparity value of each pixel. Correspondingly, the step of calculating the depth map based on the full-resolution equivalent disparity map includes: The depth value of each pixel is calculated based on the full-resolution equivalent parallax value of each pixel, the horizontal focal length of the left eye camera, and the binocular baseline, to obtain the depth map.

[0011] In one embodiment, the method further includes: Obtain the frequency division parameters; Based on the current frame number and the frequency division parameter, determine whether the current frame is a complete calculation frame; If the current frame is the complete calculation frame, then the multi-threaded execution step is executed, and the calculation result is updated to the cached result; If the current frame is not the complete calculation frame, the cached result is reused for display, and the writing of measurement records to the structured result file is abandoned.

[0012] In one embodiment, the step of calculating the fish length based on the fish instance segmentation mask and the depth map includes: Construct the target depth region of interest based on the segmentation mask of the fish body instance; Perform maximum connectivity preservation and hole filling processing on the fish body instance segmentation mask to obtain the target foreground mask; Principal component analysis was performed on the target foreground mask to extract the fish body principal axis and the two endpoints of the fish body principal axis; Calculate the pixel length based on the two endpoints; Based on the depth map, obtain the depth values ​​corresponding to the two endpoints respectively; The three-dimensional length of the fish is calculated based on the two endpoints, the corresponding depth values, and the intrinsic parameters of the left eye camera.

[0013] In one embodiment, the step of obtaining the depth values ​​corresponding to the two endpoints respectively includes: Determine the center point of the main axis of the fish body; According to a preset shrinkage ratio, each endpoint is shrunk inward toward the center point along the main axis to obtain the corresponding sampling endpoint; Median depth sampling is performed in the neighborhood of each sampling endpoint to obtain the depth value corresponding to each endpoint.

[0014] In one embodiment, the method further includes: The validity of the fish body length is determined; Based on the validity determination result, the corresponding length result is output and the corresponding measurement mode is marked, or the invalid reason is output; wherein, the measurement mode includes three-dimensional mode, approximate mode and stable mode.

[0015] In one embodiment, the first type of processor is a central processing unit; and / or, the disparity depth calculation adopts a semi-global matching algorithm; and / or, the second type of processor is a neural network processor; and / or, the fish detection and instance segmentation inference adopt an instance segmentation model running under the neural network processor framework.

[0016] Furthermore, to achieve the above objectives, this application also proposes a fish body length measurement system, which includes: The image acquisition module is used to acquire the left and right eye images captured by the binocular camera; An image correction module is used to correct the left eye image and the right eye image to obtain a corrected left eye image and a corrected right eye image; A multi-threaded execution module is used to perform disparity depth calculation based on the left-eye corrected image and the right-eye corrected image in a first thread to generate a depth map, and to perform fish body detection and instance segmentation inference on the left-eye corrected image in a second thread to obtain a fish body instance segmentation mask; wherein, the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor; The fish body length calculation module is used to calculate the fish body length based on the segmentation mask of the fish body instance and the depth map.

[0017] In addition, to achieve the above objectives, this application also proposes a fish body length measuring device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the fish body length measuring method described above.

[0018] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the fish body length measurement method described above.

[0019] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the fish body length measurement method described above.

[0020] One or more technical solutions proposed in this application have at least the following technical effects: The technical solution of this application first acquires a binocular video stream containing left and right eye images, providing the raw image data foundation for subsequent stereo vision processing. Then, by correcting the left and right eye images, the left and right images are made to satisfy epipolar constraints, eliminating the effects of lens distortion and binocular installation errors. This provides spatially aligned image data for subsequent stereo matching and depth calculation, ensuring the accuracy of disparity calculation. Next, through multi-threaded execution, two subtasks—disparity depth calculation and fish detection / instance segmentation inference—are executed in parallel within a complete calculation frame. Disparity depth calculation is performed by a first-type processor in the first thread, while fish detection / instance segmentation inference is performed by a second-type processor in the second thread. Parallel execution transforms the two time-consuming steps that previously consumed computational resources sequentially into overlapping executions within the same timeframe. This optimizes the contribution of the two main time-consuming steps in a complete computation frame to the larger of the two, significantly reducing the processing time per frame and improving the collaborative utilization efficiency of heterogeneous computing resources on the edge computing platform. Finally, the fish length is calculated based on the fish instance segmentation mask and depth map. By combining semantic segmentation foreground with binocular depth information and constraining the depth sampling range based on the fish instance segmentation mask, the interference of background and occlusion regions on disparity sampling is reduced, thus achieving accurate measurement of fish length in three-dimensional space and overcoming the lack of depth information in monocular two-dimensional image measurement. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart illustrating an embodiment of the fish body length measurement method of this application. Figure 2 This is a schematic diagram of the overall architecture of the fish body length measurement method of this application, provided in Embodiment 1. Figure 3 A schematic diagram comparing the traditional serial process and the intra-frame parallel process of this application for the fish body length measurement method embodiment 1; Figure 4 This is a flowchart illustrating Embodiment 2 of the fish body length measurement method of this application. Figure 5This is a flowchart illustrating Embodiment 3 of the fish body length measurement method of this application; Figure 6 This is a flowchart illustrating Embodiment 4 of the fish body length measurement method of this application; Figure 7 This is a flowchart illustrating the calculation and validity determination of fish body length in one embodiment of the fish body length measurement method of this application. Figure 8 A comparison chart of the average single-frame execution time of the ordinary serial operation mode and the multi-threaded enhanced version provided in any embodiment of the fish body length measurement method of this application; Figure 9 This is a flowchart illustrating the overall technical process of the fish body length measurement method described in this application. Figure 10 This is a schematic diagram of the module structure of the fish body length measurement system according to an embodiment of this application; Figure 11 This is a schematic diagram of the hardware operating environment involved in the fish body length measurement method in this application embodiment.

[0024] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0025] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0026] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0027] The main solution of this application embodiment is as follows: acquiring a binocular video stream; the binocular video stream includes multiple frames of left-eye images and multiple frames of right-eye images; performing correction processing on the left-eye images and right-eye images to obtain left-eye corrected images and right-eye corrected images; multi-threaded execution steps: within a complete calculation frame, in the first thread, disparity depth calculation is performed based on the left-eye corrected images and right-eye corrected images to generate a depth map, and in the second thread, fish body detection and instance segmentation inference are performed on the left-eye corrected images to obtain a fish body instance segmentation mask; wherein, the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor; the length of the fish body is calculated based on the fish body instance segmentation mask and the depth map.

[0028] In this embodiment, for ease of description, the following description uses an edge computing platform as the execution subject.

[0029] Currently, a typical binocular fish length measurement system usually includes multiple stages such as image correction, neural network detection and segmentation, binocular disparity calculation, depth conversion, and length post-processing. In practical deployments, object detection and segmentation tasks mainly rely on accelerated resources such as graphics processing units (GPUs) or neural processing units (NPUs) for inference, while disparity calculation mainly consumes CPU and memory bandwidth. Current systems typically execute these stages sequentially, either by completing neural network inference first, followed by disparity calculation and depth generation, or by executing them in reverse order.

[0030] However, this serial execution method results in a large processing delay per frame, making it difficult to meet real-time requirements. At the same time, although domestic edge computing platforms, represented by the RK3588, possess heterogeneous computing resources such as CPU, NPU, and RGA (Raster Graphics Acceleration), the traditional serial process cannot fully utilize their collaborative processing capabilities, resulting in a waste of computing resources.

[0031] Therefore, how to reorganize task scheduling within a complete computing frame to enable heterogeneous resources to work together efficiently and reduce single-frame processing latency has become an urgent technical problem to be solved.

[0032] Based on this, this application provides a solution by acquiring a binocular video stream, which includes multiple frames of left-eye images and multiple frames of right-eye images; performing correction processing on the left-eye and right-eye images to obtain a left-eye corrected image and a right-eye corrected image; and executing a multi-threaded process, within a complete calculation frame, performing disparity depth calculation based on the left-eye corrected image and the right-eye corrected image in the first thread to generate a depth map, and performing fish body detection and instance segmentation inference on the left-eye corrected image in the second thread to obtain a fish body instance segmentation mask; wherein, the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor; and calculating the fish body length based on the fish body instance segmentation mask and the depth map.

[0033] The technical solution of this application first acquires the left and right eye images from a binocular camera, providing the raw image data foundation for subsequent stereo vision processing. Then, by correcting the left and right eye images, the left and right images satisfy epipolar constraints, eliminating the effects of lens distortion and binocular installation errors. This provides spatially aligned image data for subsequent stereo matching and depth calculation, ensuring the accuracy of disparity calculation. Next, through multi-threaded execution, two subtasks—disparity depth calculation and fish detection / instance segmentation inference—are executed in parallel within a complete calculation frame. Disparity depth calculation is performed by a first-type processor in the first thread, while fish detection / instance segmentation inference is performed by a second-type processor in the second thread. The two threads are executed concurrently. The linear execution transforms the two time-consuming steps that previously consumed computing resources sequentially into overlapping executions within the same time period. This optimizes the contribution of the two main time-consuming steps in a complete computing frame to the overall processing time from their sum to the larger of the two, significantly reducing the processing time per frame and improving the collaborative utilization efficiency of heterogeneous computing resources on the edge computing platform. Finally, the fish length is calculated based on the fish instance segmentation mask and depth map. By combining semantic segmentation foreground with binocular depth information and constraining the depth sampling range based on the fish instance segmentation mask, the interference of background and occlusion regions on disparity sampling is reduced, thus achieving accurate measurement of fish length in three-dimensional space and overcoming the lack of depth information in monocular two-dimensional image measurement.

[0034] Overall, this embodiment executes binocular parallax depth calculation and instance segmentation inference in parallel within a complete calculation frame, utilizing the heterogeneous computing resources of the first and second types of processors to transform the two main time-consuming steps that were originally executed serially into parallel execution. This reduces single-frame processing latency and improves the real-time processing capability of the fish length measurement system on the edge computing platform. At the same time, by combining instance segmentation masking with binocular depth calculation, fish length measurement based on three-dimensional spatial information is realized, improving the accuracy of fish length measurement.

[0035] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a personal computer or mobile computing terminal, or an electronic device or edge computing platform capable of performing the above functions. The following description uses an edge computing platform as an example to illustrate this embodiment and the subsequent embodiments.

[0036] Based on this, this application provides a method for measuring fish body length, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the fish body length measurement method of this application.

[0037] In this embodiment, the fish body length measurement method includes steps S10~S40: Step S10: Obtain a stereo video stream; the stereo video stream includes multiple frames of left-eye images and multiple frames of right-eye images; It should be noted that a binocular video stream contains multiple frames of left-eye images and multiple corresponding frames of right-eye images. This binocular video stream can be a real-time RTSP (Real-Time Streaming Protocol) video stream formed by stitching together the left and right cameras of a binocular camera, or it can be acquired through USB (Universal Serial Bus) interfaces, MIPI (Mobile Industry Processor Interface) interfaces, CSI (Camera Serial Interface) interfaces, GigE (Gigabit Ethernet-based camera interface standard) interfaces, camera SDKs (Software Development Kits), local video files, or other video input methods. A binocular camera refers to a stereoscopic image acquisition device that simulates the principle of human binocular vision, consisting of two cameras fixed at a certain baseline distance. The left-eye image refers to the image captured by the left camera of the binocular camera, and the right-eye image refers to the image captured by the right camera. The left-eye and right-eye images are a pair of images captured simultaneously from different perspectives of the same scene. There is a parallax between them due to the different spatial positions of the cameras; this parallax information is the basic data source for subsequent 3D depth reconstruction.

[0038] In step S10, the system acquires real-time video streams of fish schools in underwater or land-based aquaculture scenarios using a binocular camera, obtaining a complete image frame from the video stream. If the binocular camera outputs a single stitched frame composed of left and right images, the system splits the stitched frame into two independent images, namely the left-eye image and the right-eye image, according to a preset stitching method. If the binocular camera outputs two independent video streams, the system simultaneously acquires image frames with corresponding frame numbers from both video streams, using them as the left-eye image and the right-eye image, respectively. The video stream input can be a real-time network video stream transmitted via the RTSP protocol or offline video data read from a local video file. The purpose of this step is to obtain the raw image input required for binocular stereo vision processing, providing basic data for subsequent correction, matching, and depth calculation.

[0039] In one example, the binocular camera stitches the left and right images together into a single 3840×1080 resolution frame, which is then transmitted to the edge computing platform via RTSP video stream. Upon receiving this frame, the system splits it along its width centerline into a 1920×1080 left-eye image and a 1920×1080 right-eye image.

[0040] Step S20: Perform correction processing on the left eye image and the right eye image to obtain a corrected left eye image and a corrected right eye image; It should be noted that correction processing refers to the geometric transformation of the original image using the calibration parameters of the stereo camera to eliminate lens distortion and ensure that the left and right images satisfy epipolar constraints. Lens distortion is the geometric distortion of the image caused by the optical characteristics of the camera lens, mainly including radial and tangential distortion. Epipolar constraints mean that the projection points of the same point in space in the left and right images lie on the corresponding epipolar lines. After satisfying the epipolar constraints, the matching search of the left and right images can be reduced from a two-dimensional plane to a one-dimensional row search, thus significantly reducing the computational complexity of stereo matching. Correction processing is usually implemented using a remapping method, that is, calculating a correction mapping table based on the calibration parameters, and then performing coordinate transformation and grayscale interpolation on each pixel of the original image.

[0041] In step S20, the system reads the pre-stored binocular calibration file and obtains the intrinsic parameters of the left and right cameras, the extrinsic parameters of the binoculars, and distortion parameters. The camera intrinsic parameters include focal length and principal point coordinates, while the binocular extrinsic parameters include rotation matrices and translation vectors. The system generates a left-eye calibration mapping table and a right-eye calibration mapping table based on these calibration parameters. For the left-eye image obtained in step S10, the left-eye calibration mapping table is applied to perform remapping processing to obtain the corrected left-eye image; for the right-eye image, the right-eye calibration mapping table is applied to perform remapping processing to obtain the corrected right-eye image. The corrected left-eye and right-eye images satisfy epipolar constraints, meaning that the projection points of the same spatial point in the corrected left-eye and right-eye images are located on the same horizontal row. The interpolation method used during the correction process can be bilinear interpolation or bicubic interpolation to ensure the pixel accuracy of the corrected images. The purpose of this step is to eliminate the influence of lens distortion and binocular mounting errors, aligning the left and right images spatially, and providing an accurate input image that conforms to epipolar constraints for subsequent disparity calculations.

[0042] In one example, the system reads the horizontal focal length and vertical focal length of the left camera (1200 pixels, 1200 pixels, principal point coordinates (960, 540)) and the radial and tangential distortion coefficients from the dual-camera calibration file. It also reads the rotation matrix and translation vector between the left and right cameras. Based on these parameters, the system generates left and right correction mapping tables respectively, and then performs remapping processing on the left and right images to obtain the corrected left and right images respectively. After correction, the fish targets in the corrected left and right images are on the same horizontal row, requiring only a horizontal search during stereo matching.

[0043] Step S30, multi-threaded execution step: within a complete calculation frame, in the first thread, disparity depth calculation is performed based on the left eye corrected image and the right eye corrected image to generate a depth map, and in the second thread, fish body detection and instance segmentation inference are performed on the left eye corrected image to obtain a fish body instance segmentation mask; wherein, the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor; It should be noted that a complete computation frame refers to an image frame that requires the complete execution of all processing steps, including disparity depth calculation, fish detection and instance segmentation inference, and length post-processing. A thread refers to the smallest execution unit that the operating system can schedule. Multiple threads can execute concurrently within the same process, utilizing the computing resources of multi-core processors or different types of processors to achieve parallel processing. The first type of processor refers to general-purpose processors suitable for performing intensive numerical calculations and logical branching, such as CPUs, digital signal processors (DSPs), and field-programmable gate arrays (FPGAs). They excel at handling complex algorithms such as cost aggregation and path optimization in disparity calculation and can be used to perform binocular correction, SGBM (semi-global block matching algorithm) disparity calculation, depth restoration, and fish length calculation. The second type of processor refers to dedicated processors suitable for performing large-scale parallel matrix operations, such as neural network processors (NPUs) and graphics processing units (GPUs). They contain a large number of multiplication and accumulation units, enabling efficient execution of convolution operations and activation function calculations in deep learning models. They can be used to perform neural network inference for fish detection and instance segmentation models. Disparity depth calculation refers to the process of calculating the disparity map of the left and right images based on the stereo matching algorithm, and converting the disparity map into a depth map according to the geometric relationship between disparity and depth. Fish detection and instance segmentation inference refers to the process of locating fish targets in an image and performing pixel-level foreground segmentation using a deep learning model. The output fish instance segmentation mask is a binary image, where foreground pixels represent fish regions and background pixels represent non-fish regions.

[0044] In step S30, after completing binocular calibration in step S20, the system creates two threads: a first thread and a second thread. Within a complete calculation frame, the system assigns the disparity depth calculation task to the first thread, executed by a first-type processor; and assigns the fish detection and instance segmentation inference task to the second thread, executed by a second-type processor. The first and second threads run in parallel, with no data dependency between them, and can execute independently on their respective processors. When the first thread performs disparity depth calculation, it acquires the left and right calibrated images, performs stereo matching sequentially to obtain a disparity map, and then converts the disparity map into a depth map according to the binocular visual depth calculation formula. When the second thread performs fish detection and instance segmentation inference, it inputs the left calibrated image into a pre-trained deep learning model. The model performs forward inference calculation and outputs the detection box of the fish target and the instance segmentation mask. The system waits in the main thread for both the first and second threads to complete before proceeding to the subsequent length calculation step. By parallelizing parallax depth calculation and neural network inference within a single frame, two time-consuming steps that originally consumed computational resources sequentially are transformed into overlapping executions within the same time period. The contribution of the two main time-consuming steps in a complete computation frame to the overall processing time is optimized from their sum to the larger of the two. The purpose of this step is to fully utilize the heterogeneous computing resources of different types of processors on the edge computing platform, executing two independent and time-consuming tasks in parallel, thereby reducing the processing time per frame and improving the system's real-time processing capabilities.

[0045] In one example, the system operates on the RK3588 edge computing platform, which supports heterogeneous computing. The first type of processor is the ARM (Reduced Instruction Set Computing (RISC) processor architecture developed by Acorn in 1983) central processing unit within the RK3588 chip, and the second type is the neural network processor integrated within the RK3588 chip. The system creates a first thread, which, executed by the central processing unit, performs a semi-global matching algorithm to calculate disparity depth. Specifically, this includes stereo matching of the left and right eye corrected images to generate a disparity map, and converting the disparity map into a depth map according to the depth calculation formula. Simultaneously, the system creates a second thread, which, executed by the neural network processor, performs YOLOv11-seg instance segmentation model inference within the RKNN (Rockchip's proprietary neural network model format for NPU chips) framework. Specifically, this includes inputting the left eye corrected image into the YOLOv11-seg model, which calculates the detection bounding box of the fish target and the corresponding instance segmentation mask through a convolutional neural network. The two threads execute in parallel within the same complete computation frame. The first thread's disparity depth calculation takes about 200 milliseconds, and the second thread's neural network inference takes about 180 milliseconds. The total time for the two threads to execute in parallel is determined by the first thread, which takes longer, and is about 200 milliseconds. Compared to the 380 milliseconds of serial execution, the processing time per frame is significantly reduced.

[0046] In another example, the first type of processor can also be a graphics processor, and the second type of processor can also be a digital signal processor, which can also achieve parallel execution of parallax depth calculation and neural network inference.

[0047] Step S40: Calculate the length of the fish body based on the segmentation mask of the fish body instance and the depth map.

[0048] It should be noted that the fish instance segmentation mask refers to a binary mask image that identifies the region where the fish target is located in the left-eye calibration image at the pixel level, where the value of each pixel indicates whether that pixel belongs to the fish target. The depth map refers to a two-dimensional image of the same size as the left-eye calibration image, where the value of each pixel represents the distance from the corresponding spatial scene point to the camera plane. The fish length refers to the actual physical length of the fish along its principal axis in three-dimensional space, usually measured in centimeters. The core idea of ​​calculating the fish length based on the fish instance segmentation mask and the depth map is as follows: using the instance segmentation mask to determine the pixel region of the fish target in the image, extracting the principal axis and endpoints of the fish within this region, and then combining the corresponding depth information from the depth map to convert the pixel coordinates on the image plane into three-dimensional spatial coordinates, thereby calculating the true physical length of the fish based on the coordinate difference in three-dimensional space.

[0049] In step S40, based on the fish instance segmentation mask and depth map obtained in step S30, the system first determines the target depth region of interest corresponding to the fish instance segmentation mask in the depth map. Then, it performs maximum connected component preservation and hole filling on the fish instance segmentation mask to remove segmentation noise and internal holes, obtaining a complete target foreground mask. Principal component analysis is performed on the target foreground mask to extract the principal axis direction of the fish and its two endpoints, namely the head and tail endpoints. Next, the corresponding depth values ​​are obtained from the depth map based on the pixel coordinates of the two endpoints. Finally, using the inverse camera projection transform formula, the pixel coordinates of the two endpoints and the corresponding depth values ​​are converted into three-dimensional spatial coordinates in the camera coordinate system. The Euclidean distance between the two three-dimensional coordinate points is calculated to obtain the three-dimensional length of the fish. Throughout the calculation process, the length calculation is strictly based on the fish region constrained by the instance segmentation mask, avoiding the introduction of background depth values ​​into the measurement, thereby reducing the interference of background parallax on the fish length estimation. The purpose of this step is to combine semantic segmentation foreground information with binocular depth information to achieve accurate measurement of fish length in three-dimensional space, overcoming the shortcomings of lack of depth information and background parallax interference in monocular two-dimensional image measurement.

[0050] In one example, the system uses the fish instance segmentation mask output by the YOLOv11-seg model obtained in step S30 and the depth map generated by the SGBM algorithm to locate the foreground region of the fish corresponding to the mask in the depth map. After performing maximum connected component preservation on the mask, a binary mask containing only the main target region of the fish is obtained, and the internal holes in it are filled. Principal component analysis is performed on this mask to obtain the principal axis direction of the fish, and the two endpoints of the principal axis are extracted, with pixel coordinates of (100, 200) and (400, 350), respectively. The depth values ​​at these two coordinates are found in the depth map, and the depth of the first endpoint is found to be 2.5 meters, and the depth of the second endpoint is found to be 2.55 meters. The system uses the intrinsic parameters of the left eye camera to convert the pixel coordinates and depth values ​​of the two endpoints into three-dimensional spatial coordinates, and calculates the Euclidean distance between the two three-dimensional points to be 0.274 meters, or 27.4 centimeters, as the three-dimensional length of the fish.

[0051] In another example, if there are multiple fish targets in the fish instance segmentation mask, the system independently performs the above principal component analysis, endpoint extraction and depth sampling steps for each target instance, and calculates the three-dimensional length of each fish one by one.

[0052] Furthermore, the fish length measurement method in this embodiment can also be referred to... Figure 2 , Figure 2 A schematic diagram of the system architecture corresponding to the fish body length measurement method of this embodiment is shown.

[0053] Additionally, please see Figure 3 , Figure 3 A comparative diagram of the traditional serial process and the intra-frame parallel process of this application is provided. In the traditional serial process, the total time for serial processing of a complete computation frame, T_serial, is approximately equal to the preprocessing time T_pre + disparity depth calculation time T_depth + instance segmentation neural network inference time T_nn + postprocessing time T_post + display time T_display. In the intra-frame parallel process of this application, the total time for completing a complete computation frame, T_parallel, after intra-frame parallel scheduling, is approximately equal to the preprocessing time T_pre + the larger of the time taken for disparity depth calculation and instance segmentation inference, max(T_depth, T_nn) + postprocessing time T_post + display time T_display. This reduces the processing time of this application and improves processing efficiency.

[0054] This embodiment provides a method for measuring fish body length. First, it acquires left and right eye images from a binocular camera, providing the raw image data foundation for subsequent stereo vision processing. Then, it corrects the left and right eye images to ensure they meet epipolar constraints, eliminating the effects of lens distortion and binocular installation errors. This provides spatially aligned image data for subsequent stereo matching and depth calculation, ensuring the accuracy of disparity calculation. Next, it employs multi-threaded execution, performing disparity depth calculation and fish detection / instance segmentation inference in parallel within a single calculation frame. Disparity depth calculation is executed by a first-type processor in the first thread, while fish detection / instance segmentation inference is executed by a second-type processor in the second thread. The parallel execution of two threads transforms the two time-consuming steps that would otherwise consume computational resources sequentially into overlapping executions within the same timeframe. This optimizes the contribution of the two main time-consuming steps in a complete computation frame to the larger of the two, significantly reducing the processing time per frame and improving the collaborative utilization efficiency of heterogeneous computing resources on the edge computing platform. Finally, the fish length is calculated based on the fish instance segmentation mask and depth map. By combining semantic segmentation foreground with binocular depth information and constraining the depth sampling range based on the fish instance segmentation mask, the interference of background and occlusion regions on disparity sampling is reduced, thus achieving accurate measurement of fish length in three-dimensional space and overcoming the lack of depth information in monocular two-dimensional image measurement.

[0055] Overall, this embodiment executes binocular parallax depth calculation and instance segmentation inference in parallel within a complete calculation frame, utilizing the heterogeneous computing resources of the first and second types of processors to transform the two main time-consuming steps that were originally executed serially into parallel execution. This reduces single-frame processing latency and improves the real-time processing capability of the fish length measurement system on the edge computing platform. At the same time, by combining instance segmentation masking with binocular depth calculation, fish length measurement based on three-dimensional spatial information is realized, improving the accuracy of fish length measurement.

[0056] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 In step S30, i.e., the multi-threaded execution step, the step of calculating disparity depth based on the left-eye corrected image and the right-eye corrected image in the first thread to generate a depth map includes steps S311 to S316: Step S311: Perform grayscale conversion and texture enhancement processing on the left eye correction image and the right eye correction image to obtain a left eye grayscale image and a right eye grayscale image; It should be noted that grayscale conversion refers to the process of converting a color image to a grayscale image, the purpose of which is to reduce data dimensionality and the computational burden of subsequent disparity calculations. Texture enhancement refers to the operation of locally enhancing the contrast of a grayscale image, the purpose of which is to highlight the detailed texture information in the image, so that the subsequent stereo matching process can obtain richer and more reliable pixel features, thereby improving the accuracy of disparity calculation. In this application, texture enhancement can employ a contrast-limited adaptive histogram equalization method, or other local contrast enhancement methods, such as histogram equalization, gamma correction, or denoising enhancement.

[0057] In step S311, the left and right eye corrected images are converted to grayscale to reduce data dimensionality and computational complexity in subsequent disparity calculations. Next, texture enhancement processing is applied to the grayscale images to highlight detailed texture information, enabling richer and more reliable pixel features for subsequent stereo matching and improving the accuracy of disparity calculations. Texture enhancement can employ contrast-limited adaptive histogram equalization, or it can use histogram equalization, gamma correction, denoising enhancement, or other local contrast enhancement methods. The purpose of this step is to provide high-quality input images for subsequent disparity calculations, thereby improving the accuracy and robustness of fish length measurement.

[0058] In one example, a contrast-limited adaptive histogram equalization method is used to enhance the texture of the left and right grayscale images. This method divides the image into multiple small regions, performs histogram equalization on each region separately, and limits the contrast amplification to avoid noise amplification, thereby enhancing local texture details while maintaining the overall visual consistency of the image.

[0059] Step S312: Scale the left-eye grayscale image and the right-eye grayscale image according to a preset scaling factor s to obtain a scaled left-eye grayscale image and a scaled right-eye grayscale image; where 0 <s≤1; It should be noted that the scale factor 's' refers to the scaling factor used to scale the image size, and its value ranges from greater than 0 to less than or equal to 1. A smaller scale factor 's' results in a smaller scaled image size, less data required for disparity calculation, and faster computation speed, but also a corresponding increase in the loss of image detail. The value of the scale factor 's' needs to be determined comprehensively based on the computing power of the edge computing platform, the distance range of the target fish, and the required measurement accuracy.

[0060] In step S312, the left and right grayscale images are scaled according to a preset scaling factor s. Specifically, the left and right grayscale images are multiplied by the scaling factor s in both the width and height directions to obtain scaled grayscale images for the left and right eyes. Downscaling effectively reduces the number of pixels involved in disparity calculation, thus reducing the time consumed in the disparity calculation process. The scaling factor s ranges from 0 to 1, and its specific value can be set according to the computing power of the edge computing platform, the target fish distance range, and the required measurement accuracy. The purpose of this step is to reduce the computational load by lowering the image resolution while ensuring the accuracy of subsequent disparity calculations, thereby improving the system's real-time processing capabilities.

[0061] In one example, the scaling factor s is set to 0.75, which means that the left and right grayscale images are scaled to 75% of their original size in both the width and height directions. The scaled image size is 56.25% of the original image size, thus significantly reducing the amount of data required for parallax calculation.

[0062] In another example, for edge computing platforms with strong computing power or application scenarios with high measurement accuracy requirements, the scaling factor s can be 0.9 or 1.0; for platforms with weak computing power or application scenarios with high real-time requirements, the scaling factor s can be 0.5.

[0063] Step S313: Calculate a small-scale disparity map based on the left-eye scaled grayscale image and the right-eye scaled grayscale image; It's important to note that a disparity map represents the horizontal positional difference between corresponding pixels in the left and right images. The pixel value of each pixel represents the disparity between the left and right images. A small-scale disparity map is calculated at the downscaled image resolution, with its size matching the scaled image size. The basic principle of disparity calculation is that, due to the horizontal baseline distance between binocular cameras, the pixel position of the same spatial point differs horizontally between the left and right images; this offset is the disparity.

[0064] In step S313, using the scaled grayscale images of the left and right eyes as input, stereo matching is performed at the downscaled image resolution to calculate the horizontal disparity values ​​between corresponding pixels in the left and right scaled images, thereby generating a small-scale disparity map. Disparity calculation can employ a semi-global matching algorithm. This algorithm calculates the matching cost of pixels at different disparity values ​​and aggregates the costs along multiple directions, preserving edge accuracy while exhibiting good robustness to areas with weak texture. Of course, local or global matching algorithms can also be used in other implementations. The purpose of this step is to obtain disparity distribution information in the small-scale image space, providing a foundation for subsequent full-resolution disparity recovery.

[0065] In one example, a semi-global matching algorithm is used to compute a small-scale disparity map. This algorithm searches for the pixel with the minimum matching cost in the corresponding row of the right-eye scaled grayscale image for each pixel in the left-eye scaled grayscale image, aggregating costs along 8 or 16 directions. Finally, a winner-takes-all strategy is used to determine the optimal disparity value for each pixel, thus generating the small-scale disparity map. The parameter settings of the semi-global matching algorithm, such as the disparity range and penalty coefficient, can be adjusted according to the baseline length of the binocular camera, the distance range of the target fish, and the image resolution.

[0066] Step S314: Filter the small-scale disparity map and restore the filtered small-scale disparity map to the full resolution size to obtain the restored small-scale disparity map; It's important to note that filtering refers to smoothing or correcting pixel values ​​in an image. Its purpose is to remove noise and outliers from the disparity map, making the disparity distribution more continuous and smooth. Common disparity map filtering methods include bilateral filtering, median filtering, guided filtering, and weighted least squares filtering. Restoring an image to its full resolution refers to the process of enlarging a low-resolution image to its original size using image interpolation methods. Commonly used interpolation methods include nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation.

[0067] In step S314, the small-scale disparity map is first filtered to remove noise points and outliers from the disparity calculation process, making the disparity distribution more continuous and smooth. Bilateral filtering can be used, which smooths the disparity map while preserving disparity edge information effectively. Other disparity smoothing and outlier suppression methods, such as median filtering, guided filtering, or weighted least squares filtering, can also be employed. Then, the filtered small-scale disparity map is restored to its full-resolution size using image interpolation, i.e., the same size as the original left and right corrected images, resulting in the restored small-scale disparity map. Interpolation methods can include bilinear interpolation, bicubic interpolation, or nearest neighbor interpolation. The purpose of this step is to remove noise from the disparity map and restore the full-resolution disparity map, providing accurate disparity data for subsequent depth calculations.

[0068] In one example, bilateral filtering is used to filter the small-scale disparity map. Bilateral filtering comprehensively considers the spatial proximity relationship between pixels and the similarity of pixel values, smoothing the disparity map while maintaining clear disparity edges. Subsequently, bilinear interpolation is used to enlarge the filtered small-scale disparity map to full resolution. Bilinear interpolation calculates the interpolation point value by weighted averaging the values ​​of four adjacent pixels, which can obtain a relatively smooth interpolation result while ensuring computational efficiency.

[0069] Step S315: Based on the scale factor s, convert the recovered small-scale disparity map into a full-resolution equivalent disparity map. It should be noted that the full-resolution equivalent disparity map refers to a disparity map that has the same size as the full-resolution image, and the disparity value of each pixel corresponds to the disparity value at the full-resolution scale. Since the small-scale disparity map is calculated on the downscaled image, its disparity value reflects the pixel offset in the downscaled image space, and needs to be converted according to the scale factor to obtain the equivalent disparity value at the full-resolution scale.

[0070] In step S315, the disparity values ​​of each pixel in the restored small-scale disparity map are converted into full-resolution equivalent disparity values ​​based on the scale factor s. Specifically, since the small-scale disparity map is calculated on the downscaled image, its disparity values ​​reflect the horizontal offset of pixels in the downscaled image space, while the full-resolution equivalent disparity map needs to reflect the horizontal offset of pixels in the original full-resolution image space. Therefore, the disparity values ​​of each pixel in the restored small-scale disparity map need to be divided by the scale factor s to obtain the corresponding full-resolution equivalent disparity values. The disparity values ​​of each pixel in the full-resolution equivalent disparity map are greater than the corresponding pixel values ​​in the restored small-scale disparity map because the pixel spacing is smaller and the disparity changes are more refined at full resolution, resulting in a larger pixel offset for the same spatial distance. The purpose of this step is to correctly map the disparity data in the downscaled space to the full-resolution space, enabling subsequent depth calculations and fish length measurements to be performed in a unified coordinate system.

[0071] In one example, if the disparity value of a pixel in the restored small-scale disparity map is 15 pixels and the scale factor s is 0.75, then the full-resolution equivalent disparity value for that pixel is 15 divided by 0.75, which equals 20 pixels. In another example, if the disparity value of a pixel in the restored small-scale disparity map is 8 pixels and the scale factor s is 0.5, then the full-resolution equivalent disparity value for that pixel is 8 divided by 0.5, which equals 16 pixels.

[0072] Step S316: Calculate the depth map based on the full-resolution equivalent disparity map.

[0073] It's important to note that a depth map is an image representing the distance from each pixel in a spatial point to the camera plane, where the value of each pixel represents the depth value at that point. There is an inverse relationship between depth and parallax: the greater the parallax, the closer the corresponding spatial point is to the camera; the smaller the parallax, the farther the corresponding spatial point is from the camera. Depth maps are calculated based on the principles of binocular stereo vision and provide crucial spatial depth information for subsequent 3D length measurement of fish.

[0074] In step S316, based on the disparity values ​​of each pixel in the full-resolution equivalent disparity map, and combined with the intrinsic and extrinsic parameters of the binocular camera, the disparity values ​​are converted into depth values, thereby generating a depth map. Specifically, for each pixel in the full-resolution equivalent disparity map, the depth value of the corresponding spatial point is calculated based on the disparity value of that point, the horizontal focal length of the left eye camera, and the binocular baseline. The depth values ​​of all pixels constitute the depth map, and each pixel value in the depth map represents the depth information of the corresponding pixel in the left eye calibrated image. The purpose of this step is to convert the disparity information on the two-dimensional image plane into depth information in three-dimensional space, providing a depth data foundation for the subsequent calculation of the three-dimensional length of the fish.

[0075] In one example, for a pixel at pixel coordinates in the full-resolution equivalent disparity map, if the disparity value of that pixel is 20 pixels, the horizontal focal length of the left eye camera is 1000 pixels, and the binocular baseline length is 0.12 meters, then according to the depth calculation formula, the depth value is equal to 1000 × 0.12 divided by 20, which is 6 meters.

[0076] In this embodiment, the left and right eye corrected images are first grayscaled and texture-enhanced to reduce data dimensionality, decrease the computational load of subsequent disparity calculations, and highlight detailed texture information in the images. This allows the subsequent stereo matching process to obtain richer and more reliable pixel features, improving the accuracy of disparity calculations. Then, by scaling the grayscale images according to a preset scaling factor, the number of pixels involved in disparity calculations can be effectively reduced, thereby reducing the time consumed in the disparity calculation stage and improving processing speed. Next, the small-scale disparity map is filtered and restored to remove noise points and outliers, making the disparity distribution more continuous and smooth. At the same time, the full-resolution disparity map is restored, providing accurate disparity data for subsequent depth calculations. Finally, the restored small-scale disparity map is converted into a full-resolution equivalent disparity map, enabling subsequent depth calculations and fish length measurements to be performed in a unified coordinate system, and the depth map is calculated accordingly. Overall, this embodiment significantly reduces the computational load of the disparity calculation stage by combining downscaling disparity calculation with full resolution restoration, while ensuring the accuracy of subsequent length measurement, thereby improving the real-time processing capability of the binocular vision fish length measurement system on the edge computing platform.

[0077] In one optional implementation, the step of converting the recovered small-scale disparity map into a full-resolution equivalent disparity map based on the scale factor s, i.e., step S315, includes step S3151: Step S3151: Calculate the full-resolution equivalent disparity value of each pixel based on the disparity value of each pixel in the recovered small-scale disparity map and the scale factor. In step S3151, for each pixel in the restored small-scale disparity map, the disparity value of that pixel is obtained, and this disparity value is divided by the scale factor s. The quotient is used as the full-resolution equivalent disparity value of the corresponding pixel in the full-resolution equivalent disparity map. This can be expressed by the formula: the full-resolution equivalent disparity value equals the restored small-scale disparity value divided by the scale factor s. By performing the above conversion operation pixel by pixel, the full-resolution equivalent disparity map is obtained. The purpose of this step is to accurately map the disparity values ​​calculated in the downscaled image space to the full-resolution image space to ensure the accuracy of depth calculation.

[0078] In one example, the restored small-scale disparity map has a size of 2560×1440 pixels and a scale factor of 0.8. For a pixel located at coordinates , its disparity value is 12 pixels. Therefore, the full-resolution equivalent disparity value for that point is 12 divided by 0.8, which equals 15 pixels. Performing the same division operation on all pixels in the restored small-scale disparity map yields the full-resolution equivalent disparity map.

[0079] Correspondingly, the step of calculating the depth map based on the full-resolution equivalent disparity map, i.e., step S316, includes step S3161: Step S3161: Calculate the depth value of each pixel based on the full-resolution equivalent parallax value of each pixel, the horizontal focal length of the left eye camera, and the binocular baseline to obtain the depth map.

[0080] Step S3161: Calculate the depth value of each pixel based on the full-resolution equivalent parallax value of each pixel, the horizontal focal length of the left eye camera, and the binocular baseline to obtain the depth map.

[0081] In step S3161, for each pixel in the full-resolution equivalent disparity map, the full-resolution equivalent disparity value of that pixel is obtained. Combined with the horizontal focal length of the left-eye camera and the binocular baseline length, the depth value of the corresponding spatial point is calculated. Specifically, the depth value is calculated using the depth calculation formula for binocular stereo vision, i.e., the depth value equals the horizontal focal length of the left-eye camera multiplied by the binocular baseline and then divided by the full-resolution equivalent disparity value. The horizontal focal length of the left-eye camera and the binocular baseline length are obtained through the binocular calibration process. This calculation is performed pixel-by-pixel to obtain the depth value of each pixel, and the depth values ​​of all pixels constitute the depth map. When the full-resolution equivalent disparity value of a pixel is invalid, zero, or non-finite, the depth value of that pixel is marked as invalid. The purpose of this step is to convert disparity information into depth information in three-dimensional space, providing accurate spatial depth data for subsequent fish length measurement.

[0082] In one example, the horizontal focal length of the left eye camera is 1200 pixels, the binocular baseline length is 0.15 meters, and the full-resolution equivalent disparity value of a pixel in the full-resolution equivalent disparity map is 24 pixels. Therefore, the depth value corresponding to this pixel is equal to 1200 multiplied by 0.15 divided by 24, which is 7.5 meters. Performing the above calculation for each pixel in the full-resolution equivalent disparity map yields the complete depth map.

[0083] In this embodiment, through precise parallax conversion and depth calculation, the parallax values ​​calculated in the downscaled space are correctly mapped to the full-resolution space, and the depth values ​​of each pixel are calculated based on the principle of binocular stereo vision, providing an accurate depth data basis for subsequent three-dimensional length measurement of the fish.

[0084] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 After step S40, the fish body length measurement method further includes steps S51-S54: Step S51: Obtain the frequency division parameters; It should be noted that the frequency division parameter refers to the parameter used to control the execution frequency of the complete calculation frame. It determines how many frames are between each complete fish detection, disparity depth calculation, and length measurement. The frequency division parameter is an integer greater than or equal to 1. A frequency division parameter of 1 means that a complete calculation is performed every frame, a frequency division parameter of 2 means that a complete calculation is performed every two frames, and so on. The frequency division parameter can be dynamically set according to the target fish's movement speed, display smoothness requirements, and length measurement refresh rate requirements.

[0085] In step S51, the frequency division parameters are read from environment variables, configuration files, or obtained through program parameters, user input, etc. These parameters are used in subsequent steps to determine whether the current frame needs to undergo complete computation. The purpose of this step is to obtain the control parameters of the frequency division multiplexing strategy, providing a basis for subsequent frame determination and result multiplexing.

[0086] In one example, the frequency division parameter is stored in a configuration file, and the value of the frequency division parameter is read from the configuration file as 2 when the system starts up.

[0087] In another example, the frequency division parameter is passed in as a command-line argument when the program is started, for example, by specifying the frequency division parameter as 3 using "--interval=3".

[0088] Step S52: Determine whether the current frame is a complete calculation frame based on the current frame number and the frequency division parameter; It should be noted that a complete computation frame refers to an image frame that requires the full execution of all processing steps, including image correction, disparity and depth calculation, neural network detection, segmentation and inference, mask post-processing, and fish length calculation. In contrast to a complete computation frame is a reused frame, which does not repeat the detection, disparity, and length calculations; instead, it reuses the calculation results from the most recent complete computation frame for display.

[0089] In step S52, the frame number of the current frame is obtained, and it is determined whether the current frame is a complete calculation frame based on the current frame number and the frequency division parameter. Specifically, a determination rule can be set: when the cached result is empty, i.e., no complete calculation has been performed, the current frame is a complete calculation frame; when the cached result is not empty, if the remainder of the current frame number divided by the frequency division parameter is zero, the current frame is a complete calculation frame; otherwise, it is a reused frame. This can be expressed by the formula: the condition for determining a complete calculation frame is that the cached result is empty or the current frame number modulo the frequency division parameter equals 0; when the above conditions are met, the current frame is a complete calculation frame; otherwise, it is a reused frame. The purpose of this step is to determine the processing mode of the current frame to decide whether to execute a complete calculation process or reuse existing results.

[0090] In one example, the frequency division parameter is 2, the current frame number is 0, and the buffer result is empty, satisfying the condition that the buffer result is empty. Therefore, the current frame is a complete computation frame. When the current frame number is 1, the buffer result is not empty and the modulo 1 divided by 2 is not equal to 0. Therefore, the current frame is a multiplexed frame. When the current frame number is 2, the buffer result is not empty, but the modulo 2 divided by 2 is equal to 0. Therefore, the current frame is a complete computation frame.

[0091] Step S53: If the current frame is the complete calculation frame, then execute the multi-threaded execution step and update the calculation result to the cached result; In step S53, when step S52 determines that the current frame is a complete calculation frame, a multi-threaded execution step is performed. This means that disparity depth calculation and neural network detection, segmentation, and inference are performed in parallel within a complete calculation frame, and fish length calculation is performed after both are completed. All calculation results obtained from this complete calculation, including detection boxes, instance segmentation masks, fish length, and labeled images, are stored in memory as cached results for display in subsequent reused frames. The purpose of this step is to concentrate computational resources on a subset of key frames through a frequency-division calculation strategy, without repeatedly counting the same target, thereby reducing the load on continuous video processing.

[0092] In one example, the current frame is determined to be a complete calculation frame. The system starts the first thread to calculate disparity depth, while the second thread performs neural network detection, segmentation, and inference. The two threads execute in parallel. After both threads are completed, the fish length is calculated based on the segmentation mask and depth map. The detection results, length information, and labeled images are stored as cached results, and the measurement records are written to a structured results file.

[0093] Step S54: If the current frame is not the complete calculation frame, the cached result is reused for display, and the writing of measurement records to the structured result file is abandoned.

[0094] In step S54, when step S52 determines that the current frame is a reused frame (i.e., not a complete calculation frame), the system does not repeatedly execute neural network inference, disparity depth calculation, and length calculation. Instead, it directly reads the cached result of the most recent complete calculation frame from memory and uses the labeled image, detection box, and length information from the cached result for the display output of the current frame. Simultaneously, since no new measurements are performed in the current frame, the system abandons writing measurement records to the structured result file, avoiding duplicate statistics for the same target. The purpose of this step is to improve the smoothness of real-time display while ensuring that the measurement data recorded in the structured result file is not redundant.

[0095] In one example, the current frame is determined to be a reused frame. The system reads the cached results of the previous complete calculation frame from memory, overlays the same detection box, fish length and annotation information as the previous complete calculation frame onto the display screen of the current frame, and overlays the "Reuse" status text on the screen to indicate that the current frame displays a reused result, but does not write any new measurement records into the structured result file.

[0096] In this embodiment, firstly, frequency division parameters are obtained to provide a basis for subsequent calculation frame determination and result reuse. Then, based on the current frame number and frequency division parameters, it is accurately determined whether the current frame is a complete calculation frame, realizing on-demand allocation of computing resources. Next, for complete calculation frames, a complete detection, disparity, and length measurement process is executed, and the calculation results are updated to cached results, ensuring the measurement accuracy of key frames and saving reusable calculation results. Finally, for reused frames, the cached results are directly reused for display, and writing measurement records to the structured result file is abandoned, significantly reducing the load of continuous video processing. Overall, this embodiment, through the mechanism of separating complete calculation frames and reused frames, improves the real-time display smoothness while ensuring that length measurement statistics are not repeated, and reduces the processing load of the edge computing platform.

[0097] Based on the first embodiment of this application, in the fourth embodiment of this application, the content that is the same as or similar to that in the first embodiment can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 The step of calculating the length of the fish body based on the segmentation mask of the fish body instance and the depth map, i.e., step S40, includes steps S41 to S46: Step S41: Construct a target depth region of interest based on the segmentation mask of the fish body instance; It should be noted that the target depth region of interest refers to the region in the depth map corresponding to the segmentation mask of the fish instance, that is, the pixel area covered by the foreground of the fish, used to define the range of subsequent depth sampling and length calculation. By constructing the target depth region of interest, depth calculation and length measurement can be concentrated in the area where the fish target is located, reducing the interference of the background area on the measurement results.

[0098] In step S41, based on the fish instance segmentation mask obtained in step S30, a pixel region corresponding to the mask is determined in the depth map, and this region is designated as the target depth region of interest. The pixel coordinates of the target depth region of interest correspond exactly to the pixel coordinates of the fish instance segmentation mask, and the depth value of each pixel within its range is used for subsequent fish length calculation. The purpose of this step is to separate the foreground and background regions of the fish, so that subsequent length calculations can focus on the fish target itself, improving measurement accuracy.

[0099] In one example, the fish instance segmentation mask is a binary image of the same size as the left eye corrected image, where the pixel value of the foreground region of the fish is 1 and the pixel value of the background region is 0. The system uses all pixels in the depth map corresponding to the regions with a pixel value of 1 in the fish instance segmentation mask as the target depth region of interest.

[0100] Step S42: Perform maximum connectivity preservation and hole filling processing on the fish body instance segmentation mask to obtain the target foreground mask; It should be noted that maximum connected component preservation refers to detecting all connected regions in a binary image and retaining only the connected region with the largest number of pixels, while deleting other smaller connected regions. A connected region is a region composed of adjacent pixels with a value of 1. Hole filling refers to the operation of filling the holes (background pixels) inside the foreground region of a binary image with foreground pixels.

[0101] In step S42, connected component analysis is performed on the fish instance segmentation mask to detect all connected regions in the mask. The number of pixels in each connected region is calculated, and only the connected region with the largest number of pixels is retained, while other connected regions with smaller number of pixels are deleted to remove isolated small regions caused by occlusion, noise, or inaccurate segmentation. Next, hole-filling processing is performed on the retained largest connected region to fill the holes inside the foreground region with foreground pixels, making the target foreground mask a complete, hole-free connected region. The purpose of this step is to remove noise and isolated regions from the segmentation result, obtain a more complete and stable fish foreground mask, and improve the accuracy of subsequent principal axis extraction and length calculation.

[0102] In one example, the fish body instance segmentation mask contains a large main connected region of the fish body and several small isolated regions caused by segmentation noise. The system calculates the area of ​​all connected regions, retains the largest connected region, and deletes all other small connected regions. Then, it detects holes inside the retained connected regions and fills these holes with foreground pixels to obtain a complete, hole-free target foreground mask.

[0103] Step S43: Perform principal component analysis on the target foreground mask to extract the fish body principal axis and the two endpoints of the fish body principal axis; It should be noted that Principal Component Analysis (PCA) is a data dimensionality reduction and feature extraction method. It determines the main distribution direction of the data by calculating the eigenvectors of the covariance matrix of a data point set. In image processing, performing PCA on the pixels of a target foreground mask can yield the target's principal axis direction. The principal axis of a fish refers to the central axis along its length, with its two endpoints corresponding to the head and tail.

[0104] In step S43, the coordinate set of all foreground pixels in the target foreground mask is obtained. Principal component analysis is performed on these coordinate points, i.e., the covariance matrix of the coordinate point set is calculated, and the eigenvalues ​​and eigenvectors of the covariance matrix are obtained. The direction of the eigenvector corresponding to the largest eigenvalue is the principal axis direction of the fish. The two endpoints of the fish are determined along the principal axis direction. The foreground pixels are projected onto the principal axis, and the two points with the smallest and largest projection values ​​are taken as the two endpoints of the fish's principal axis. The purpose of this step is to accurately extract the principal axis direction and the two endpoints of the fish based on the fish foreground mask, providing basic data for subsequent pixel length calculation and 3D length calculation.

[0105] In one example, the target foreground mask contains 5000 foreground pixels. The system uses the coordinates of these pixels as input data for principal component analysis, and the direction of the first principal component is the direction of the fish's principal axis. All foreground pixels are projected along the principal axis, and the point with the smallest projection value and the point with the largest projection value are found, corresponding to the head and tail ends of the fish, respectively.

[0106] Step S44: Calculate the pixel length based on the two endpoints; In step S44, based on the pixel coordinates of the two endpoints of the fish's main axis extracted in step S43, the Euclidean distance between the two endpoints is calculated to obtain the pixel length of the fish. The pixel length is calculated using the Euclidean distance formula, that is, by calculating the straight-line distance between the two endpoints on the image plane. The pixel length is expressed in pixels and serves as an intermediate parameter for subsequent approximate length calculations. The purpose of this step is to obtain the length information of the fish on the image plane, providing a basis for subsequent approximate length calculations.

[0107] In one example, the coordinates of the two endpoints of the main axis of the fish are (100, 200) and (400, 350). The pixel length is equal to the Euclidean distance between the two endpoints, which is the square root of (400-100) plus the square root of (350-200), resulting in 335.41 pixels.

[0108] Step S45: Based on the depth map, obtain the depth values ​​corresponding to the two endpoints respectively; In step S45, based on the pixel coordinates of the two endpoints of the fish's main axis extracted in step S43, the depth values ​​at the corresponding locations are found in the depth map to obtain the depth values ​​corresponding to each of the two endpoints. The depth value represents the distance of the spatial point where the endpoint is located from the camera plane. The purpose of this step is to obtain the depth information of the endpoints of the fish's main axis, providing depth data for subsequent 3D length calculations.

[0109] In one example, the coordinates of the first endpoint of the fish's main axis are (100, 200) and the coordinates of the second endpoint are (400, 350). The system finds the depth value at coordinates (100, 200) in the depth map as the depth value of the first endpoint and the depth value at coordinates (400, 350) as the depth value of the second endpoint.

[0110] Step S46: Calculate the three-dimensional length of the fish body based on the two endpoints, the corresponding depth values, and the intrinsic parameters of the left eye camera.

[0111] It should be noted that the intrinsic parameters of the left-eye camera include the horizontal focal length, vertical focal length, and principal point coordinates of the left-eye camera, which are used to convert pixel coordinates in the image coordinate system to spatial coordinates in the camera coordinate system. The three-dimensional length of the fish body refers to its length along the principal axis of the fish body in three-dimensional space, measured in centimeters or meters.

[0112] In step S46, based on the pixel coordinates of the two endpoints of the fish's main axis extracted in step S43, the depth values ​​corresponding to the two endpoints obtained in step S45, and the intrinsic parameters of the left-eye camera, the two endpoints are transformed from image pixel coordinates to three-dimensional spatial coordinates in the camera coordinate system through camera coordinate system transformation. Then, the Euclidean distance between the two three-dimensional spatial points is calculated to obtain the three-dimensional length of the fish. Specifically, for any endpoint, its pixel coordinates and corresponding depth value are used as input, and the X, Y, and Z coordinates of that endpoint in the camera coordinate system are calculated using the inverse camera projection transformation formula. Then, the Euclidean distance between the two endpoints in three-dimensional space is calculated as the three-dimensional length of the fish. The purpose of this step is to accurately calculate the length of the fish in real three-dimensional space based on the principle of three-dimensional reconstruction using binocular vision, overcoming the deficiency of lack of depth information in monocular two-dimensional image measurement.

[0113] In one example, the pixel coordinates of the first endpoint are (100, 200), with a depth of 2.5 meters; the pixel coordinates of the second endpoint are (400, 350), with a depth of 2.55 meters; the horizontal focal length of the left camera is 1200 pixels, the vertical focal length is 1200 pixels, and the principal point coordinates are (960, 540). According to the inverse transformation formula of camera projection, the three-dimensional coordinates of the first endpoint in the camera coordinate system are calculated to be (-0.717 meters, -0.283 meters, 2.5 meters), and the three-dimensional coordinates of the second endpoint in the camera coordinate system are (-0.467 meters, -0.158 meters, 2.55 meters). The Euclidean distance between the two three-dimensional points is 0.274 meters, or 27.4 centimeters. This length is converted to centimeters and taken as the three-dimensional length of the fish.

[0114] In this embodiment, a target depth region of interest is first constructed based on a fish instance segmentation mask, concentrating depth calculation and length measurement on the area where the fish target is located, reducing the interference of background areas on the measurement results. Then, maximum connected component preservation and hole filling are performed on the fish instance segmentation mask to remove noise and isolated regions from the segmentation results, obtaining a more complete and stable fish foreground mask. Next, principal component analysis is performed on the target foreground mask to accurately extract the fish's principal axis and its two endpoints, providing a reliable basis for principal axis endpoint localization and length calculation. Then, pixel lengths are calculated based on the two endpoints, and the corresponding depth values ​​are obtained from the depth map, preparing depth data for 3D length calculation. Finally, the 3D length of the fish is calculated based on the two endpoints, the corresponding depth values, and the intrinsic parameters of the left-eye camera, realizing 3D fish length measurement based on binocular vision. Overall, this embodiment combines semantic segmentation foreground with binocular depth information, reducing the impact of background parallax and occlusion areas on fish length estimation, achieving accurate measurement of fish length in 3D space.

[0115] In one optional implementation, step S45, the step of obtaining the depth values ​​corresponding to the two endpoints respectively, includes steps S451 to S453: Step S451: Determine the center point of the main axis of the fish body; In step S451, based on the two endpoints of the fish's main axis extracted in step S43, the midpoint of the line connecting the two endpoints is calculated, and this midpoint is determined as the center point of the fish's main axis. The coordinates of the center point are the average of the coordinates of the two endpoints. The purpose of this step is to determine the center position of the fish's main axis, providing a reference point for subsequent endpoint inward contraction.

[0116] In one example, the coordinates of the two endpoints of the main axis of the fish are (100, 200) and (400, 350), and the coordinates of the center point are (250, 275).

[0117] Step S452: According to the preset shrinkage ratio, each endpoint is shrunk inward along the main axis towards the center point to obtain the corresponding sampling endpoint; It should be noted that "shrinkage" refers to moving the endpoint a certain distance along the main axis towards the center point. The purpose is to move the sampling point away from the fish's boundary, reducing the instability of depth values ​​at the boundary due to parallax noise and occlusion. The shrinkage ratio is the ratio of the moving distance to the distance from the endpoint to the center point, typically ranging from 0 to 0.2.

[0118] In step S452, for each endpoint, the vector from that endpoint to the center point is calculated. This vector is multiplied by a preset shrinkage ratio to obtain a shrinkage offset. Then, the coordinates of that endpoint are added to the shrinkage offset to obtain the shrinkage-adjusted sampling endpoint coordinates. The same operation is performed on both endpoints to obtain two sampling endpoints. The shrinkage ratio can be set according to the fish size, mask boundary quality, and parallax noise. The purpose of this step is to ensure that the sampling points avoid the parallax noise region at the fish boundary, thereby improving the stability and accuracy of endpoint depth sampling.

[0119] In one example, the indentation ratio is 0.08, the coordinates of the first endpoint are (100, 200), the coordinates of the center point are (250, 275), the vector from the center point to the first endpoint is (-150, -75), the indentation offset is (-150×0.08, -75×0.08) = (-12, -6), and the coordinates of the first sampling endpoint are (100+12, 200+6) = (112, 206). The coordinates of the second endpoint are (400, 350), the vector from the center point to the second endpoint is (150, 75), the indentation offset is (12, 6), and the coordinates of the second sampling endpoint are (400-12, 350-6) = (388, 344).

[0120] Step S453: Perform median depth sampling in the neighborhood of each sampling endpoint to obtain the depth value corresponding to each endpoint.

[0121] It should be noted that the neighborhood refers to a local area centered on the sampling endpoint, typically a rectangular or circular region. Median depth sampling involves collecting all valid depth values ​​within the neighborhood, sorting these values ​​by size, and taking the median value as the depth value of the sampling point. Median sampling is highly robust to outliers and can effectively suppress the impact of disparity noise on depth estimation.

[0122] In step S453, for each sampling endpoint, a neighborhood range, such as a square region with an odd number of pixels on each side, is defined centered on that endpoint. All valid depth values ​​within this neighborhood range are obtained from the depth map. These depth values ​​are then sorted, and the median value is taken as the depth value corresponding to that sampling endpoint. The above operation is performed on both sampling endpoints to obtain the depth values ​​corresponding to each endpoint. If the number of valid depth values ​​in the neighborhood is insufficient, the depth value of that endpoint is marked as invalid. The purpose of this step is to reduce the impact of boundary noise on the 3D length through local statistical sampling, thereby improving the robustness and stability of depth sampling.

[0123] In one example, the first sampling endpoint is located at coordinates (112, 206), and its neighborhood is defined as a 7×7 pixel square region, i.e., 49 pixels centered at (112, 206). The effective depth values ​​of these 49 pixels are obtained from the depth map, sorted in ascending order, and the 25th value, i.e., the median, is taken as the depth value corresponding to the first endpoint. The second sampling endpoint is processed in the same way to obtain its corresponding depth value.

[0124] In this embodiment, by determining the center point on the main axis of the fish body, shrinking the endpoints inward toward the center point along the main axis, and performing median depth sampling in the neighborhood of the sampling endpoints, the region with large parallax noise at the fish body boundary is effectively avoided. Furthermore, the impact of noise on depth estimation is reduced through local statistical sampling, thereby improving the stability and accuracy of depth sampling at the fish body endpoints and ultimately enhancing the robustness of the three-dimensional length measurement of the fish body.

[0125] In one optional implementation, after step S40, the fish body length measurement method further includes steps S61-S62: Step S61: Determine the validity of the fish body length; It should be noted that validity determination refers to the process of checking whether the calculated fish length is reasonable and reliable according to preset determination rules. Preset determination rules may include, but are not limited to: whether the target box is close to the image edge, whether the principal axis pixel length of the fish is valid, whether the center parallax is valid, whether the endpoint depth is within a reasonable range, whether the difference between the two endpoint depths is too large, whether the mask constraint depth sampling failed, and whether the measured length is less than the minimum threshold or greater than the maximum threshold, etc.

[0126] In step S61, the three-dimensional length of the fish calculated in step S46 is verified according to preset validity judgment rules. Specifically, one or more of the following judgment conditions can be checked: whether the target box of the fish is close to the image edge; whether the main axis pixel length of the fish is greater than the preset lower limit of the effective pixel length; whether the endpoint depth value is a valid finite value and is within the preset depth range; whether the difference between the two endpoint depth values ​​is within the preset reasonable difference range; whether the depth sampling under the mask constraint successfully obtains a valid depth value; and whether the three-dimensional length of the fish is between the preset minimum length threshold and the maximum length threshold. If all of the above judgment conditions are met or a preset number of conditions are met, the fish length is determined to be valid; otherwise, the fish length is determined to be invalid, and the corresponding invalidation reason is recorded. The purpose of this step is to eliminate length misjudgment caused by abnormal targets and incorrect depths, and to ensure that the output fish length has high reliability.

[0127] In one example, the preset judgment rules include: the fish target bounding box is not close to the image edge, i.e., the edge distance is greater than 5 pixels; the main axis pixel length of the fish is greater than 25 pixels; the depth values ​​of both endpoints are within the range of 1 meter to 10 meters and the difference between them is less than 0.5 meters; and the three-dimensional length of the fish is between 8 centimeters and 55 centimeters. The three-dimensional length of the fish in the current frame is 27.4 centimeters, which meets all the above conditions, and therefore is judged as valid.

[0128] In another example, if the three-dimensional length of the fish is only 3 centimeters, which is less than the preset minimum length threshold of 8 centimeters, it is determined to be invalid, and the reason for invalidity is recorded as "too_short".

[0129] Step S62: Based on the validity determination result, output the corresponding length result and mark the corresponding measurement mode, or output the invalid reason; wherein, the measurement mode includes three-dimensional mode, approximate mode and stable mode.

[0130] It should be noted that the 3D mode refers to the 3D length mode of the fish body, which is directly calculated based on the 3D spatial coordinates of the two endpoints of the fish's main axis. The length value output in this mode is the most accurate and reliable. The approximate mode refers to the approximate length mode estimated based on the pixel length of the fish's main axis and the center depth when the endpoint depth is unavailable but the fish's center depth is valid. The stable mode refers to the historically stable length mode output when the measurement of the current frame is unstable but there is a historically stable length for the same tracked target.

[0131] In step S62, when step S61 determines that the fish length is valid and the depth values ​​at both endpoints are valid, the system outputs the calculated 3D length of the fish in 3D mode and marks the measurement mode as "3D". When the fish length is determined to be valid but the endpoint depths are unavailable while the center depth is valid, the system calculates an approximate length based on the fish's main axis pixel length and center depth, outputs the approximate length, and marks the measurement mode as "approx". When the current frame measurement is invalid but there is a historical stable length for the same tracked target, the system outputs the historical stable length and marks the measurement mode as "stable". When the fish length is determined to be invalid and there is no available historical stable length, the system sets the fish length to an invalid value and outputs the corresponding invalid reason. Invalid reasons may include "near_edge", "bad_disp", "center_depth_out_of_range", "depth_mismatch", "masked_depth_fail", "too_short", or "too_long", etc. The purpose of this step is to provide interpretable length measurement results in complex fish school scenarios, rather than simply outputting a single length value, making it easier for on-site operators to judge the reliability of the length measurement results.

[0132] In one example, the fish body length in the current frame is valid and the three-dimensional length of 27.4 cm is calculated directly based on the depths of the two endpoints. The system outputs a length of 27.4 cm and the measurement mode is marked as "3D".

[0133] In another example, the current frame endpoint depth sampling failed, but the fish's center depth was valid at 2.5 meters. The main axis pixel length of the fish was 335.41 pixels. The system calculated the approximate length according to the approximate length calculation formula as 335.41 × 2.5 divided by 1200 and then multiplied by 100, which is approximately 69.88 centimeters. The system output this approximate length, and the measurement mode was marked as "approx".

[0134] In another example, the fish target in the current frame is close to the edge of the image, the length of the fish is determined to be invalid, and there is no historical tracking record for the target. The system sets the length to an invalid value and outputs the invalid reason as "near_edge".

[0135] In addition, please see Figure 7 , Figure 7A flowchart illustrating the calculation and validity determination of fish body length is provided. Invalid reasons can include, but are not limited to: near edge, bad dissipation, center depth out of range, depth mismatch, masked depth failure, too short, or too long.

[0136] In this embodiment, by determining the validity of the fish length and outputting the corresponding length result and marking the measurement mode or outputting the invalid reason according to the determination result, the technical effect of providing interpretable length measurement results in complex fish school scenarios is achieved. This makes it easier for on-site operators to judge the credibility of the length measurement results and improves the engineering usability and reliability of the system.

[0137] Based on the first embodiment of this application, in the fifth embodiment of this application, the same or similar content as the first embodiment can be referred to the above description, and will not be repeated hereafter. On this basis, the first type of processor is a central processing unit; and / or, the disparity depth calculation adopts a semi-global matching algorithm; and / or, the second type of processor is a neural network processor; and / or, the fish detection and instance segmentation inference adopt an instance segmentation model running under the neural network processor framework.

[0138] It's important to note that the Central Processing Unit (CPU) is the general-purpose computing core in edge computing platforms, adept at handling complex, multi-branch computational tasks, and suitable for executing algorithms such as disparity depth calculation. The Neural Processing Unit (NPU) is a processor specifically optimized for neural network inference computation, characterized by high computing power and low power consumption, suitable for performing inference tasks of deep learning models such as fish detection and instance segmentation. The Semi-Global Matching algorithm is a stereo matching algorithm that calculates disparity maps by aggregating costs along multiple directions, achieving a good balance between computational accuracy and efficiency. The Neural Processing Unit framework refers to the inference runtime environment provided for the Neural Processing Unit, such as the RKNN framework, used to deploy trained deep learning models to the Neural Processing Unit for inference.

[0139] In one implementation, the first type of processor is a central processing unit (CPU), and the disparity depth calculation is performed on the CPU using a semi-global matching algorithm. Specifically, the CPU runs the program code for the semi-global matching algorithm, performs stereo matching calculations on the left and right corrected images, generates a disparity map, and further calculates a depth map. The CPU has a flexible instruction set and strong logic processing capabilities, making it suitable for handling the complex cost aggregation and path optimization calculations in the semi-global matching algorithm.

[0140] In another implementation, the second type of processor is a neural network processor, and fish detection and instance segmentation inference employ an instance segmentation model running within the neural network processor framework. Specifically, the pre-trained instance segmentation model is converted into a format recognizable by the neural network processor, and the model is loaded and inference is performed within the neural network processor framework. The neural network processor contains a large number of multiplication and accumulation units, enabling it to efficiently perform convolution operations and matrix multiplication operations in convolutional neural networks, significantly accelerating the fish detection and instance segmentation inference process.

[0141] In another implementation, the two implementation methods described above can be used in combination. Specifically, the disparity depth calculation is performed on the central processing unit using a semi-global matching algorithm, while the fish body detection and instance segmentation inference are performed on the neural network processor using an instance segmentation model running under the neural network processor framework. The two are executed in parallel on different processors, thereby achieving the collaborative utilization of heterogeneous computing resources.

[0142] In this embodiment, by designating the first type of processor as a central processing unit, using a semi-global matching algorithm for disparity depth calculation, using a neural network processor as the second type of processor, and using an instance segmentation model running under the neural network processor framework for fish detection and instance segmentation inference, the division of labor and algorithm implementation of heterogeneous computing resources are clarified, providing a specific implementation plan for making full use of the computing advantages of different types of processors on the edge computing platform.

[0143] Furthermore, in one example, internal tests were performed on fixed stereo video and real-time RTSP video streams to verify the effectiveness of the method of the present invention. Performance comparisons of different optimized versions are shown in Table 1.

[0144] Table 1 Performance Comparison of Different Optimized Versions

[0145] Table 1 shows the number of video frames per second that different versions of the method of the present invention can process when processing the same fixed video.

[0146] In addition, please see Figure 8 , Figure 8The figure shows a comparison of the average single-frame processing time between the normal serial operation mode and the multi-threaded enhanced version. As can be seen from the figure, the multi-threaded enhanced version of this application has a significant reduction in processing time compared to the skip2 normal serial operation mode.

[0147] For example, to help understand the implementation process of the fish length measurement method obtained by combining this embodiment with any of the above embodiments, please refer to... Figure 9 , Figure 9 A general technical flowchart for a method of measuring fish body length is provided, specifically: First, acquire the stereo video stream; the stereo video stream contains the left eye image and the right eye image.

[0148] Then, the binocular calibration parameters are read and image correction is performed on the left and right eye images. The system constructs a correction mapping table based on camera intrinsic parameters, distortion parameters, binocular extrinsic parameters, focal length, principal point, and binocular baseline B, and performs remapping correction processing on the left and right eye images to make the left and right images satisfy the epipolar correction relationship, thus obtaining the corrected left and right eye images.

[0149] Next, fish detection and instance segmentation are performed on the left-eye corrected image. The system converts the left-eye corrected image to RGB format and inputs it into a fish detection or instance segmentation model deployed on RKNN or other neural network inference frameworks. In a preferred embodiment, the model is a YOLOv11-seg fish instance segmentation model, which obtains the fish category, confidence score, bounding box, and instance segmentation mask.

[0150] Next, within a complete computation frame, the first thread and the second thread are generated: First thread: Fish body detection and instance segmentation processing; The second thread: disparity calculation and depth map generation; which includes performing grayscale and texture enhancement processing on the left and right eye corrected images, and calculating the depth map based on the full-resolution equivalent disparity.

[0151] Then, within a complete computation frame, the disparity depth calculation worker thread (i.e., the second thread) is started, while the main thread executes the neural network fish detection and segmentation inference thread (i.e., the first thread). Before entering length post-processing, the main thread (i.e., the first thread) waits for the disparity depth calculation thread (the second thread) to complete, ensuring that the full-resolution equivalent disparity map and depth map have been generated. In one feasible implementation, the disparity depth calculation thread executes SGBM / depth, and the main thread executes RKNN inference.

[0152] Next, mask processing of the main axis endpoints: construct the target depth ROI (region of interest) based on the segmentation mask of the fish body instance, perform maximum connectivity preservation and hole filling processing on the mask, and use PCA to extract the main axis endpoints of the fish body within the mask region.

[0153] Then, the three-dimensional length is calculated by combining the median depth sampling of the endpoints, and the pattern is determined.

[0154] Next, the frequency division calculation strategy is set according to the frequency division parameters, frequency division reuse is performed, and the cached results are updated. When the frequency division parameter compute_interval=N, the system performs a complete detection, disparity, and length measurement every N frames, and the remaining frames reuse the annotation results of the previous complete calculation frame for display.

[0155] Finally, the system displays and outputs structured data.

[0156] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the fish length measurement method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0157] This application also provides a fish body length measurement system; please refer to... Figure 10 The fish body length measurement system includes: Image acquisition module 10 is used to acquire the left and right eye images captured by the binocular camera; Image correction module 20 is used to perform correction processing on the left eye image and the right eye image to obtain a corrected left eye image and a corrected right eye image; The multi-threaded execution module 30 is used to perform disparity depth calculation based on the left eye correction image and the right eye correction image in a first thread to generate a depth map, and to perform fish body detection and instance segmentation inference on the left eye correction image in a second thread to obtain a fish body instance segmentation mask; wherein, the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor. The fish body length calculation module 40 is used to calculate the fish body length based on the segmentation mask of the fish body instance and the depth map.

[0158] The fish length measurement system provided in this application, employing the fish length measurement method in the above embodiments, can solve the technical problem of how to reorganize task scheduling within a complete computation frame, enabling heterogeneous resources to work together efficiently and reducing single-frame processing latency. Compared with the prior art, the beneficial effects of the fish length measurement system provided in this application are the same as those of the fish length measurement method provided in the above embodiments, and other technical features of the fish length measurement system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0159] This application provides a fish length measuring device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the fish length measuring method in the above embodiment 1.

[0160] The following is for reference. Figure 11 The diagram illustrates a structural schematic of a fish length measuring device suitable for implementing embodiments of this application. The fish length measuring device in this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 11 The fish length measuring device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0161] like Figure 11As shown, the fish length measuring device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the fish length measuring device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the fish length measuring device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows fish length measuring devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0162] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0163] The fish length measuring device provided in this application, employing the fish length measuring method in the above embodiments, can solve the technical problem of how to reorganize task scheduling within a complete computing frame, enabling heterogeneous resources to work together efficiently and reducing single-frame processing latency. Compared with the prior art, the beneficial effects of the fish length measuring device provided in this application are the same as those of the fish length measuring method provided in the above embodiments, and other technical features in this fish length measuring device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0164] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0165] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0166] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the fish body length measurement method in the above embodiments.

[0167] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0168] The aforementioned computer-readable storage medium may be included in the fish length measuring device; or it may exist independently and not assembled into the fish length measuring device.

[0169] The aforementioned computer-readable storage medium carries one or more programs that, when executed by the fish length measuring device, cause the fish length measuring device to: acquire left and right eye images captured by a binocular camera; perform correction processing on the left and right eye images to obtain left-eye corrected images and right-eye corrected images; perform multi-threaded execution steps, within a complete calculation frame, in the first thread performing disparity depth calculation based on the left-eye corrected images and right-eye corrected images to generate a depth map, and in the second thread performing fish body detection and instance segmentation inference on the left-eye corrected image to obtain a fish body instance segmentation mask; wherein the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor; and calculate the fish length based on the fish body instance segmentation mask and the depth map.

[0170] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0171] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0172] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0173] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described fish length measurement method. This solves the technical problem of how to reorganize task scheduling within a complete computation frame, enabling heterogeneous resources to work efficiently together and reducing single-frame processing latency. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the fish length measurement method provided in the above embodiments, and will not be repeated here.

[0174] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the fish body length measurement method described above.

[0175] The computer program product provided in this application can solve the technical problem of how to reorganize task scheduling within a complete computing frame, enabling heterogeneous resources to work together efficiently and reducing single-frame processing latency. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the fish length measurement method provided in the above embodiments, and will not be repeated here.

[0176] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for measuring the body length of a fish, characterized in that, The method includes: Acquire a binocular video stream; the binocular video stream includes multiple frames of left-eye images and multiple frames of right-eye images; The left and right eye images are corrected to obtain a corrected left eye image and a corrected right eye image. The multi-threaded execution steps are as follows: within a complete calculation frame, in the first thread, disparity depth is calculated based on the left eye correction image and the right eye correction image to generate a depth map, and in the second thread, fish body detection and instance segmentation inference are performed on the left eye correction image to obtain a fish body instance segmentation mask; wherein, the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor. The length of the fish is calculated based on the segmentation mask of the fish instance and the depth map.

2. The method as described in claim 1, characterized in that, The step of calculating disparity depth based on the left-eye corrected image and the right-eye corrected image in the first thread to generate a depth map includes: The left and right eye corrected images are subjected to grayscale conversion and texture enhancement processing to obtain a left eye grayscale image and a right eye grayscale image; The left-eye grayscale image and the right-eye grayscale image are scaled according to a preset scaling factor s to obtain a scaled left-eye grayscale image and a scaled right-eye grayscale image; where 0 <s≤1; Calculate a small-scale disparity map based on the zoomed grayscale image of the left eye and the zoomed grayscale image of the right eye; The small-scale disparity map is filtered, and the filtered small-scale disparity map is restored to the full resolution size to obtain the restored small-scale disparity map. Based on the scale factor s, the recovered small-scale disparity map is converted into a full-resolution equivalent disparity map; The depth map is calculated based on the full-resolution equivalent disparity map.

3. The method as described in claim 2, characterized in that, The step of converting the recovered small-scale disparity map into a full-resolution equivalent disparity map based on the scale factor s includes: Based on the disparity values ​​of each pixel in the restored small-scale disparity map and the scale factor, calculate the full-resolution equivalent disparity value of each pixel. Correspondingly, the step of calculating the depth map based on the full-resolution equivalent disparity map includes: The depth value of each pixel is calculated based on the full-resolution equivalent parallax value of each pixel, the horizontal focal length of the left eye camera, and the binocular baseline, to obtain the depth map.

4. The method as described in claim 1, characterized in that, The method further includes: Obtain the frequency division parameters; Based on the current frame number and the frequency division parameter, determine whether the current frame is a complete calculation frame; If the current frame is the complete calculation frame, then the multi-threaded execution step is executed, and the calculation result is updated to the cached result; If the current frame is not the complete calculation frame, the cached result is reused for display, and the writing of measurement records to the structured result file is abandoned.

5. The method as described in claim 1, characterized in that, The step of calculating the length of the fish body based on the segmentation mask of the fish body instance and the depth map includes: Construct the target depth region of interest based on the segmentation mask of the fish body instance; Perform maximum connectivity preservation and hole filling processing on the fish body instance segmentation mask to obtain the target foreground mask; Principal component analysis was performed on the target foreground mask to extract the fish body principal axis and the two endpoints of the fish body principal axis; Calculate the pixel length based on the two endpoints; Based on the depth map, obtain the depth values ​​corresponding to the two endpoints respectively; The three-dimensional length of the fish is calculated based on the two endpoints, the corresponding depth values, and the intrinsic parameters of the left eye camera.

6. The method as described in claim 5, characterized in that, The step of obtaining the depth values ​​corresponding to the two endpoints respectively includes: Determine the center point of the main axis of the fish body; According to a preset shrinkage ratio, each endpoint is shrunk inward toward the center point along the main axis to obtain the corresponding sampling endpoint; Median depth sampling is performed in the neighborhood of each sampling endpoint to obtain the depth value corresponding to each endpoint.

7. The method as described in claim 5, characterized in that, The method further includes: The validity of the fish body length is determined; Based on the validity determination result, the corresponding length result is output and the corresponding measurement mode is marked, or the invalid reason is output; wherein, the measurement mode includes three-dimensional mode, approximate mode and stable mode.

8. The method according to any one of claims 1 to 7, characterized in that, The first type of processor is a central processing unit; and / or, the disparity depth calculation adopts a semi-global matching algorithm; and / or, the second type of processor is a neural network processor; and / or, the fish detection and instance segmentation inference adopt an instance segmentation model running under the neural network processor framework.

9. A fish body length measurement system, characterized in that, The system includes: The image acquisition module is used to acquire the left and right eye images captured by the binocular camera; An image correction module is used to correct the left eye image and the right eye image to obtain a corrected left eye image and a corrected right eye image; A multi-threaded execution module is used to perform disparity depth calculation based on the left eye correction image and the right eye correction image in a first thread to generate a depth map, and to perform fish body detection and instance segmentation inference on the left eye correction image in a second thread to obtain a fish body instance segmentation mask; wherein, the first thread and the second thread are executed in parallel, the first thread is executed by a first type of processor, and the second thread is executed by a second type of processor. The fish body length calculation module is used to calculate the fish body length based on the segmentation mask of the fish body instance and the depth map.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the fish body length measurement method as described in any one of claims 1 to 8.