Underwater micropterus salmoides body length and mass measurement method based on YOLOv8 and monocular depth estimation

By applying YOLOv8 and monocular depth estimation technology in aquaculture, real-time, non-contact body length and weight measurement of largemouth bass is achieved, solving the problems of low accuracy and labor-intensiveness in traditional methods, and supporting intelligent aquaculture management.

CN120107769APending Publication Date: 2025-06-06HUAZHONG AGRI UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510166508.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In traditional aquaculture, obtaining the body length and weight information of largemouth bass requires manual operation, resulting in fish stress, low accuracy and labor-intensive, which cannot meet the needs of intelligent breeding.

Method used

The coordinates of the fish head and tail are obtained through target detection and key point detection, combined with the monocular depth estimation calculation method, and the fish's mass is estimated through power function regression model to achieve real-time, non-contact body length and weight measurement.

Benefits of technology

Real-time, non-contact body length and weight measurement of largemouth bass is realized, which improves the accuracy and efficiency of measurement, reduces stress on fish, and supports intelligent management under the factory-based circulating water breeding model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107769A_ABST
    Figure CN120107769A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater micropterus salmoides body length and mass measurement method, and relates to the technical field of aquaculture. Comprising the following steps: 1) target detection and key point detection: performing target detection and key point detection on underwater largemouth bass by using a YOLOv8 model to obtain key point coordinate information of a fish head and a fish tail; and 2) obtaining depth information: obtaining the depth information of the underwater image through a monocular depth estimation algorithm, and calculating the body length of the largemouth bass in combination with a key point detection result. And 3) mass estimation: based on the body length information, estimating the mass of the micropterus salmoides through a power function regression model. And 4) model deployment: deploying the model in a recirculating aquaculture monitoring system, and displaying and recording the body length and quality information of the micropterus salmoides in real time. According to the method, the body length and the body weight of the largemouth bass can be obtained in real time in a non-contact mode, the method is suitable for a factory-like circulating water culture mode, data support can be provided for culture management, and the practical value and the popularization prospect are high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aquaculture, and in particular to a method for estimating the length and mass of underwater largemouth bass based on YOLOv8 and monocular depth estimation. The method is suitable for a factory-scale recirculating aquaculture mode, can monitor the growth status of largemouth bass in real time, and provide data support for aquaculture management. Background Art

[0002] Largemouth bass ( Micropterus salmoides ) is a common farmed fish. It is ferocious by nature, likes to chase, and has a large appetite. When fed improperly, it is easy to become hungry, and then the big ones will eat the small ones. During breeding management, grading is generally carried out every two months, and the basis for dividing the ponds is the growth information of the fish at different growth stages. Body length and weight information are key growth data, representing the growth status of the fish. By judging the growth status of fish, possible problems in the breeding process can be discovered, and appropriate adjustments can be made to the parameters of the breeding equipment in a timely manner, such as the frequency of the circulation pump, the light cycle and light intensity of the lighting system, and the oxygen supply pressure of the oxygen generator. The growth information of fish can also provide basic data for accurate feeding of feed, adjustment of the flow rate of aquaculture water inlet, and estimation of aquaculture output.

[0003] The traditional method of obtaining body length and weight data is to use manual instruments to weigh the mass, measure the body length with a ruler and other contact methods. The fish needs to be salvaged from the water, which will cause stress to the fish and cause damage to the fish. The accuracy cannot be guaranteed due to human subjective judgment, and a large amount of labor is required. Decisions are subject to the subjective influence of managers and their lags, and are no longer suitable for the increasingly intelligent breeding model.

[0004] Using non-contact methods to automatically obtain the length and weight data of largemouth bass is an important research topic. With the development of artificial intelligence technology and Internet of Things technology, the construction of intelligent recirculating aquaculture systems has become a new development direction, and modern information technology provides support for the refined management of aquaculture. In short, the land-based factory recirculating aquaculture model is one of the directions for the development of sustainable aquaculture, and the automatic measurement technology of largemouth bass size is one of the key technologies to realize unmanned aquaculture in recirculating water systems.

[0005] Largemouth bass in the recirculating aquaculture model grow in aquaculture tanks with clean water. Due to their habits, they are mostly located in the lower layer of the water body, and there are intervals between individuals, which provides a relatively good environment for visual measurement of fish length. The method of measuring fish body vital signs information with binocular vision has been widely used, but considering the high cost of binocular underwater cameras, it is not conducive to market popularization. Monocular cameras are inexpensive and widely used in monitoring systems. The accuracy of the algorithm for estimating depth with only a single image is getting higher and higher. Combined with the real-time changing data transmitted to the intelligent decision-making platform, it can make more detailed evaluations and decisions on aquaculture management. Therefore, this study takes largemouth bass in the recirculating aquaculture model as the research object, uses deep learning and monocular camera depth estimation methods to realize target detection and body length and weight measurement of underwater largemouth bass, and combines the monitoring platform provided by Hikvision to deploy the model to the terminal device for real-time display and recording, which is of great significance to promoting the development of recirculating aquaculture towards intelligence. Summary of the invention

[0006] The purpose of the present invention is to provide a method for measuring the body length and mass of underwater largemouth bass based on YOLOv8 and monocular depth estimation, which can obtain the body length and weight information of largemouth bass in real time and non-contact, and provide data support for aquaculture management.

[0007] To achieve the above object, the present invention provides the following technical solutions: The present invention provides a method for measuring the length and mass of underwater largemouth bass based on YOLOv8 and monocular depth estimation, comprising the following steps: 1) Target detection and key point detection: The YOLOv8 model is used to perform target detection and key point detection on underwater largemouth bass to obtain the key point coordinate information of the fish head and tail.

[0008] 2) Depth information acquisition: The depth information of the underwater image is obtained through a monocular depth estimation algorithm, and the body length of the largemouth bass is calculated by combining the key point detection results.

[0009] 3) Mass estimation: Based on body length information, the mass of largemouth bass was estimated using a power function regression model.

[0010] 4) Model deployment: The model is deployed into the recirculating aquaculture monitoring system to display and record the length and mass information of largemouth bass in real time.

[0011] The YOLOv8 model is trained by the following steps: 1) Collect image data of underwater largemouth bass, perform data enhancement, and generate training and validation sets; 2) Use transfer learning methods to train based on pre-trained models and optimize model parameters; 3) Evaluate model performance through the loss function and adjust the learning rate and optimizer parameters until the model converges.

[0012] The monocular depth estimation algorithm adopts the Depth Anything model and is implemented by the following steps: 1) Using a small amount of data with deep labels to assign pseudo labels to unlabeled datasets; 2) Combine the labeled dataset and the pseudo-labeled dataset to train the model and improve the generalization ability of the model; 3) Fine-tune the model by measuring depth to improve the accuracy of the model.

[0013] The power function regression model is used to estimate the mass of largemouth bass, and its formula is:

[0014] Among them, W is the mass of the fish, L is the body length of the fish, and a and b are regression parameters.

[0015] Furthermore, the method also includes calibrating the underwater camera and correcting image distortion, the specific steps are: 1) Use Zhang Zhengyou calibration method to calibrate the monocular camera to obtain the camera's intrinsic parameter matrix, extrinsic parameter matrix and distortion coefficient; 2) Use the OpenCV library to correct image distortion and reduce errors in the 3D reconstruction process.

[0016] Furthermore, the method also includes converting the model into ONNX format and deploying it on a terminal device using an ONNX Runtime inference engine to achieve real-time reasoning and display.

[0017] Furthermore, the method performs real-time body length and mass estimation by the following steps: 1) Obtain real-time video stream from the monitoring system and extract images frame by frame; 2) Input the image into the YOLOv8 model and the monocular depth estimation model to obtain the key point coordinates and depth information of the fish head and tail; 3) Calculate the three-dimensional distance between the fish head and tail to get the estimated value of body length; 4) Calculate the mass estimate based on the body length estimate using a power function regression model.

[0018] Furthermore, the method further includes performing error analysis on the estimation result, the specific steps are: 1) Compare the manual measurement results with the model estimation results and calculate the relative errors of body length and mass; 2) Reduce estimation errors and improve model accuracy by adjusting model parameters and optimizing algorithms.

[0019] Furthermore, the method is applicable to the factory-scale recirculating aquaculture model, and can monitor the growth status of largemouth bass in real time, providing data support for aquaculture management.

[0020] The beneficial effects of the present invention are: The present invention provides a method for measuring the length and mass of underwater largemouth bass based on YOLOv8 and monocular depth estimation, which can obtain the length and weight information of largemouth bass in real time and contactlessly. The method is suitable for the factory-scale recirculating aquaculture mode, can provide data support for aquaculture management, and has high practical value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 : Underwater images of largemouth bass in a recirculating aquaculture environment.

[0022] Figure 2 : Data annotation interface of labelme annotation tool.

[0023] Figure 3 : Underwater image of largemouth bass after shading and noise processing.

[0024] Figure 4 : The curve of loss changing with the number of training rounds during training.

[0025] Figure 5 : Change curve of target detection evaluation index under different confidence levels.

[0026] Figure 6 : The curve of the change of the average accuracy of target detection in the number of training rounds.

[0027] Figure 7 : The curve of the change of the average accuracy of key point detection in the number of training rounds.

[0028] Figure 8 : Underwater images of the calibration plate at different angles.

[0029] Fig. 9 : Monocular camera calibration error.

[0030] Fig.10 : Relationship diagram between the actual positions of the monocular camera and the calibration plate.

[0031] Fig.11 : Image comparison of largemouth bass underwater images before and after correction.

[0032] Fig.12 : Capture the real-time prediction result image. DETAILED DESCRIPTION

[0033] The present invention is further described in detail below in conjunction with specific embodiments.

[0034] The test process of the embodiment is as follows: 1. Image acquisition: In the factory-scale recirculating aquaculture mode, an underwater webcam was used to collect image data of largemouth bass with a resolution of 2560×1440 and a frame rate of 25fps.

[0035] Real-time video is displayed through the Hikvision platform, and sufficient underwater largemouth bass image data is obtained to build a deep learning dataset.

[0036] 2. Dataset annotation: Use the labelme annotation tool to annotate the image and mark the key point coordinate information of the fish head and tail.

[0037] Data augmentation was performed on 500 images to generate a dataset of 2500 images, which was divided into training and validation sets in a ratio of 8:2.

[0038] 3. Model training: Model training is performed on the cloud platform Featurize, using the transfer learning method based on the pre-trained model.

[0039] During the training process, the changes in the loss function are recorded, and the learning rate and optimizer parameters are adjusted until the model converges.

[0040] 4. Experimental results: The training set loss function gradually decreases, and after 50 rounds of training, the total training loss drops to 0.8184.

[0041] The mean average precision (mAP) for object detection and keypoint detection is 0.9942 and 0.993 respectively.

[0042] The non-maximum suppression (NMS) method is used to remove prediction boxes with confidence levels lower than 0.25 to achieve accurate positioning and identification of largemouth bass.

[0043] 5. Camera calibration and image correction: The monocular camera is calibrated using Zhang Zhengyou calibration method to obtain the camera's intrinsic parameter matrix, extrinsic parameter matrix and distortion coefficient.

[0044] The OpenCV library is used to correct image distortion and reduce errors in the 3D reconstruction process.

[0045] 6. Monocular depth estimation and body length measurement: The Depth Anything model is used to obtain the depth information of the image, and the body length of the largemouth bass is calculated based on the key point detection results.

[0046] Estimation of mass of largemouth bass by power regression model.

[0047] 7. Model deployment and real-time estimation: Convert the model to ONNX format and deploy it on the terminal device using the ONNX Runtime inference engine.

[0048] The real-time video stream frame rate reaches 25fps, which can display and record the body length and quality information of largemouth bass in real time.

[0049] 8. Error analysis: Comparing the manual measurement results with the model estimation results, the average relative error of body length was 6.4%, and the average relative error of mass was 14.8%.

[0050] By adjusting model parameters and optimizing algorithms, estimation errors can be reduced and model accuracy can be improved.

[0051] Example 1. Detection of underwater target information and key point information of largemouth bass Deploying the model to a recirculating aquaculture monitoring system requires high speed. The purpose of fish length estimation is to provide a reference for feeding and management by aquaculture operators, and accuracy is also required. Therefore, the YOLOv8 method with a more balanced speed and accuracy is chosen to train the underwater largemouth bass target detection and key point detection model.

[0052] YOLOv8 is an open source project that supports image classification, object detection, and instance segmentation tasks. The YOLO series of algorithms have been widely used in tasks such as target detection. The main innovations of YOLOv8 are the construction of a new backbone network, the change of the detection head, and the introduction of a new loss function. Its main calculation process is as follows: (1) After the image is input, the image is divided into several grids through image preprocessing, and object detection is performed on each grid. The bounding box and confidence are predicted for each grid; (2) The category probability of the target and the location information of the key points are predicted in each bounding box; (3) The confidence and category probability score of each bounding box are calculated; (4) Overlapping bounding boxes are deleted through non-maximum suppression (NMS), and only the box with the highest confidence value is retained; (5) The target detection and key point detection results are output, which include the bounding box coordinates, category, confidence, and coordinate information of the key points in the box.

[0053] 1. Real-time image acquisition of California bass 1.1 Image acquisition platform The research object is the largemouth bass cultured in an indoor factory-scale circulating water mode. Each culture barrel is equipped with an underwater network camera to form an intelligent underwater monitoring system, and the real-time video is displayed through the Hikvision platform. Relying on this platform, sufficient underwater largemouth bass image data can be obtained to build a deep learning dataset.

[0054] The entire image acquisition unit includes a breeding pond, a monocular camera, a fixed bracket, a biological lighting lamp, etc. The diameter of the breeding pond is 2m, the effective water depth is 0.8m, and the material is fiberglass, which is not easy to reflect light. The camera model is underwater HIKVISION-Model 2, which supports wipers to ensure clear images. The image resolution is 2560×1440, the frame rate is 25fps, the output format is HEVC, and the camera focal length is 2mm. The light source is a biological lighting lamp, which is installed 50cm above the center of the breeding pond.

[0055] 1.2 Image acquisition process 50 largemouth bass were placed in the breeding pond in advance, and real-time video was recorded by Hikvision NVR with a resolution of 2560×1440@25fps. The light cycle of the biological lighting source was 8:00-20:00 for a total of 12 hours. 24 hours of video was collected, including 12 hours of MP4 video under the built-in light source and 12 hours of MP4 video under supplementary light. The video was decoded by PotPlayer professional video decoder, and the image data with clear target was manually screened and intercepted. Each image contains about 5 largemouth bass. Some of the collected images are shown in the figure below. Figure 1 shown.

[0056] 1.3 Dataset Annotation Use labelme to annotate. This tool supports rectangle, polygon, point and line segment annotation. After annotation, a json file with the same name as the image can be generated. The annotation interface is as follows: Figure 2 As shown. This article uses a rectangular box to annotate target detection, and uses a point tool to annotate key points. Only complete fish including the head and tail are labeled. The rectangular box records the location information of the fish, and the point records the location information of the head and tail. The label of the largemouth bass is jiazhoulu, and the labels of the head and tail are head and tail respectively. The json file records the normalized upper left and lower right corner coordinate information of the rectangular box. The point coordinates in the rectangular box represent the location information of the end point of the fish snout and the midpoint of the base of the caudal fin. A total of 500 images containing largemouth bass were annotated.

[0057] 1.4 Data enhancement and dataset construction Data augmentation is a method to expand existing small sample images and directly and effectively solve the problem of insufficient samples. Slight changes can make the neural network regard them as different samples, thereby improving the robustness of the model. Since the actual farming scenes are complex and most of the time they are in low-light and high-noise environments, it is necessary to process the images with light, dark, Gaussian noise, and salt and pepper noise. This paper uses four methods, namely brightness enhancement, brightness reduction, Gaussian noise, and salt and pepper noise, to transform a small sample of 500 images. During the processing, a txt annotation information file with the same name as the processed image is generated. Figure 3 Results for different treatments are shown.

[0058] The number of images after data enhancement is added to the original images to obtain 5 times the amount of data. The data set is increased from 500 images to 2500 images, which is enough to train an ideal model. Use the data shuffle tool to randomly shuffle the images, and then divide them into training set and validation set in a ratio of 8:2. Put the txt annotation files into the corresponding folders to obtain two folders named images and labels, which store images and annotation files respectively. Each folder contains two subfolders, train and val, in which the data correspond one to one. The two folders together constitute the largemouth bass image dataset.

[0059] 2. Model Training 2.1 Preprocessing Image preprocessing is the process of processing images of different resolutions into the same size through cropping and other methods. Preprocessing can speed up the convergence of the loss function during training. The image size in the dataset is 2560×1440 pixels, and the input image size is set to 640×640 pixels. First, the scaling ratio is calculated. Due to the difference in length and width, the scaling ratios in both directions are calculated separately, and the smaller one is selected as its scaling ratio. Therefore, the selected scaling ratio is 0.25. The length and width of the image after scaling are calculated to be 640×360 pixels. Finally, the blank part outside the image is filled with 0 to obtain an image of a size that the network can accept. In order to correspond to the input size of the yolov8 model, the image is first converted to RGB format, and after standardization, one dimension is added to convert it into a four-dimensional tensor format.

[0060] 2.2 Training starts The model training was run on the cloud platform Featurize. Featurize provides a configured instance environment and high-performance servers. The configuration of this training is a 14-core AMD EPYC 7453 processor, a graphics card is an RTX 4090 (24GB video memory), and a memory of 64.4GB. A cloud disk is used to store data sets and other files required for training. The cloud disk space is 50GB. The instance system environment is a development version system of the Linux kernel. The deep learning framework environment is PyTorch2.0.1+CUDA11.8, and Python3.10.12 is used to run the program. The training method uses transfer learning and starts training under the parameters of the pre-trained model. Transfer learning is a machine learning strategy that can use a library of already trained models to reduce the amount of training data required for the target task, improve learning efficiency, and in many cases improve the performance of the model.

[0061] 3. Experimental results and analysis 3.1 Changes in the training set loss function The training set training round epochs is 50, 16 samples are taken for each training, the initial learning rate lr0 is 0.001, the optimizer uses the AdamW optimizer, its momentum is set to 0.999, the weight decay is set to 0.0005, and other parameters use the default settings of yolov8. The loss is recorded during the training process, and the loss value change curve of the training set is as follows Figure 4 As shown in the figure. The initial value of the overall loss is low. Thanks to the use of the pre-training model, the loss converges quickly, but the overall loss decreases more and more slowly. The posture loss weight decreases significantly slower after the 15th round, and the key point loss weight decreases significantly slower after the 20th round. After 50 rounds of training, the total training loss drops to 0.8184, the bounding box loss is 0.3325, the classification loss is 0.1929, the posture loss is 0.0102, and the key point loss is 0.0048.

[0062] 3.2 Object Detection and Keypoint Detection The results of target detection usually include the location coordinates, category label and confidence score of each detected target. The location information of the target is determined by the coordinates of the upper left corner and the lower right corner of the bounding box. The quality of a target detection model can be evaluated from three aspects: positioning accuracy, classification accuracy and running speed of the detection results. The evaluation index of positioning accuracy is the intersection over union (IoU) of the predicted box and the labeled box. The classification accuracy can be evaluated by accuracy, precision, recall rate, PR curve, AP (Average Precision) and mAP (meanAverage Precision). AP is the area of ​​the graph enclosed by the PR curve and the coordinate axis, and mAP represents the average value of AP. The target detection process records the change curves of precision, recall, PR and the harmonic mean of the two under different confidence levels, such as Figure 5 As shown in the figure, the accuracy of the model trained this time approaches 1 as the confidence requirement increases. The recall rate decreases significantly when the confidence is around 0.9. The area enclosed by the PR curve and the coordinate axis is close to 1. It can be seen that the model has a higher accuracy when the confidence is 0.9. The running speed is mainly based on the number of frames predicted in real time for the network video stream, and its speed is affected by multiple factors such as the performance of the terminal device, the inference engine, and whether the model is lightweight.

[0063] The result of key point detection contains the coordinates and category information of the feature points in the target detection frame. The model uses a common evaluation index to measure the performance of the model, and calculates the evaluation index mAP based on the key point similarity OKS (Object Keypoint Similarity).

[0064] 3.3 Results Analysis In the process of validating the mean average precision (mAP) of object detection and keypoint detection, 500 test set images were used. Figure 6 The figure describes the trend of the object detection mAP changing with the number of training rounds under different IoU values. It can be observed that when the IoU is 0.5, the mAP tends to be stable after the 25th round. After 50 rounds of training, the final mAP is 0.9942. When the IoU is 0.5-0.95, the mAP gradually increases. After 50 rounds of training, the final mAP is 0.945.

[0065] Figure 7The following figure describes the trend of the key point detection mAP curve with the number of training rounds under different I / O ratio values. It can be observed from the figure that after 32 rounds of training, the mAP value tends to be stable under different I / O ratio values. After 50 rounds of training, when the I / O ratio is 0.5, the final mAP is 0.994, and when the I / O ratio is 0.5-0.95, the final mAP is 0.993.

[0066] The recirculating aquaculture system is a high-density farming model. Often, many fish appear in a camera image at the same time, which inevitably results in occlusion and incomplete display of the fish body. Therefore, this model only recognizes and detects fish with complete heads and tails in the image. By using the non-maximum suppression (NMS) method, the prediction boxes with confidence values ​​lower than 0.25 and intersection-over-union ratios lower than 0.7 are removed, which enables the positioning and recognition of largemouth bass.

[0067] 2. Distortion correction of underwater images of largemouth black bass The present invention studies an algorithm for estimating the body length of largemouth bass based on image detection and three-dimensional reconstruction, and the imaging effect of the camera is the basis of all image processing. Compared with the above-water camera, the underwater camera has greater distortion during the imaging process. This is because the density of water is much higher than that of air, resulting in a large deviation between the refractive index of underwater light and the refractive index in the air. Therefore, it is necessary to calibrate the underwater camera. The accuracy of the calibration directly affects the accuracy of the later body length estimation, so image distortion correction is a very critical link.

[0068] 1. Camera calibration and image correction Camera calibration can obtain the camera's intrinsic and extrinsic matrix and distortion coefficients. The distortion coefficients are used to correct the intrinsic matrix, and then correct the image distortion. The intrinsic matrix and extrinsic matrix are used to convert the coordinates of the world coordinate system point and the pixel coordinate system point. This section uses Zhang Zhengyou's calibration method to calibrate the monocular camera and write the parameters into OpenCV to achieve image distortion correction.

[0069] 1.1 Calibration method DLT solves the linear equations by the least squares method to obtain the camera's internal and external parameter matrix. It is better used in cases where there are obvious signs such as traffic signs and road markings (Zheng Baofeng et al. 2018, Shi et al. 2021), but its accuracy is not high. Zhang Zhengyou's calibration method obtains the world coordinates and image coordinates of the corner points through the known chessboard. It only requires multiple calibration plate images from different perspectives, and then solves the homography matrix to calculate the internal and external parameter matrix. It is simple to operate, has high accuracy, and has strong robustness. The purpose of this application scenario is to accurately detect the coordinates of key points. There are requirements for accuracy and there are no obvious 2D-3D corresponding point pairs in the scene. Therefore, Zhang Zhengyou's calibration method is used to calibrate the monocular camera.

[0070] 1.2 Monocular Camera Calibration Water and air are two media through which light passes. The refractive index of light passing through them is quite different. This paper uses a calibration plate made of waterproof material for underwater calibration. The calibration plate is made of an alumina checkerboard on a float glass substrate with a size of 400mm×300mm. It contains 12×9 black and white grids with a side length of 30mm and an accuracy of ±0.01mm. The network camera recording function is turned on. By posing underwater for 10 minutes, video stream data with a resolution of 2560×1440 and a frame rate of 30 frames is obtained. 14 images are selected for calibration. The images are as follows: Figure 8 shown.

[0071] Matlab provides a monocular camera calibration toolbox Camera Calibrator. Load the selected image, set the parameters, and perform focus detection. Fig. 9 is the calibration error diagram, Fig.10 This is the position relationship diagram of the camera and the calibration plate.

[0072] The camera's intrinsic parameter matrix and other parameters are obtained through calibration, as shown in Table 1: Table 1 Camera properties

[0073] 1.3 Writing calibration parameters into OpenCV and implementing distortion correction Matlab calibration obtains the camera's internal and external parameter matrix and distortion parameters. Correction refers to using the distortion coefficient to restore the true position of the coordinate point and reduce the error in the 3D reconstruction process. Fig.11 As shown in (a), the image taken by the camera has obvious radial distortion. The new intrinsic parameter matrix of the camera after correction can be obtained through the getOptimalNewCameraMatrix function in OpenCV. The mapping table is generated by the initUndistortRectifyMap function, and the image can be remapped to the corrected image 11 (b) through the remap function.

[0074] 3. Application of Monocular Depth Estimation and Size Measurement of Largemouth Bass The process of factory-scale recirculating aquaculture requires regular acquisition of fish length and weight information. Combined with the information on the amount of stocking, the total amount of organisms in the breeding stage can be calculated, and then the output can be estimated, providing a reference for subsequent sales and processing production. At the same time, during the entire breeding cycle of recirculating aquaculture of largemouth bass, screening is required. Screening can ensure that the facilities for fish activities have sufficient space and flowing water sources. The timing of screening is generally based on experience or screening according to the cycle, and the degree of difference in body length and weight is the key to determining whether to screen. Therefore, it is necessary to understand the growth differences of largemouth bass in the breeding process in real time, which helps to understand the growth of fish and the breeding effect, and provide a reference for subsequent management. The second part introduces image distortion correction. We know that to use two-dimensional images to obtain the information of three-dimensional space points lacks a constraint condition, depth information. This paper uses a monocular depth estimation algorithm to obtain the depth information of the image, combined with the key point detection algorithm in the first part, to obtain the world coordinate information of the head and tail of the largemouth bass, and calculates the estimated body length information of the largemouth bass through the distance formula, and further estimates the weight information based on the body length.

[0075] 1. Monocular Depth Estimation Algorithm The labeled dataset for depth estimation is not as simple as image annotation. Deep image annotation requires professional equipment to collect depth information, and it is difficult to build a dataset with tens of millions of depth labels. Traditional depth data comes from devices such as sensors and depth cameras. The cost of collecting data is high, and the collected depth information is difficult to match with the image, which increases the difficulty of processing. In contrast, Depth Anything focuses on a large amount of unlabeled monocular image data, which is easy to collect and involves scenes in various fields of various industries, with strong generalization ability. The specific implementation idea is to use a small amount of data with depth labels to assign pseudo labels to the unlabeled dataset, and then combine the labeled dataset with the assigned label dataset to train a model. The model inherits rich semantic priors from the pre-trained encoder, forcing the model to learn to use additional knowledge. After obtaining the model, it is fine-tuned using metric depth to improve the accuracy of the new model.

[0076] 2. Real-time estimation algorithm for largemouth bass body length This section combines the key point detection algorithm in the first part with the monocular depth estimation algorithm to write a real-time estimation algorithm for the body length of largemouth bass in a circulating water system. The monocular depth estimation is used to estimate the depth value of each pixel in the underwater image of largemouth bass. The pixel coordinates of the key points of the fish head and tail obtained by YOLO detection are combined, and the real three-dimensional coordinates of the fish head and tail are obtained through the conversion relationship between the camera coordinate system and the world coordinate system.

[0077] The detection results of the model include the coordinate information, category and confidence level of each predicted bounding box. Each bounding box contains the key point coordinate information of the fish head and tail. A judgment is added to the process. When the overall deviation of the estimated body length due to monocular depth estimation is greater than the threshold, it is discarded and not recorded as sample data.

[0078] 2.1 Largemouth bass length-to-weight regression model For largemouth bass, the corresponding relationship between body length and mass can be found by regression analysis, and the power function regression model is better than the linear regression model (Tong Minghang 2022). Therefore, this section uses the power function regression method to establish a regression model of fish body length and mass. The regression model formula is: (4-1) In the formula, is the mass of the fish in g, L is the length of the fish in cm, and b are solution parameters.

[0079] First, the body length and mass of 50 largemouth bass were measured and recorded manually. The body length was measured using a triangle ruler with an error of less than 1 mm. The eyes of the largemouth bass were covered with a towel during measurement. The fish was kept in a quiet state before measurement, and the reading was estimated to two decimal places. The mass was measured using an electronic scale with an error of less than 1 g. The fish was placed in a basin of water that could completely immerse the fish body to avoid human errors caused by large jumps due to stress. The reading was recorded after stabilization. The regression model of body length and mass of largemouth bass was established as follows: .

[0080] 3. ONNXRuntime inference engine and deployment In order to simplify the model and make it easier to deploy it in the circulating water monitoring system, the depth estimation model and target detection model are uniformly converted into ONNX format files and deployed on the terminal using the ONNXRuntime inference engine. ONNXRuntime supports a variety of operating backends including CPU, GPU, TensorRT, DML, etc. The video stream display of the monitoring system adopts a multi-threaded frame extraction mode, which separates the acquired image from the displayed image. After the image is acquired, it is input into the algorithm for calculation, and then transmitted to the display end for display. This method can clear the cache in real time, prevent the cache from being full due to the algorithm processing image speed not keeping up with the network video recorder recording speed, and enable the monitoring video to be displayed stably.

[0081] 4. Experiment on real-time estimation of body length and mass of largemouth bass 4.1 Test materials and test methods Ten largemouth bass cultured in a laboratory factory-scale circulating water mode were randomly selected as test samples. The body length of each fish was measured using a triangle with an error of less than 1 mm to obtain the true value of manual measurement; the mass was measured using an electronic scale with an error of less than 1 g to obtain the true value of manual weighing. The image was used as the input of the model to obtain the corresponding body length estimation value and mass estimation value, and the error was calculated by comparing the estimation result with the manual measurement result to verify the accuracy of the largemouth bass body length estimation algorithm. After the model was optimized through format conversion and inference structure optimization, the size of the target detection model was only 12.21Mb, and the size of the monocular depth estimation model was 96.65Mb. It was deployed in the monitoring system to play the picture in real time. 4.2 Results and Analysis In this section, a video of a largemouth bass swimming freely underwater is collected as a validation set. The real-time video stream is displayed as follows: Fig.12 As shown in Figure 2. Multiple images of a fish are input into the largemouth bass body length and mass estimation algorithm to obtain multiple estimation results, and the average value is taken as the body length and mass estimation value of the largemouth bass. The comparison between the manual measurement results and the visual estimation results is shown in Table 2.

[0082] Table 2 Estimation results of body length and mass of largemouth bass

[0083] The test results show that the average relative error of the body length of 10 largemouth bass is 6.4%, of which the minimum relative error of body length is 2.5% and the maximum is 13.8%; the average relative error of mass is 14.8%, the minimum relative error of mass is 5.4% and the maximum is 31.9%. Since the estimated value of monocular depth estimation at the edge position is sensitive to changes and there is an overall offset phenomenon, the generalization ability of the model needs to be improved, which may be the main reason for the error. The number of frames of real-time video stream reaches about 25, which can achieve the expected effect. The experiment shows that the body length and mass estimation algorithm can estimate the average body length and mass information of largemouth bass groups under the circulating water mode.

[0084] This section develops a body length estimation algorithm for largemouth bass in a recirculating aquaculture mode based on monocular depth estimation and YOLOv8 key point detection branch network. A highly robust depth estimation algorithm Depth Anything is introduced to achieve three-dimensional reconstruction based on the depth estimation information of two-dimensional images. A body length mass regression model for largemouth bass is constructed based on the power function regression method. , wrote a mass estimation algorithm for largemouth bass in recirculating aquaculture mode. Combined with the ONNXRuntime inference framework provided by Microsoft, it was deployed in the monitoring system. Finally, the accuracy and speed of the largemouth bass body length mass estimation algorithm were verified through experiments. Compared with the manual measurement data, the average error of body length was 6.4%, the average error of weight was 13.8%, and the number of real-time picture frames was about 25. At the same time, it also verified the feasibility of monocular depth estimation applied to the three-dimensional reconstruction of key points of fish bodies. Although the error is relatively large at present, with the continuous expansion of supervised depth data sets and the training and learning of unsupervised image data sets, the absolute accuracy of the depth estimation model will continue to improve. The application of monocular depth estimation in aquaculture, especially in recirculating aquaculture mode, as a tool for fish bioinformatics measurement has great development prospects.

Claims

1. A method for measuring the length and mass of underwater largemouth bass based on YOLOv8 and monocular depth estimation, characterized in that: The following steps are involved: 1) Use the YOLOv8 model to perform target detection and key point detection on underwater largemouth bass to obtain the key point coordinate information of the fish head and tail; 2) Obtain the depth information of the underwater image through a monocular depth estimation algorithm, and calculate the body length of the largemouth bass by combining the key point detection results; 3) Based on body length information, the mass of largemouth bass was estimated using a power function regression model; 4) The model was deployed into the recirculating aquaculture monitoring system to display and record the length and mass information of largemouth bass in real time.

2. The method according to claim 1, characterized in that The YOLOv8 model is trained by the following steps: 1) Collect image data of underwater largemouth bass, perform data enhancement, and generate training and validation sets; 2) Use transfer learning methods to train based on pre-trained models and optimize model parameters; 3) Evaluate model performance through the loss function and adjust the learning rate and optimizer parameters until the model converges.

3. The method according to claim 1, characterized in that The monocular depth estimation algorithm adopts the DepthAnything model and is implemented by the following steps: 1) Using a small amount of data with deep labels to assign pseudo labels to unlabeled datasets; 2) Combine the labeled dataset and the pseudo-labeled dataset to train the model and improve the generalization ability of the model; 3) Fine-tune the model by measuring depth to improve the accuracy of the model.

4. The method according to claim 1, characterized in that: The power function regression model is used to estimate the mass of largemouth bass, and its formula is: Among them, W is the mass of the fish, L is the body length of the fish, and a and b are regression parameters.

5. The method according to claim 1, characterized in that The method further includes calibrating the underwater camera and correcting image distortion, the specific steps of which are: 1) Use Zhang Zhengyou calibration method to calibrate the monocular camera to obtain the camera's intrinsic parameter matrix, extrinsic parameter matrix and distortion coefficient; 2) Use the OpenCV library to correct image distortion and reduce errors in the 3D reconstruction process.

6. The method according to claim 1, characterized in that The method also includes converting the model into ONNX format and deploying it on a terminal device using the ONNX Runtime inference engine to achieve real-time inference and display.

7. The method according to claim 1, characterized in that The method performs real-time body length and mass estimation by the following steps: 1) Obtain real-time video stream from the monitoring system and extract images frame by frame; 2) Input the image into the YOLOv8 model and the monocular depth estimation model to obtain the key point coordinates and depth information of the fish head and tail; 3) Calculate the three-dimensional distance between the fish head and tail to get the estimated value of body length; 4) Calculate the mass estimate based on the body length estimate using a power function regression model.

8. The method according to claim 1, characterized in that The method further includes performing error analysis on the estimation result, and the specific steps are: 1) Compare the manual measurement results with the model estimation results and calculate the relative errors of body length and mass; 2) Reduce estimation errors and improve model accuracy by adjusting model parameters and optimizing algorithms.

9. The method according to claim 1, characterized in that: The method is applicable to the factory-scale recirculating aquaculture model, and can monitor the growth status of largemouth black bass in real time, providing data support for aquaculture management.

Citation Information

Cited By

  • Deep learning-based multi-modal phenotype determination method for economic traits of lateolabrax japonicus

    CN121617134A

  • Fish body mass non-contact estimation method based on binocular vision

    CN121810695A

  • A non-contact fish body weight estimation method based on binocular vision

    CN121810695B

  • Non-contact fish body length measurement method based on multi-scale anti-shielding dynamic optimization

    CN122368518A