Juvenile fish phenotypic feature detection method and device, electronic equipment and storage medium

By combining RGB images and depth images, using the improved YOLOv8-Pose model and feature pyramid network, the problem of low phenotype detection accuracy of juvenile fish is solved, and accurate measurement of body length, width and weight of juvenile fish is achieved.

CN120298748APending Publication Date: 2025-07-11CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271124.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing methods have low accuracy in the detection of phenotypic characteristics of juvenile fish, especially in the case of small size, large breeding density, serious overlap and fast swimming speed, making it difficult to achieve accurate measurement.

Method used

Using a combination of RGB images and depth images, the improved YOLOv8-Pose model and feature pyramid network ContextGuideFPN were used to extract fish body key points, and the body length, width and weight data of juvenile fish were calculated using a three-dimensional coordinate system.

Benefits of technology

It realizes accurate extraction of phenotypic features of juvenile fish, improves the accuracy and accuracy of detection, and can accurately identify key points of fish in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298748A_ABST
    Figure CN120298748A_ABST
Patent Text Reader

Abstract

The invention provides a juvenile fish phenotypic feature detection method and device, electronic equipment and a storage medium, and relates to the technical field of computer vision, and the method comprises the steps: obtaining an RGB image and a depth image of a to-be-detected juvenile fish; based on the RGB image, key point extraction is carried out on the to-be-detected juvenile fish, and a plurality of fish body key points are determined; determining three-dimensional coordinates of the fish body key points based on the two-dimensional coordinates of the fish body key points and the depth image; and determining phenotypic characteristics of the to-be-detected juvenile fish based on the three-dimensional coordinates. And through the RGB image and the depth image, fish body key points are extracted, three-dimensional coordinates of the juvenile fish are calculated, and then phenotypic features of the juvenile fish are obtained. The distance information of the real world is provided based on the depth image, so that the measurement result has a real physical scale, the three-dimensional coordinates of the key points of the fish body can be accurately calculated in combination with the RGB image, accurate extraction of the juvenile fish phenotypic features is realized, and the accuracy of phenotypic feature detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device, electronic device and storage medium for detecting the phenotypic characteristics of juvenile fish. Background Art

[0002] In the process of detecting fish phenotypic characteristics, existing methods generally implement fish phenotypic characteristic detection based on machine learning or deep learning. Specifically, through means such as image segmentation and feature extraction, measurements of phenotypic characteristics such as fish length and width are realized.

[0003] The existing method of directly predicting based on machine learning or deep learning has low accuracy. Especially for the situation of juvenile fish (small size), high breeding density, serious overlap, and fast swimming speed, it is extremely difficult to accurately measure, and the accurate determination process of the phenotypic characteristics of juvenile fish cannot be realized. Summary of the Invention

[0004] The present invention provides a method, device, electronic device and storage medium for detecting the phenotypic characteristics of juvenile fish, so as to improve the accuracy of detecting the phenotypic characteristics of juvenile fish.

[0005] The present invention provides a method for detecting the phenotypic characteristics of juvenile fish, including the following steps: Obtain the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected; Based on the RGB image, perform key point extraction on the juvenile fish to be detected, and determine multiple fish body key points of the juvenile fish to be detected; Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system; Based on the three-dimensional coordinates, determine the phenotypic characteristics of the juvenile fish to be detected, and the phenotypic characteristics include body length data, body width data and body weight data.

[0006] According to the method for detecting the phenotypic characteristics of juvenile fish provided by the present invention, the step of performing key point extraction on the juvenile fish to be detected based on the RGB image and determining multiple fish body key points of the juvenile fish to be detected includes: Input the RGB image into a key point detection model, and obtain multiple fish body key points of the juvenile fish to be detected output by the key point detection model. The key point detection model is trained based on juvenile fish image samples and the fish body key point labels corresponding to the juvenile fish image samples.

[0007] According to the method for detecting the phenotypic characteristics of juvenile fish provided by the present invention, the construction process of the key point detection model includes: Using a detection module to replace the C2f module in the backbone network of the YOLOv8-Pose model, an improved YOLOv8-Pose model is obtained. The detection module is constructed based on a residual connection between a collaborative multi-attention SMA module and a convolutional gated linear unit CGLU module; Embedding a Feature Pyramid Network ContextGuideFPN into the neck network of the improved YOLOv8-Pose model to obtain the key point detection model.

[0008] According to a method for detecting the phenotypic characteristics of juvenile fish provided by the present invention, determining the three-dimensional coordinates of the key points of the fish body in a three-dimensional world coordinate system based on the two-dimensional coordinates of the key points of the fish body in the RGB image and the depth image includes: Based on the two-dimensional coordinates of the key points of the fish body in the RGB image and the depth image, determining the depth value of the key points of the fish body; Connecting multiple key points of the fish body in the RGB image to obtain multiple connecting lines; Based on the depth values of a preset number of pixel points near the key points of the fish body and on the connecting lines, adjusting the depth value of the key points of the fish body to obtain an adjusted depth value; Based on the adjusted depth value and the two-dimensional coordinates corresponding to the key points of the fish body, determining the three-dimensional coordinates of the key points of the fish body in a three-dimensional world coordinate system.

[0009] According to a method for detecting the phenotypic characteristics of juvenile fish provided by the present invention, determining the phenotypic characteristics of the juvenile fish to be detected based on the three-dimensional coordinates includes: Based on the three-dimensional coordinates, determining the body length data and body width data of the juvenile fish to be detected; Inputting the body length data and the body width data into a weight prediction model to obtain the weight data output by the weight prediction model. The weight prediction model is trained based on a sample of length and width data and the weight data label corresponding to the sample of length and width data.

[0010] According to a method for detecting the phenotypic characteristics of juvenile fish provided by the present invention, the key points of the fish body include the fish mouth point, the midpoint of the tail, the midpoint of the caudal fin, the dorsal fin point, and the pelvic fin point.

[0011] The present invention also provides a device for detecting the phenotypic characteristics of juvenile fish, including the following modules: An image acquisition module for acquiring an RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected; A key point extraction module, configured to extract key points of the to-be-detected juvenile fish based on the RGB image, and determine multiple fish body key points of the to-be-detected juvenile fish; A three-dimensional coordinate determination module, configured to determine three-dimensional coordinates of the fish body key points in a three-dimensional world coordinate system based on two-dimensional coordinates of the fish body key points in the RGB image and the depth image; A phenotypic feature extraction module, configured to determine phenotypic features of the to-be-detected juvenile fish based on the three-dimensional coordinates, where the phenotypic features include body length data, body width data, and body weight data.

[0012] The present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the method for detecting phenotypic features of juvenile fish as described in any one of the above is implemented.

[0013] The present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for detecting phenotypic features of juvenile fish as described in any one of the above is implemented.

[0014] The present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for detecting phenotypic features of juvenile fish as described in any one of the above is implemented.

[0015] The method, device, electronic device, and storage medium for detecting phenotypic features of juvenile fish provided by the present invention extract fish body key points through an RGB image and a depth image, calculate three-dimensional coordinates of the juvenile fish, and further obtain phenotypic features of the juvenile fish, including body length, body width, and body weight data. Using the depth image provides distance information in the real world, making the measurement results have a real physical scale. By combining the RGB image, the three-dimensional coordinates of the fish body key points can be accurately calculated, realizing the accurate extraction of phenotypic features of juvenile fish and improving the accuracy of detecting phenotypic features of juvenile fish. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the method for detecting phenotypic features of juvenile fish provided by the present invention.

[0018] Figure 2 It is a schematic structural diagram of the detection module provided by the present invention.

[0019] Figure 3 It is a schematic structural diagram of the SMA module provided by the present invention.

[0020] Figure 4 It is a schematic structural diagram of the GGLU module provided by the present invention.

[0021] Figure 5 It is a schematic structural diagram of the feature pyramid network provided by the present invention.

[0022] Figure 6 It is a schematic structural diagram of the key point detection model provided by the present invention.

[0023] Figure 7 It is a schematic diagram of the phenotypic feature extraction process provided by the present invention.

[0024] Figure 8 It is a schematic diagram of the fish body key points provided by the present invention.

[0025] Figure 9 It is a schematic structural diagram of the juvenile fish phenotypic feature detection device provided by the present invention.

[0026] Figure 10 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0027] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] Figure 1 It is a schematic flowchart of the juvenile fish phenotypic feature detection method provided by the present invention. As Figure 1 shown, the method includes the following: Step 110, obtaining the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected; Step 120, based on the RGB image, performing key point extraction on the juvenile fish to be detected to determine multiple fish body key points of the juvenile fish to be detected; Step 130, based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determining the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system; Step 140: Based on the three-dimensional coordinates, determine the phenotypic characteristics of the juvenile fish to be detected, where the phenotypic characteristics include body length data, body width data, and body weight data.

[0029] The execution subject of the juvenile fish phenotypic characteristic detection method provided by the present invention can be an electronic device, a component in the electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), or a personal computer (PC), etc. The present invention does not make specific limitations.

[0030] Taking the computer executing the juvenile fish phenotypic characteristic detection method provided by the present invention as an example, the technical solution of the present invention will be described in detail below.

[0031] In step 110, obtain the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected.

[0032] The RGB image is an image representation method using three color channels: Red, Green, and Blue. In this image, each pixel is determined by the values of these three color channels. Specifically, based on the binocular camera system, the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected can be obtained.

[0033] The binocular camera system consists of two cameras (left camera and right camera). By simulating the parallax principle of the human eye, the same scene is photographed from different angles, and the depth information is calculated using the difference between the two images, so as to determine the depth image of the juvenile fish to be detected.

[0034] In step 120, based on the RGB image, perform key point extraction on the juvenile fish to be detected, and determine multiple fish body key points of the juvenile fish to be detected.

[0035] After obtaining the RGB image, the extraction of fish body key points can be realized based on the key point detection model.

[0036] Before performing keypoint detection, it is necessary to obtain video or image data of juvenile fish swimming freely underwater, construct a high-quality dataset, and accurately label the surface keypoints of the juvenile fish in the dataset. These keypoints mainly include the coordinates of the main parts of the juvenile fish, such as the head, tail, abdomen, etc. These keypoints are crucial for subsequent detection of the phenotypic characteristics of juvenile fish. When constructing the dataset, data on the length, width, and weight of juvenile fish at different growth stages are also collected for subsequent fitting and prediction of body weight in phenotypic characteristics.

[0037] After constructing the dataset, the initial keypoint detection model is trained. After training is completed, based on the trained keypoint detection model, the extraction process of multiple fish body keypoints of the juvenile fish to be detected can be realized.

[0038] Optionally, the keypoint detection model can detect and extract the keypoints of juvenile fish based on the YOLOv8-Pose model. As an efficient pose estimation model, YOLOv8-Pose can accurately extract the surface keypoints of juvenile fish in a complex underwater environment through an optimized deep learning network.

[0039] Optionally, considering the characteristics of fast swimming speed and small body size of juvenile fish, to further improve the performance of the keypoint model in detecting complex backgrounds and small targets, the YOLOv8-Pose model can be improved to enhance the model's focusing ability on the characteristics of juvenile fish. Specifically, the feature fusion method of the ContextGuideFPN of the feature pyramid network can be combined in the YOLOv8-Pose model to further enhance the multi-scale feature fusion ability, enabling the model to accurately detect the fish body keypoints of juvenile fish at different scales. Especially in the case where the size of juvenile fish is small and the background is complex, it can still maintain high accuracy.

[0040] In step 130, based on the two-dimensional coordinates of the fish body keypoints in the RGB image and the depth image, determine the three-dimensional coordinates of the fish body keypoints in the three-dimensional world coordinate system.

[0041] Combining the two-dimensional coordinates of the fish body keypoints in the RGB image with the depth information in the depth image can be converted into coordinates in the three-dimensional world coordinate system.

[0042] First, camera calibration is required to obtain the internal parameters (focal length, principal point, etc.) and external parameters (camera position and attitude) of the camera.

[0043] Based on the two-dimensional coordinates of the fish body keypoints, obtain the corresponding depth value (Z coordinate) from the depth image. Using the camera model and depth information, convert the two-dimensional coordinates into three-dimensional coordinates, thereby realizing the determination process of the three-dimensional coordinates of the fish body keypoints in the three-dimensional world coordinate system.

[0044] For example, for the depth image obtained by a binocular camera system, 3D reconstruction can be performed based on binocular stereo vision technology. The binocular camera system obtains left and right view images and performs stereo matching to generate a depth map. This depth map contains the depth information of juvenile fish in three-dimensional space, which is crucial for subsequent phenotypic feature measurement. Combining the internal and external parameters of the camera and the two-dimensional coordinates of the fish body key points in the image, these two-dimensional coordinates can be converted into three-dimensional coordinates, thereby realizing the 3D reconstruction of each fish body key point on the surface of the juvenile fish. Through the three-dimensional coordinates, important phenotypic features such as the body length and body width of the juvenile fish can be accurately calculated, providing basic data for subsequent weight prediction.

[0045] In step 140, based on the three-dimensional coordinates of the fish body key points, the phenotypic features of the to-be-detected juvenile fish are determined, and the phenotypic features include body length data, body width data, and weight data.

[0046] It should be noted that the extracted fish body key points may include the fish mouth point, the midpoint of the tail, the midpoint of the caudal fin part, the dorsal fin point, and the ventral fin point.

[0047] After obtaining the three-dimensional coordinates of the key points on the surface of the juvenile fish, a multi-segment measurement method can be adopted to accurately calculate the body length and body width of the juvenile fish. Through the Euclidean distance calculation formula, based on the spatial positions between the fish body key points, the accurate body length and body width data of the juvenile fish can be obtained. At the same time, in order to realize the prediction of weight, the body length and body width data of the juvenile fish can be combined, and a neural network model is used to predict the weight. This neural network is a fitting model obtained through training based on the existing juvenile fish growth data set, which can predict the weight of the juvenile fish according to its body length and body width, and has a high prediction accuracy.

[0048] However, in practical applications, in the case of juvenile fish being juvenile fish, due to the relatively fast swimming speed of juvenile fish, the situation of fish body key point deviation or abnormal depth value may occur, affecting the measurement accuracy. To address this problem, it can be optimized based on the depth value correction method to solve the problem of sudden change in the depth value of fish body key points caused by swimming. The correction method is to correct based on the depth difference at the edge of the connection line of the fish body key points. Specifically, at the edge of the connection line of the body length and body width of the juvenile fish, a certain number of pixel points are taken, the depth difference is calculated, and interpolation is performed through the depth value information of the previous moment to correct the depth value of the current fish body key point. In the actual experimental performance of this method, it not only effectively reduces the problem of sudden change in depth value caused by swimming, but also improves the measurement accuracy of the body length, body width, and weight of juvenile fish.

[0049] The method for detecting the phenotypic characteristics of juvenile fish provided by the present invention extracts the key points of the fish body through RGB images and depth images, calculates the three-dimensional coordinates of the juvenile fish, and then obtains the phenotypic characteristics of the juvenile fish, including body length, body width, and weight data. The depth image provides the distance information of the real world, making the measurement results have a real physical scale. By combining the RGB image, the three-dimensional coordinates of the key points of the fish body can be accurately calculated, realizing the accurate extraction of the phenotypic characteristics of the juvenile fish and improving the accuracy of the detection of the phenotypic characteristics of the juvenile fish.

[0050] In one embodiment, based on the RGB image, key point extraction is performed on the juvenile fish to be detected, and multiple key points of the fish body of the juvenile fish to be detected are determined, including: inputting the RGB image into a key point detection model to obtain multiple key points of the fish body of the juvenile fish to be detected output by the key point detection model, and the key point detection model is trained based on juvenile fish image samples and the corresponding fish body key point labels of the juvenile fish image samples.

[0051] Based on the key point detection model, the extraction of the key points of the fish body is realized.

[0052] Before performing key point detection, it is necessary to obtain video or image data of juvenile fish swimming freely underwater, construct a high-quality data set, and accurately label the surface key points of the juvenile fish in the data set. These key points mainly include the coordinates of the main parts of the juvenile fish, such as the head, tail, abdomen, etc. These key points are crucial for the subsequent detection of the phenotypic characteristics of the juvenile fish. When constructing the data set, the length, width, and weight data of juvenile fish at different growth stages are also collected for subsequent fitting and prediction of the weight in the phenotypic characteristics.

[0053] After constructing the data set, the initial key point detection model is trained. After the training is completed, based on the trained key point detection model, the extraction process of multiple key points of the fish body of the juvenile fish to be detected can be realized.

[0054] Optionally, based on the detection process of the key point detection model, it can be to detect and extract the key points of the juvenile fish based on the YOLOv8-Pose model. As an efficient pose estimation model, YOLOv8-Pose can accurately extract the surface key points of the juvenile fish in a complex underwater environment through an optimized deep learning network.

[0055] In one embodiment, the construction process of the key point detection model includes: using a detection module to replace the C2f module in the backbone network of the YOLOv8-Pose model to obtain an improved YOLOv8-Pose model, where the detection module is constructed based on a residual connection between a synergistic multi-attention (SMA) module and a convolutional gated linear unit (CGLU) module; embedding a feature pyramid network ContextGuideFPN into the neck network of the improved YOLOv8-Pose model to obtain the key point detection model.

[0056] Specifically, the key point detection model for detecting the key points of juvenile fish is determined based on the improvement of the YOLOv8-Pose model.

[0057] The YOLOv8-Pose model is a model specifically developed for pose estimation based on the YOLOv8 architecture, which further realizes accurate prediction of poses on the basis of object detection.

[0058] The YOLOv8-Pose model specifically includes a backbone network, a neck network, and a head network. Among them, the backbone network is used to efficiently extract multi-scale features of images. The neck network is used to fuse features of different scales, enhance the feature expression ability, and enable the model to better capture details and context information of different parts. The head network is used to predict the positions of key points.

[0059] Specifically, a detection module is pre-constructed.

[0060] The detection module is constructed based on a residual connection between a synergistic multi-attention (SMA) module and a convolutional gated linear unit (CGLU) module. The structural schematic diagram of the constructed detection module can be as Figure 2 shown in the structural schematic diagram of the detection module provided by the present invention. Among them, LayerNorm refers to normalization processing.

[0061] It should be noted that the SMA module integrates multiple attention mechanisms, including position encoding (PE), channel attention (CA), spatial attention (SA), and pixel attention (PA), thereby effectively enhancing the model's ability to capture features of different scales. The structural schematic diagram of the constructed SMA module can be as Figure 3As shown in the structural schematic diagram of the SMA module provided by the present invention. Among them, Embedded Modulator is a mechanism for adjusting or modulating feature representations. It is usually used to embed additional information (such as location information or context information) into features to enhance the expressive power of the model. The multi-head self-attention mechanism (MSA) calculates self-attention in parallel through multiple attention heads, thereby capturing the relationships between different positions in the input sequence. Each attention head can focus on different parts of the input, and finally the results are combined.

[0062] At the same time, the CGLU module constructs a more complex feature mapping through an enhanced multi-layer perceptron, significantly improving the expressive power of the model. By adopting residual connections, these two modules effectively reduce the number of parameters and stabilize the training process.

[0063] The constructed detection module has fewer parameters, improves the expressive power of features, and significantly enhances the model's sensitivity to local features. In addition, adopting the design idea of ResNet to make an efficient connection between the SMA module and the CGLU module makes the entire network perform better in complex scenarios. The improved detection module architecture has been significantly optimized compared to the original structure, improving the detection accuracy and performing well in multiple performance indicators.

[0064] Specifically, the main component of the SMA module is the attention mechanism. The attention mechanism can optimize the model's ability to capture core feature information in images, and it mainly includes two categories: spatial attention mechanism and channel attention mechanism. The spatial attention mechanism focuses on evaluating the importance of image spatial positions. By focusing on specific positions and ignoring secondary feature information, the model performance is improved. For example, in an image classification task, the model will pay more attention to the area where the target object is located rather than the background information. On the other hand, the channel attention mechanism focuses on differentiating the importance of different feature channels and assigns different weights to each feature channel. In this embodiment, a combination of multiple attentions is used to improve the C2f module so that it can integrate the advantages of multiple attention mechanisms and has a more sensitive extraction ability for the detection of juvenile fish.

[0065] Specifically, pixel attention enables the model to focus on the importance of each pixel in the juvenile fish image, which is crucial for distinguishing the key points of juvenile fish from other background features. Channel attention allows the model to assign different weights between different channels of the feature map, further highlighting the features of the key points. Spatial attention focuses on the relationships between different positions on the feature map, which is very helpful for identifying the spatial layout of different parts of the juvenile fish body.

[0066] The GGLU module utilizes the powerful feature extraction ability of convolutional neural networks and the dynamic gating mechanism of GLU to achieve efficient extraction and selection of features from juvenile fish images. The schematic diagram of the GGLU module structure can be as shown in Figure 4 the schematic diagram of the GGLU module structure provided by the present invention. Among them, Activation is the activation function, and Linear is the linear transformation operation. This module maps the input feature map to twice the dimension of the hidden feature space through a 1x1 convolutional layer and divides it into two parts along the channel dimension: x and v. Among them, x is passed as the input feature to the subsequent depthwise separable convolutional layer, and v is used as the gating signal, which is combined with the output after depthwise separable convolution of x through element-wise multiplication to achieve selective enhancement or suppression of features. The depthwise separable convolutional layer can use a 3x3 convolutional kernel, with a stride of 1, padding of 1, and the number of groups set to the number of hidden features, which helps to reduce the number of parameters and improve the computational efficiency. In addition, GELU (Gaussian Error Linear Unit) can be introduced as the activation function to enhance the non-linear expression ability of the network. Subsequently, the features are mapped back to the output feature space through a 1x1 convolutional layer, and the risk of overfitting is reduced through the Dropout layer.

[0067] In the task of detecting the phenotypic characteristics of juvenile fish, the CGLU module can not only effectively extract the detailed features in the image, but also dynamically adjust the importance of features through the gating mechanism, thereby improving the sensitivity to key phenotypic characteristics.

[0068] It should be noted that applying the SMA module to the task of detecting key points of juvenile fish can significantly improve the detection accuracy and robustness. The SMA module provides a multi-dimensional feature extraction framework for the analysis of juvenile fish images by integrating pixel attention, channel attention, and spatial attention. In the detection of key points of juvenile fish, this means that the model can understand the pose characteristics of juvenile fish at different levels and improve the detection accuracy more precisely.

[0069] Specifically, pixel attention enables the model to focus on the importance of each pixel in the juvenile fish image, which is crucial for distinguishing the key points of juvenile fish from other background features. Channel attention allows the model to assign different weights among different channels of the feature map, further highlighting the features of the key points. Spatial attention focuses on the relationship between different positions on the feature map, which is very helpful for identifying the spatial layout of different parts of the juvenile fish body.

[0070] The combination of such multi-attention mechanisms enables the SMA module to accurately identify and locate the key points of juvenile fish in a complex underwater environment, even when the postures of juvenile fish change, are occluded, or the lighting conditions are not ideal. For example, when identifying the position of the eyes of juvenile fish, the improved model can not only capture the local features of the eye region but also understand the position of the eyes relative to other parts of the juvenile fish's body through spatial attention, thereby improving the accuracy of detection.

[0071] In addition, the multi-head self-attention mechanism of the SMA module can further analyze the mutual relationships between the key points of juvenile fish, which is particularly important for constructing a detailed body map of juvenile fish. This mechanism enables the model to capture the relative positions and spatial relationships between various parts of the juvenile fish's body, providing rich information for subsequent growth monitoring and morphological analysis. However, the enhanced multi-layer perceptron E-MLP in the SMA module has a large number of parameters and a large amount of computation. The CGLU module can be used to replace this module, which not only reduces the number of parameters of the module, speeds up the detection speed, but also improves the detection accuracy.

[0072] The detection module constructed by the CGLU module and the SMA module can be applied to the detection of the key points of juvenile fish, which can significantly improve the performance of the model. The SMA module enables the model to capture multi-scale features in the image by fusing pixel, channel, and spatial attention mechanisms, which is crucial for identifying the detailed features of juvenile fish. The fusion of such multi-attention mechanisms allows the model to simultaneously focus on the global structural information and local detailed features when processing juvenile fish images, thereby more accurately identifying the key points of juvenile fish.

[0073] The introduction of the CGLU module further enhances the sensitivity of the model to local features. By adding a 3x3 depthwise separable convolution before the gated linear unit, the CGLU module provides a gating signal for each token based on its nearest neighbor features, which enables the model to capture more fine-grained feature changes in the local area. Especially when identifying the tiny structures of juvenile fish, such as key parts like eyes, mouths, and fins, the sensitivity to local features is particularly important.

[0074] At the same time, the CGLU module does not weaken the global perception ability of the SMA module but complements it. The combination of CGLU and the multi-head self-attention mechanism in the SMA module enables the model to maintain sensitivity to global context information while also conducting a more in-depth analysis of local features. The organic combination of such global and local information provides a more comprehensive perspective for the model, enabling it to accurately locate and identify the key points of juvenile fish in a complex environment.

[0075] In addition, the high computational efficiency of the CGLU module helps reduce the computational burden of the overall model, especially when dealing with high-resolution images. This, combined with the multi-scale feature capture ability of the SMA module, enables the model to maintain high accuracy while also meeting the requirements for real-time or near-real-time detection.

[0076] Based on the YOLOv8-Pose model, a detection module was constructed, and the C2f module in the backbone network of the original YOLOv8-Pose model was replaced with the constructed detection module to obtain the improved YOLOv8-Pose model.

[0077] It can be understood that there is a residual connection between the SMA module and the CGLU module in the detection module, allowing information to be directly transmitted between the modules, which helps alleviate the problem of gradient disappearance and enhance the feature expression ability. What is received is the feature output by the upper convolutional module. The MSA improves the model's ability to capture and utilize information in the feature space by introducing a multi-scale and multi-dimensional attention mechanism. The CGLU controls the feature information flow by introducing a gating mechanism. It contains a convolutional layer to extract features and a gating unit to regulate the transmission of these features. This structure helps the model adaptively select important features and suppress unimportant features. Finally, the features at specific positions of the key points are output.

[0078] It should be noted that replacing the C2f module in the backbone network of the original YOLOv8-Pose model with the detection module can specifically be to replace the C2f module between CBS and CBS in the backbone network, as well as the C2f module between CBS and SPPF.

[0079] Furthermore, the improved YOLOv8-Pose model was further adjusted by specifically embedding the ContextGuideFPN of the feature pyramid network into the neck network of the improved YOLOv8-Pose model to obtain the key point detection model.

[0080] The role of the Neck layer in the neck network is to achieve multi-scale fusion. In the module structure of YOLOv8-Pose, its role is to further fuse features and enhance the context based on the features extracted by the backbone network. The design of the neck helps the network perceive targets at different scales. The key part during fusion at different scales is the Concat operation. The Concat module is applied between different levels of the network to achieve the fusion of feature maps. Specifically, the Concat module stitches together the feature map with low resolution but rich semantic information and the feature map with high resolution but less semantic information, thus forming a larger output feature map. This cross-layer connection method can balance details and the perception range to a certain extent, contributing to improving the accuracy of object detection. However, it does not have particularly outstanding performance in the feature fusion of small targets.

[0081] Based on this, the Feature Pyramid Network ContextGuideFPN can be embedded into the position of feature fusion in the Neck layer, specifically between the C2f module and the Upsample module in the Neck layer. ContextGuideFPN receives features of different scales from the feature extraction module, sets different weights for the importance of the features, and after processing, suppresses the weights at non-key point positions and enhances the feature weights at key point positions.

[0082] For ContextGuideFPN, specifically, it can be divided into three parts: multi-modal feature fusion, attention-guided weight assignment, and feature interaction enhancement. For multi-modal feature fusion, ContextGuideFPN receives features from different network layers and resolves the information inconsistency problem between different features through feature alignment and fusion operations. For the attention-guided weight assignment module, ContextGuideFPN uses the attention mechanism (Squeeze-and-Excitation Networks, SE) to weight the features, weights the feature maps according to the importance of different feature maps, and highlights the features that contribute to the key point detection of juvenile fish in the early stage. Regarding feature interaction enhancement, the fused features not only contain the direct information of the input features but also enhance the complementarity and context information expression ability of the two-way features through the weight interaction mechanism. In ContextGuideFPN, if the number of channels of the two-way input features is different, the number of channels can be adjusted through a convolutional layer, and the two aligned features are concatenated in the channel dimension to form a joint feature representation. Subsequently, the concatenated features are weighted by the attention mechanism SE to generate a channel weight vector, realizing the dynamic adjustment of the two-way features, thereby achieving the effect of highlighting key features and suppressing irrelevant noise. The weighted features are enhanced through interaction to enhance the semantic expression ability, and the enhanced features are fused here again. In this module, SE plays an important role in weight assignment and noise data suppression.

[0083] Specifically, the structural schematic diagram of the feature pyramid network ContextGuideFPN can be as Figure 5 shown in the structural schematic diagram of the feature pyramid network provided by the present invention. Among them, Channel1 and Channel2 are two different feature channels, and OutChannel refers to the output channel.

[0084] ContextGuideFPN demonstrates many significant advantages in the task of juvenile fish key point detection. These advantages are mainly reflected in the accuracy of feature fusion, excellent scale adaptation ability, and flexible processing ability for multi-dimensional inputs. For the specific task of juvenile fish key point localization, the juvenile fish in the image are often small in size, have variable postures, and have a complex background, which is likely to cause significant interference. Therefore, it is particularly important to effectively fuse information of different scales and accurately lock the key points.

[0085] ContextGuideFPN, with the Squeeze-and-Excitation (SE) attention mechanism, achieves adaptive weighted adjustment of the feature map channels. This mechanism can dynamically evaluate the contribution of each channel and adjust the weights accordingly based on its importance, effectively suppressing unimportant features while enhancing key information. In the detection of key points of juvenile fish, due to the complex background and diverse postures of juvenile fish, especially when the details of juvenile fish are blurred or overlap with the background, redundant information will seriously interfere with the accurate detection of key points. The SE attention mechanism can intelligently adjust the weights of each channel in the feature map, effectively suppressing background noise and irrelevant features, making key features more prominent, thus helping the model to more easily focus on the key parts of juvenile fish, such as fins, eyes and other key areas.

[0086] ContextGuideFPN demonstrates excellent ability in fusing multi-scale features. The images of juvenile fish often contain information at multiple scales. Especially for different parts of the fish body, due to differences in perspective or shooting distance, their scales will also change. Feature fusion methods in related methods may lose detail information or lead to inaccurate key point localization when dealing with these different scale features. However, ContextGuideFPN can automatically adjust the weights of each scale to ensure that all parts of juvenile fish, whether it is the overall contour or small parts such as fins and eyes, can be fully emphasized. Therefore, in a multi-scale complex scenario, this module can still efficiently fuse features at different scales, thereby improving the detection accuracy of key points.

[0087] ContextGuideFPN also has a powerful 1x1 convolution channel adjustment ability. In the task of detecting key points of juvenile fish, it is usually necessary to fuse image information from different levels and different feature dimensions. The feature maps extracted by different convolutional layers may have inconsistent numbers of channels, while this module can easily adjust the number of channels through 1x1 convolution to achieve the unification of different feature map dimensions, so as to perform more efficient fusion. This mechanism enables the model to flexibly handle the situation of dimension mismatch, ensuring that the feature fusion effect is not affected by the different numbers of channels of the input feature maps. At the same time, 1x1 convolution helps to reduce the computational amount and improve the operation efficiency, enabling the model to perform the task of detecting key points of juvenile fish more efficiently while maintaining high accuracy.

[0088] The design of ContextGuideFPN can also effectively suppress redundant information, which is particularly important when dealing with complex backgrounds. When juvenile fish swim in water, they are easily affected by factors such as water flow and light, resulting in complex and variable background information. Through the attention mechanism, this module can accurately distinguish which features are necessary for key point localization, thus effectively eliminating irrelevant features in the background. In practical applications, the suppression of such redundant information can significantly reduce the occurrence of false detections and missed detections, especially in complex backgrounds, ensuring the accurate detection of key points of juvenile fish.

[0089] The multi-level fusion ability of ContextGuideFPN enables the model to effectively integrate low-level detail information (such as fish scales and the outline of juvenile fish) and high-level semantic information (such as the overall shape of the fish) when processing juvenile fish images. This fusion not only significantly improves the model's representation ability of juvenile fish but also enhances its adaptability to key points in different environments and poses. Optionally, when improving the YOLOv8-Pose model, the structural schematic diagram of the constructed key point detection model can be as Figure 6 shown in the structural schematic diagram of the key point detection model provided by the present invention. Among them, CBS represents the convolutional layer, batch normalization layer, and activation function (SiLU), and SPPF represents spatial pyramid pooling fast, which is used to extract features at different scales to improve the model's detection ability for targets of different sizes. UpSample is upsampling, which is used to increase the spatial resolution of the feature map. Pose-head is the head network for pose estimation, which is used to extract key point information from the feature map and predict the pose. SMA-CGLU is the constructed detection module.

[0090] In one embodiment, based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determining the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system includes: based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determining the depth value of the fish body key points; connecting multiple fish body key points in the RGB image to obtain multiple connecting lines; based on the depth values of a preset number of pixel points near the fish body key points and on the connecting lines, adjusting the depth value of the fish body key points to obtain an adjusted depth value; based on the adjusted depth value and the two-dimensional coordinates corresponding to the fish body key points, determining the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system.

[0091] Connect multiple fish body key points in the RGB image according to biological characteristics to form connection lines. The purpose of the connection lines is to utilize the spatial relationship between the key points to optimize the accuracy of depth values.

[0092] Extract a preset number (such as 5 - 10) of pixel points near the connection lines. These pixel points are located on or near the connection lines and are used for the optimization of local depth values. The depth values of the key points can be adjusted using the statistical quantities (such as the mean or median) of the local depth values.

[0093] Utilizing the information of the connection lines and the local area can improve the stability of the depth values.

[0094] Specifically, 5 pixel points can be taken at the connection line and the depth difference between two points can be calculated from the inside out. If the depth interpolation between the two points with the farthest depth does not exceed 5 cm, then calculate according to the original key points. If the depth value exceeds 5 cm, then use the depth value of the innermost key point with the smallest difference between the two calculated points as the depth value of the fish body key point, and calculate the length and width. Compared with the clustering algorithm, this method greatly reduces the calculation time and can achieve accurate measurement of the length and width.

[0095] In one embodiment, based on the three - dimensional coordinates, determine the phenotypic characteristics of the to - be - detected juvenile fish, including: based on the three - dimensional coordinates, determine the body length data and body width data of the to - be - detected juvenile fish; input the body length data and the body width data into a weight prediction model to obtain the weight data output by the weight prediction model, and the weight prediction model is trained based on the length - width data samples and the weight data labels corresponding to the length - width data samples.

[0096] Specifically, the extraction process of the phenotypic characteristics of the to - be - detected juvenile fish can be as Figure 7 shown in the schematic diagram of the phenotypic characteristic extraction process provided by the present invention. After obtaining the RGB image and the depth image, first determine the three - dimensional coordinates of the fish body key points. Based on the three - dimensional coordinates, determine the body length data and body width data.

[0097] The body length data of the juvenile fish usually refers to the straight - line distance from the fish head to the fish tail. The body width data of the juvenile fish usually refers to the width at the widest part of the juvenile fish.

[0098] A large amount of body length, body width data and their corresponding weight data of juvenile fish can be collected through actual measurement or public data sets. Each piece of data includes three fields: body length, body width, and weight.

[0099] Divide the data set into a training set and a test set (such as 80% training set, 20% test set). Use the training set to train the model and adjust the hyperparameters to optimize the performance. Input the calculated body length and body width data into the trained weight prediction model. The model predicts the weight of the juvenile fish according to the input data.

[0100] Among them, the weight prediction model can be a WBP weight fitting model based on a BP neural network, and the weight of juvenile fish is obtained by inputting the length and width. The weight prediction model is trained based on the length-width data samples and the weight data labels corresponding to the length-width data samples. The length-width data samples are determined based on the body length and body width data of juvenile fish.

[0101] Precisely calculating the body length and body width data based on the three-dimensional coordinates of the key points of the fish body, and efficiently estimating the weight data using the weight prediction model, can achieve the precise determination of the phenotypic characteristics of juvenile fish.

[0102] In one embodiment, the key points of the fish body include the fish mouth point, the midpoint of the tail, the midpoint of the caudal fin part, the dorsal fin point, and the pelvic fin point.

[0103] The key points of the fish body refer to a series of important points used to describe and locate the characteristics of juvenile fish in fish morphology, ethology, and related applied research. These key points of the fish body can be as Figure 8 shown in the schematic diagram of the key points of the fish body provided by the present invention. Specifically, they include the fish mouth point, the midpoint of the tail, the midpoint of the caudal fin part, the dorsal fin point, and the pelvic fin point.

[0104] The fish mouth point is located at the very front end of the fish head, at the opening of the fish mouth. The midpoint of the tail is located at the central position of the tail of the juvenile fish, usually at the base of the caudal fin. The midpoint of the caudal fin part is located at the center point of the caudal fin, usually the point with the largest amplitude of caudal fin swing. The dorsal fin point is located at the very front end of the dorsal fin on the back of the juvenile fish. The pelvic fin point is located at the base or central position of the pelvic fin.

[0105] It should be noted that in the case of juvenile fish, due to the small body size of juvenile fish, the direct head-to-tail measurement scheme in related methods cannot meet the precise measurement of curved fish. The measurement scheme for curved fish in related methods is a curve fitting scheme, which involves curve fitting modeling based on key points, and the computational time complexity is relatively high, and it is more accurate for measuring medium and large-sized fish.

[0106] The present invention implements a multi-segment measurement method for juvenile fish. By identifying the fish mouth point, the midpoint of the tail, and the midpoint of the caudal fin, and on the basis of these points, taking the midpoints of the pelvic fin and the dorsal fin as transition points. The multi-segment measurement method can simply, quickly, and accurately measure the length of juvenile fish. This is related to the characteristics of juvenile fish themselves. Mainly during swimming, the middle part of the body will not show a large bend, and only the tail will have a large deformation during swimming, resulting in a large change in depth. Through experimental comparison of various measurement effects, the correct rate of the calculated length is 95%.

[0107] The present invention also provides a complete detection system for the phenotypic characteristics of juvenile fish. The system includes multiple functional modules, specifically including an image acquisition module, a feature extraction and detection module, a 3D reconstruction and measurement module, a neural network prediction module, and a depth value correction module. The image acquisition module obtains the RGB image and depth map data of the juvenile fish through a binocular camera system. The data acquisition module also includes the calibration of the camera and the function of image preprocessing to ensure that the acquired images have sufficient accuracy and quality. The feature extraction and detection module accurately extracts the key point data of the juvenile fish and classifies and locates them through the trained YOLOv8-Pose model in combination with the detection module and the ContextGuideFPN fusion method. The 3D reconstruction and measurement module accurately calculates the body length and body width of the juvenile fish based on the binocular stereovision technology in combination with the depth information of the image. The neural network prediction module predicts the weight of the juvenile fish using the body length and body width data. The depth value correction module corrects the key point offset and depth value mutation caused by swimming to ensure the accuracy of the measurement results.

[0108] The following describes the detection device for the phenotypic characteristics of juvenile fish provided by the present invention. The detection device for the phenotypic characteristics of juvenile fish described below can be mutually corresponding and referred to the detection method for the phenotypic characteristics of juvenile fish described above.

[0109] As Figure 9 shown, the device includes: An image acquisition module 910, configured to acquire the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected; A key point extraction module 920, configured to extract key points of the juvenile fish to be detected based on the RGB image and determine multiple fish body key points of the juvenile fish to be detected; A three-dimensional coordinate determination module 930, configured to determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image; A phenotypic characteristic extraction module 940, configured to determine the phenotypic characteristics of the juvenile fish to be detected based on the three-dimensional coordinates, where the phenotypic characteristics include body length data, body width data, and weight data.

[0110] The detection device for the phenotypic characteristics of juvenile fish provided by the present invention extracts the fish body key points through the RGB image and the depth image, calculates the three-dimensional coordinates of the juvenile fish, and further obtains the phenotypic characteristics of the juvenile fish, including body length, body width, and weight data. The depth image provides the distance information of the real world, making the measurement results have a real physical scale. By combining the RGB image, the three-dimensional coordinates of the fish body key points can be accurately calculated, realizing the accurate extraction of the phenotypic characteristics of the juvenile fish and improving the accuracy of the detection of the phenotypic characteristics of the juvenile fish.

[0111] In one embodiment, the key point extraction module 920 is specifically configured to: Based on the RGB image, perform key point extraction on the to-be-detected juvenile fish, and determine multiple fish body key points of the to-be-detected juvenile fish, including: Input the RGB image into a key point detection model to obtain multiple fish body key points of the to-be-detected juvenile fish output by the key point detection model, where the key point detection model is trained based on juvenile fish image samples and fish body key point labels corresponding to the juvenile fish image samples.

[0112] In one embodiment, the key point extraction module 920 is further specifically configured to: Determine that the key point detection model is obtained by embedding a Feature Pyramid Network ContextGuideFPN in the neck layer on the basis of the YOLOv8-Pose model and replacing the C2f module with a C2f-SMA-CGLU module.

[0113] In one embodiment, the three-dimensional coordinate determination module 930 is specifically configured to: Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system, including: Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the depth value of the fish body key points; Connect multiple fish body key points in the RGB image to obtain multiple connection lines; Based on the depth values of a preset number of pixel points near the fish body key points and on the connection lines, adjust the depth value of the fish body key points to obtain an adjusted depth value; Based on the adjusted depth value and the two-dimensional coordinates corresponding to the fish body key points, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system.

[0114] In one embodiment, the phenotypic feature extraction module 940 is specifically configured to: Based on the three-dimensional coordinates, determine the phenotypic features of the to-be-detected juvenile fish, including: Based on the three-dimensional coordinates, determine the body length data and body width data of the to-be-detected juvenile fish; Input the body length data and the body width data into a weight prediction model to obtain weight data output by the weight prediction model, where the weight prediction model is trained based on length-width data samples and weight data labels corresponding to the length-width data samples.

[0115] In one embodiment, the key point extraction module 920 is further specifically configured to: Determine that the key points of the fish body include the fish mouth point, the midpoint of the tail, the midpoint of the caudal fin part, the dorsal fin point, and the pelvic fin point.

[0116] Figure 10 Illustrates a schematic physical structure diagram of an electronic device, such as Figure 10 As shown, the electronic device may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040. Among them, the processor 1010, the communication interface 1020, and the memory 1030 complete mutual communication through the communication bus 1040. The processor 1010 can call the logical instructions in the memory 1030 to execute the method for detecting the phenotypic characteristics of juvenile fish, and the method includes: obtaining the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected; Based on the RGB image, perform key point extraction on the juvenile fish to be detected, and determine multiple key points of the fish body of the juvenile fish to be detected; Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system; Based on the three-dimensional coordinates, determine the phenotypic characteristics of the juvenile fish to be detected, and the phenotypic characteristics include body length data, body width data, and body weight data.

[0117] In addition, when the logical instructions in the above-mentioned memory 1030 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0118] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the juvenile fish phenotypic feature detection method provided by the above-mentioned various methods. The method includes: obtaining an RGB image of the juvenile fish to be detected and a depth image of the juvenile fish to be detected; Based on the RGB image, perform key point extraction on the juvenile fish to be detected to determine multiple fish body key points of the juvenile fish to be detected; Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system; Based on the three-dimensional coordinates, determine the phenotypic features of the juvenile fish to be detected. The phenotypic features include body length data, body width data, and body weight data.

[0119] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the juvenile fish phenotypic feature detection method provided by the above-mentioned various methods. The method includes: obtaining an RGB image of the juvenile fish to be detected and a depth image of the juvenile fish to be detected; Based on the RGB image, perform key point extraction on the juvenile fish to be detected to determine multiple fish body key points of the juvenile fish to be detected; Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system; Based on the three-dimensional coordinates, determine the phenotypic features of the juvenile fish to be detected. The phenotypic features include body length data, body width data, and body weight data.

[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting the phenotypic characteristics of juvenile fish, characterized in that, Including: Obtain the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected; Based on the RGB image, perform key point extraction on the juvenile fish to be detected, and determine multiple fish body key points of the juvenile fish to be detected; Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system; Based on the three-dimensional coordinates, determine the phenotypic characteristics of the juvenile fish to be detected, where the phenotypic characteristics include body length data, body width data, and body weight data.

2. The method for detecting phenotypic characteristics of juvenile fish according to claim 1, wherein The step of performing key point extraction on the juvenile fish to be detected based on the RGB image and determining multiple fish body key points of the juvenile fish to be detected includes: Input the RGB image into a key point detection model to obtain multiple fish body key points of the juvenile fish to be detected output by the key point detection model. The key point detection model is trained based on juvenile fish image samples and the fish body key point labels corresponding to the juvenile fish image samples.

3. The juvenile fish phenotypic trait detection method according to claim 2, characterized in that, The construction process of the key point detection model includes: Use a detection module to replace the C2f module in the backbone network of the YOLOv8-Pose model to obtain an improved YOLOv8-Pose model. The detection module is constructed by connecting a collaborative multi-attention SMA module and a convolutional gated linear unit CGLU module through a residual connection; Embed a Feature Pyramid Network ContextGuideFPN into the neck network of the improved YOLOv8-Pose model to obtain the key point detection model.

4. The fry phenotype feature detection method according to claim 1, wherein The step of determining the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image includes: Based on the two-dimensional coordinates of the fish body key points in the RGB image and the depth image, determine the depth value of the fish body key points; Connect multiple fish body key points in the RGB image to obtain multiple connecting lines; Based on the depth values of a preset number of pixel points near the fish body key points and on the connecting lines, adjust the depth value of the fish body key points to obtain an adjusted depth value; Based on the adjusted depth value and the two-dimensional coordinates corresponding to the fish body key points, determine the three-dimensional coordinates of the fish body key points in the three-dimensional world coordinate system.

5. The method for detecting the phenotypic characteristics of juvenile fish according to claim 1, wherein The step of determining the phenotypic characteristics of the juvenile fish to be detected based on the three-dimensional coordinates includes: Based on the three-dimensional coordinates, determine the body length data and body width data of the juvenile fish to be detected; Input the body length data and the body width data into a body weight prediction model to obtain the body weight data output by the body weight prediction model. The body weight prediction model is trained based on length-width data samples and the body weight data labels corresponding to the length-width data samples.

6. The method for detecting the phenotypic characteristics of juvenile fish according to claim 1, wherein The fish body key points include the fish mouth point, the midpoint of the tail, the midpoint of the caudal fin part, the dorsal fin point, and the pelvic fin point.

7. A detection device for the phenotypic characteristics of juvenile fish, characterized in that, Including: An image acquisition module for obtaining the RGB image of the juvenile fish to be detected and the depth image of the juvenile fish to be detected; A key point extraction module, configured to extract key points of the to-be-detected juvenile fish based on the RGB image, and determine a plurality of fish body key points of the to-be-detected juvenile fish; A three-dimensional coordinate determination module, configured to determine three-dimensional coordinates of the fish body key points in a three-dimensional world coordinate system based on two-dimensional coordinates of the fish body key points in the RGB image and the depth image; A phenotypic feature extraction module, configured to determine phenotypic features of the to-be-detected juvenile fish based on the three-dimensional coordinates, where the phenotypic features include body length data, body width data, and body weight data.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the juvenile fish phenotypic feature detection method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the juvenile fish phenotypic feature detection method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the juvenile fish phenotypic feature detection method according to any one of claims 1 to 6.