Oat flag leaf phenotype calculation method and system based on RGB-D camera
Through the method based on RGB-D camera, the image and depth information of oat flag leaves are collected, image preprocessing, key point detection and segmentation are performed, and the flag leaves angle and length are calculated, which solves the problem of low accuracy in the calculation of oat flag leaves phenotype parameters in the prior art, and achieves higher calculation accuracy and applicability.
Patent Information
- Application Number
- CN202510148242.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-13
AI Technical Summary
The existing methods for obtaining crop phenotypic parameters have problems with low accuracy, especially in the calculation of oat flag leaf phenotypic parameters.
Using an RGB-D camera-based method, the RGB images and depth images of oat flag leaves were collected, and image preprocessing, key point detection, image segmentation and phenotypic calculations were performed, including calculations of flag leaves angle and flag leaves length.
The calculation accuracy of oat flag leaf phenotype parameters is improved, and the problem of large errors in traditional methods is overcome. This method is not harmful to the plants and is suitable for large-scale applications.
Smart Images

Figure CN120147349A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of crop phenotype calculation, and particularly relates to a method for calculating the phenotype of oat flag leaves based on an RGB-D camera. Background Art
[0002] Oat is an important miscellaneous grain crop that can be used as both food and feed, and is widely planted in temperate regions north of 40 degrees north latitude in Europe, North America, and Asia. The phenotypic traits manifested during the growth process are an intuitive reflection of the genetic characteristics of genes and are also an important reference for oat breeding work. As the main organ for plants to obtain energy, the leaf angle is a key determinant of plant structure and grain yield. The oat flag leaf is the first leaf below the ear. Although the flag leaf has a lifespan of only about 40 days, it is crucial for the later flowering, filling, and grain formation of oats. The flag leaf angle is the angle between the oat stem and the flag leaf. The flag leaf angle is an important parameter for describing the crop canopy structure, has a significant impact on the photosynthesis intensity, yield, etc. of the crop, and is also an important reference basis for breeding oat varieties with a compact plant type. Therefore, accurately and non-destructively obtaining the phenotypic parameters of oat flag leaves is crucial for the research of oat plants and the breeding of high-yield varieties.
[0003] Traditional methods for obtaining crop phenotypic parameters include manual measurement, two-dimensional image measurement, and precise instrument measurement. Manual measurement is time-consuming and laborious, and depends on manual experience and sensory judgment, with problems such as low efficiency and large errors, and it is easy to damage the plants during the measurement process. When measuring crop phenotypic parameters based on two-dimensional image technology, key spatial and volume information is easily lost during the data conversion process from three-dimensional (3D) to two-dimensional state, and there are large errors when measuring parameters such as leaf area, leaf angle, and leaf azimuth angle from different angles. In crop phenotype research, laser scanners with high resolution are often used for 3D reconstruction of plants and estimation of plant traits, but their scanning 3D imaging speed is slow and the equipment is expensive. In addition, it is difficult to correctly extract and process the three-dimensional point clouds generated by laser scanners from a large amount of three-dimensional data. The high cost and limited availability of laser scanning equipment have hindered its widespread application. Therefore, RGB-D depth cameras that can provide color images, depth, and infrared emitters are widely used in the field of 3D plant phenotype analysis.
[0004] However, the existing methods for obtaining crop phenotypic parameters currently have the problem of low accuracy. Therefore, how to improve the accuracy of oat flag leaf phenotypic parameters is an urgent problem to be solved in this field. Summary of the Invention
[0005] The object of the present invention is to provide an oat flag leaf phenotype calculation method, system and device based on an RGB-D camera in view of the defects of the prior art. The oat flag leaf phenotype calculation method based on an RGB-D camera of the present invention uses an RGB-D camera to collect RGB images and depth images of oat flag leaves; performs image preprocessing on the RGB images of oat flag leaves; inputs the preprocessed RGB images into a key point detection model to perform flag leaf key point detection to obtain oat flag leaf key point information; based on the key point information, performs flag leaf segmentation based on an image segmentation model to obtain a flag leaf segmentation result; performs flag leaf phenotype calculation based on the depth image, key point information and flag leaf segmentation result, and the flag leaf phenotype includes flag leaf angle and flag leaf length, thereby solving the problem of low accuracy of existing oat flag leaf phenotype calculation methods.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] The present invention provides an oat flag leaf phenotype calculation method based on an RGB-D camera, which is characterized by including the following steps:
[0008] S1. Use an RGB-D camera to collect RGB images and depth images of oat flag leaves;
[0009] S2. Perform image preprocessing on the RGB images of the oat flag leaves;
[0010] S3. Input the preprocessed RGB images into a key point detection model to perform flag leaf key point detection to obtain oat flag leaf key point information;
[0011] Among them, the key point detection model is an improved RTMPose model, and the improved RTMPose model is used for flag leaf key point detection; the improved RTMPose model is specifically: using StarNet as the backbone network and introducing a two-dimensional selective scanning SS2D reconstruction Head part;
[0012] S4. Based on the key point information, perform flag leaf segmentation based on an image segmentation model to obtain a flag leaf segmentation result;
[0013] S5. Perform flag leaf phenotype calculation based on the depth image, key point information and flag leaf segmentation result, and the flag leaf phenotype includes flag leaf angle and flag leaf length.
[0014] Further, in step S2, the image preprocessing includes distortion correction and stereo correction.
[0015] Further, in step S3, specifically including performing flag leaf key point detection using an improved RTMPose model: taking the RGB image of the oat flag leaf as input, using the backbone network StarNet for feature extraction, outputting a multi-scale feature map, inputting the feature map of the last layer into the Head module, first passing through a DwConv layer which can reduce the computational amount and effectively extract high-level features, then linearly segmenting the output feature map. The feature map of the first branch is input into the SS2D module which scans the local area of the feature image horizontally, vertically and in their reverse directions through the Cross-Scan Module, processes the scanning results using the Selective State Space Block, and combines with the SiLU activation function to increase the non-linear expression ability of the model; the second branch passes through a SiLU activation layer and the results output by the first branch are combined through feature combination, and finally the classification feature maps of the X-axis and Y-axis are output to represent the key point coordinate prediction.
[0016] Further, in step S4, the image segmentation model is SegmentAnythingModel.
[0017] Further, after step S4, it includes: based on the flag leaf segmentation result, performing depth information interpolation and filtering to obtain the oat flag leaf point cloud; the performing depth information interpolation includes: using the oat plant mask obtained by image segmentation to extract the non-zero depth information in the corresponding depth image area, calculating the gradient and angle of the depth information, evaluating the interpolation contribution of the known depth value points to the points with zero depth information, constructing the zero-based gradient direction weight ω angle,i and the spatial distance weight ω dinstance,i ; ω angle,i represents the similarity between the gradient direction of the known depth point and the gradient direction of the target point, and ω dinstance,i represents the influence of the spatial distance between the known depth point and the target point on the interpolation contribution; then depth prediction is performed on the points with zero depth information within the mask area, and finally bilateral filtering is used to make the predicted depth information smoother and more natural;
[0018]
[0019] where, d i represents the Euclidean distance between the target point (x t , y t ) and the known point (x i , y i ), θ t and θ i respectively represent the gradient directions of the target point and the known point; ω combined,iThe larger it is, the greater the contribution of the known point to the interpolation of the target point, and the depth value z of the target point to be interpolated t is as follows:
[0020]
[0021] where w i is the weight of the i-th known point to the interpolation point, and z i is the depth value of the i-th known point
[0022] Furthermore, in step S5, the flag leaf angle calculation step includes: based on the depth image and key point information, obtaining the depth information of each key point, and using the parameters of the camera to convert the two-dimensional key point coordinates and depth information into three-dimensional coordinates (x, y, z) through the projection model of the camera, specifically as follows:
[0023]
[0024] where (c x , c y ) is the principal point of the camera, (f x , f y ) is the focal length of the camera, d is the depth value obtained from the depth map, and (u, v) are the key point coordinates;
[0025] The flag leaf angle calculation formula is as follows:
[0026]
[0027] where point C is the point on the leaf sheath and is the vertex of the flag leaf angle. The point on the flag leaf closest to the leaf sheath is point B, and point A is on the stem above the leaf sheath.
[0028] Furthermore, in step S5, the flag leaf length calculation step includes: obtaining the three-dimensional coordinates of the key points by means of three-dimensional reconstruction using the interpolated depth image, extracting the three-dimensional coordinates of all key points representing the flag leaf length, and calculating the distance between the key points using the Euclidean distance formula in three-dimensional space; for any two three-dimensional coordinate points P i (X i , Y i , Z i ) and P i+1 (X i+1 , Y i+1 , Z i+1 ), their three-dimensional distance d i,i+1 can be calculated by the following formula:
[0029]
[0030] Sum the distances between key points along the key points from the leaf tip to the leaf sheath in sequence to obtain the true leaf length L leaf ,
[0031]
[0032] The present invention also provides an oat flag leaf phenotype calculation system based on an RGB-D camera, which is characterized in that the oat flag leaf phenotype calculation system executes the oat flag leaf phenotype calculation method based on the RGB-D camera, including: an image acquisition module, an image preprocessing module, a flag leaf key point detection module, a flag leaf segmentation module, and a flag leaf phenotype calculation module.
[0033] The image acquisition module uses an RGB-D camera to acquire the RGB image and depth image of the oat flag leaf;
[0034] The image preprocessing module performs image preprocessing on the RGB image of the oat flag leaf;
[0035] The flag leaf key point detection module inputs the preprocessed RGB image into the key point detection model to perform flag leaf key point detection to obtain oat flag leaf key point information;
[0036] The flag leaf segmentation module performs flag leaf segmentation based on the key point information and the image segmentation model to obtain a flag leaf segmentation result;
[0037] The flag leaf phenotype calculation module performs flag leaf phenotype calculation based on the depth image, key point information, and flag leaf segmentation result, and the flag leaf phenotype includes flag leaf angle and flag leaf length.
[0038] The present invention also provides a computer device, which includes a memory and a processor, and a computer program is stored on the memory. When the processor executes the computer program, the above method is implemented.
[0039] Compared with the prior art, it has the following beneficial effects:
[0040] The oat flag leaf phenotype calculation method based on an RGB-D camera of the present invention uses an RGB-D camera to collect RGB images and depth images of oat flag leaves; performs image preprocessing on the RGB images of oat flag leaves; inputs the preprocessed RGB images into a key point detection model to detect the key points of the flag leaves, obtaining the key point information of oat flag leaves; based on the key point information, performs flag leaf segmentation using an image segmentation model to obtain the flag leaf segmentation result; calculates the flag leaf angle and flag leaf length based on the depth image, key point information, and flag leaf segmentation result; the present invention uses an RGB-D camera to collect RGB images and depth images of oat flag leaves, employs an improved RTMPose model for flag leaf key point detection, uses an image segmentation model for flag leaf segmentation, and then combines the depth information image to calculate the oat flag leaf phenotype parameters according to the pinhole imaging principle, improving the accuracy of the oat flag leaf phenotype calculation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0042] Figure 1 Schematic diagram of the oat flag leaf phenotype calculation method based on an RGB-D camera provided by an embodiment of the present invention.
[0043] Figure 2 Frame diagram of the oat flag leaf phenotype calculation based on an RGB-D camera provided by an embodiment of the present invention.
[0044] Figure 3 Schematic diagram of the improved RTMPose provided by an embodiment of the present invention.
[0045] Figure 4 Flow chart of depth information interpolation provided by an embodiment of the present invention.
[0046] Figure 5 Schematic diagram of the oat flag leaf phenotype calculation system based on an RGB-D camera provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0048] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0049] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but it is not intended to limit the present invention.
[0050] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0051] It should be further understood that the term " / and" used in this specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0052] The present invention proposes a method for calculating the phenotypic traits of oat flag leaves based on an RGB-D camera. As Figure 1 shown, the method for calculating the phenotypic traits of oat flag leaves based on an RGB-D camera includes the following steps S1 to S5. The framework diagram for calculating the phenotypic traits of oat flag leaves based on an RGB-D camera is as Figure 2 shown.
[0053] S1. Use an RGB-D camera to collect the RGB image and depth image of the oat flag leaf.
[0054] The non-destructive detection of phenotypic traits of field crops is very important for crop breeding. A ground mobile platform equipped with sensors can efficiently and accurately obtain the phenotypic traits of crops. The present invention proposes a method for collecting three-dimensional data of the top leaves in the field applicable to Gramineae. Taking oats as the research object, a consumer-grade RGB-D camera installed on a ground mobile platform is used to automatically collect the RGB image and depth image of the oat flag leaf.
[0055] The sowing time of oats in the National Germplasm Repository (Southwest) (30°47′N, 103°76′E) was early November 2023. Image data of the flag leaf angle at three stages, namely the heading stage, flowering stage, and milk ripening stage of oats, were mainly collected. The flag leaf is the first leaf near the ear after the heading of oats or wheat and is an important vegetative organ of gramineous plants. The collection time was from March 2024 to mid-May 2024, the shooting time was from 8:00 to 12:30, and the collection frequency was once a week. The collection device was the Orbbec GminiPro depth camera. Compared with the depth cameras of the kinect and realsensor series, this camera has a lower cost. With this camera, RGB images can be obtained for object detection and image segmentation, and at the same time, the obtained depth information facilitates our 3D reconstruction. The true oat phenotypic values can be directly calculated without the aid of markers. At the same time, based on the SDK open source by the camera community, a software package was developed in combination with pyqt5 to develop a ROS (Robot Operating System) data acquisition system, which can be deployed on smaller-sized and lower-power-consuming embedded devices, making it more convenient to collect data in the field environment.
[0056] When collecting data, place the depth camera on a horizontally placed bracket, adjust the height of the bracket to the position of the flag leaf of the oat plant. With this system, a pair of color and depth images of the oat flag leaf angle data can be collected at one time, or a bag file in the ROSBAG file format can be recorded. Set the single recording time to 30 seconds, and color and depth image data can be obtained by playback. The resolution of the collected color and depth images is 640x480 and 640x400 pixels, and the images are saved in the "png" format. All data images were collected in the natural field environment. In order to overcome the influence of the field environment on data quality and improve the accuracy of object detection and image segmentation, a black cardboard was used as the background for data collection. To enrich the types of the dataset, the early training dataset also came from the color images of the oat flag leaf angle collected by a XiaoMi12 smartphone, with a resolution of 1080x1920. The validation data came from oats planted in Huzhu County, Qinghai. The image collection scheme was the same as above. At the same time, tools such as a tape measure and a protractor were used to collect the true oat flag leaf phenotypic data for validation.
[0057] S2. Perform image preprocessing on the RGB image of the oat flag leaf.
[0058] In step S2, the image preprocessing includes distortion correction and stereo correction.
[0059] Specifically, collect the RGB and depth images of the oat flag leaf, and then perform image correction on the collected RGB image, including preprocessing operations such as distortion correction and stereo correction.
[0060] S3. Input the preprocessed RGB image into the key point detection model to perform the key point detection of the flag leaf, and obtain the oat flag leaf key point information. Among them, the key point detection model is an improved RTMPose model, and the improved RTMPose model is used for the key point detection of the flag leaf. The specific improved RTMPose model is as follows: StarNet is used as the backbone network, and the two-dimensional selective scanning SS2D is introduced to reconstruct the Head part.
[0061] In step S3, using the improved RTMPose model for the key point detection of the flag leaf specifically includes:
[0062] The more compact and efficient network model StarNet is used as the backbone network. In order to further reduce computing resources, the structure of the head is reconstructed. The RGB image of the oat flag leaf is used as the input, and the backbone network StarNet is used for feature extraction to output a multi-scale feature map. The feature map of the last layer is input into the Head module. First, it passes through a DwConv layer, which can reduce the amount of calculation and effectively extract high-level features. Then, the output feature map is linearly segmented. The feature map of the first branch is input into the SS2D module, which scans the local area of the feature image horizontally, vertically, and in the reverse direction through the Cross-Scan Module, and uses the Selective StateSpace Block to process the scanning results, combined with the SiLU activation function to increase the non-linear expression ability of the model. The second branch passes through a SiLU activation layer, and the result output by the first branch undergoes feature combination. Finally, the classification feature maps of the X-axis and Y-axis are output to represent the key point coordinate prediction. The improved RTMPose model is as Figure 3 shown.
[0063] The present invention uses 8 points to characterize the length of the flag leaf. In addition, two points are marked on the stem above and below the leaf sheath. Taking the point on the leaf sheath as the vertex C of the flag leaf angle, this point, the point B closest to the leaf sheath, and the point A on the stem above the leaf sheath form and two vectors. The flag leaf angle is obtained by calculating the included angle degree between the vectors. Similarly, by connecting all the key points passed from the point on the leaf sheath to the leaf tip of the flag leaf, the distance between points is calculated to characterize the length of the flag leaf.
[0064] In terms of flag leaf key point detection, RTMPose uses a top-down method. It uses existing object detection algorithms to obtain bounding boxes and then performs key point detection. RTMPose is an efficient and accurate pose estimation algorithm mainly used for human pose estimation tasks. Its design goal is to improve the accuracy and speed of pose estimation so that it can perform well in real-time applications. RTMPose uses CSPNeXt as the backbone network. The head consists of a large kernel convolutional layer of 7x7, an MLP layer, and a gated attention unit (GAU). Then, SimCC is used to classify the representation of key point detection as the coordinates of the x-axis and y-axis of each key point to predict the horizontal and vertical positions of the key points. The model uses KL divergence as the loss function to optimize the model parameters. Although this model can achieve over 90 FPS on an Intel i7-11700 CPU, over 430 FPS on an NVIDIA GeForce GTX 1660Ti GPU, and over 35 FPS on a Snapdragon 865 chip, and the detection results have good performance, due to the limited computing resources and cost of the field phenotype platform developed based on embedded devices, the present invention uses a more compact and efficient network model, StarNet, as the backbone network.
[0065] Benefiting from the fact that mamba reduces the computational complexity from quadratic to linear and also has excellent performance compared with the dominant Transformer models, the Head part of RTMPose is redesigned in this paper. Two-dimensional selective scan (2D Selective Scan, SS2D) is introduced to replace the original Transformer module to improve the accuracy and efficiency of the model. First, the Cross-Scan Module (CSM) adopts a multi-directional scan strategy, scanning the local area of the image horizontally, vertically, and in their reverse directions. Each pixel can effectively integrate the information of all other pixels from different directions, and can better capture the important global information in the image. Second, the Selective State Space Block is used to process the scan results, capturing the long-range dependencies in the input features through linear transformation, further enhancing the feature representation ability. This improvement not only maintains the real-time performance of the RTMPose model but also further improves its performance in complex visual scenarios, achieving a better balance between accuracy and efficiency.
[0066] S4. Based on the key point information, perform flag leaf segmentation based on the image segmentation model to obtain the flag leaf segmentation result.
[0067] In step S4, the image segmentation model is the SegmentAnything Model.
[0068] In smart agriculture, rich image information can be obtained with the help of sensor devices such as RGB cameras and drones. Applying image segmentation can extract important information from images. For example, by segmenting plant leaves, fruits, or roots, phenotypic data can be extracted to analyze the growth status of plants. Common threshold-based methods require consistent image quality to use global thresholds for segmentation to improve segmentation efficiency. However, threshold segmentation requires a clear distinction between the target and the background and is sensitive to changes in illumination, making it suitable for image information obtained under controlled indoor environments. It is difficult to ensure consistent imaging conditions for images obtained from outdoor observations, and the illumination conditions vary greatly. Threshold-based segmentation methods need to frequently adjust the threshold to segment each image. Therefore, a deep learning-based segmentation method is used to automatically segment the flag leaves, which does not require manual feature design and can significantly improve the accuracy and robustness of segmentation. Deep learning models can also be seamlessly integrated with tasks such as detection and classification to form multi-task models for handling more tasks.
[0069] Flag leaf segmentation. SAM is a general image segmentation model that can automatically identify and segment any object in an image. SAM is based on ViT (Vision Transformer) and has a powerful global information modeling ability. Through multi-task learning and large-scale data training, it has excellent segmentation performance and robustness. It supports interactive segmentation, and users can generate accurate target segmentation results by inputting images and prompt information (such as key points, boxed areas, etc.). Therefore, after obtaining the key points of the flag leaf, to further extract the morphological features of the flag leaf, the detection box and key points on the flag leaf are used as prompts to enhance its recognition and segmentation ability for the flag leaf area, and the SegmentAnythingModel (SAM) model is further combined for accurate segmentation of the flag leaf. The segmentation result output by the SAM model is binarized, where the white area represents the flag leaf and the black area represents the background. The binarized segmentation result will be input into the phenotypic calculation module for interpolating missing depth information and for filtering to obtain the oat flag leaf point cloud and directly calculating the leaf angle and length of the flag leaf.
[0070] After step S4, it includes: based on the flag leaf segmentation result, performing depth information interpolation and filtering to obtain the oat flag leaf point cloud.
[0071] Depth compensation. Since the leaf tips and stems of oat plants are relatively slender, it is easy to cause missing depth information due to weak reflection of infrared light and the influence of natural light in the field. It is necessary to compensate for the missing depth information of the leaf tips and stems of oat plants.
[0072] The depth information interpolation is as Figure 4As shown, it includes: extracting the non-zero depth information in the corresponding depth image region using the oat plant mask obtained by image segmentation, calculating the gradient and angle of the depth information, evaluating the interpolation contribution of known depth value points to points with zero depth information, and constructing a zero-based gradient direction weight ω angle,i and a spatial distance weight ω dinstance,i ; ω angle,i represents the similarity between the gradient direction of the known depth point and the gradient direction of the target point, and ω dinstance,i represents the influence of the spatial distance between the known depth point and the target point on the interpolation contribution; then, depth prediction is performed on the points with zero depth information in the mask region, and finally, bilateral filtering is used to make the predicted depth information smoother and more natural;
[0073]
[0074] where d i represents the Euclidean distance between the target point (x t , y t ) and the known point (x i , y i ), θ t and θ i represent the gradient directions of the target point and the known point respectively; the larger ω combined,i , the greater the interpolation contribution of the known point to the target point, and the depth value z t of the target point to be interpolated is as follows:
[0075]
[0076] where w i is the weight of the i-th known point for the interpolation point, and z i is the depth value of the i-th known point.
[0077] S5. Calculate the flag leaf phenotype based on the depth image, key point information, and flag leaf segmentation result, where the flag leaf phenotype includes the flag leaf angle and the flag leaf length.
[0078] In step S5, the flag leaf angle calculation step includes: based on the depth image and key point information, obtaining the depth information of each key point, and using the parameters of the camera to convert the two-dimensional key point coordinates and depth information into three-dimensional coordinates (x, y, z) through the camera's projection model, specifically as follows:
[0079]
[0080] where (c x , c y ) is the principal point of the camera, and (f x , f y) is the focal length of the camera, d is the depth value obtained from the depth map, and (u, v) are the coordinates of the key points.
[0081] Specifically, through the improved RTMPose algorithm, the key points (u, v) representing the flag leaf angle on the color image are detected. Then, the depth information (i.e., the z-axis coordinate) of each key point is obtained according to the corresponding depth image. Using the parameters of the camera, the two-dimensional key point coordinates and depth information are converted into three-dimensional coordinates (x, y, z) through the camera's projection model without reconstructing the entire scene.
[0082] The calculation formula for the flag leaf angle is as follows:
[0083]
[0084] Among them, point C is the point on the leaf sheath, which is the vertex of the flag leaf angle. The point B on the flag leaf that is closest to the leaf sheath, and point A is on the stem above the leaf sheath.
[0085] In addition, the two-dimensional angle of the flag leaf angle in a single RGB image can also be calculated simultaneously.
[0086] In step S5, the calculation steps for the flag leaf length include: obtaining the three-dimensional coordinates of the key points by means of three-dimensional reconstruction using the interpolated depth image, extracting the three-dimensional coordinates of all key points representing the flag leaf length, and using the Euclidean distance formula in three-dimensional space to calculate the distance between the key points; for any two three-dimensional coordinate points P i (X i , Y i , Z i ) and P i+1 (X i+1 , Y i+1 , Z i+1 ), their three-dimensional distance d i,i+1 can be calculated by the following formula:
[0087]
[0088] Sum up the distances between the key points along the key points from the leaf tip to the leaf sheath in sequence to obtain the true leaf length L leaf ,
[0089]
[0090] Specifically, using key points can better represent the length of non-erect flag leaves and reduce the error caused by leaf bending. Compared with the method of using a reference object, it is necessary to convert from pixels through the reference object standard to indirectly obtain the leaf length. In the present invention, the three-dimensional coordinates of the key points are obtained by means of three-dimensional reconstruction using the interpolated depth image, and by summing up the distances between the key points, the true leaf length can be directly measured.
[0091] In this paper, the flag leaf angle is obtained by calculating the vector angle formed by the key points on the flag leaf and leaf sheath, and the stem and leaf sheath. The flag leaf length is obtained by calculating the lengths of the lines connecting the tip of the flag leaf to 8 key points on the leaf sheath. Specifically, color images and depth images of oats at the booting stage, heading stage, flowering stage, and milk ripening stage are collected using an RGB-D camera. Two-dimensional key point coordinates representing the flag leaf angle and length are obtained through key point detection. Three-dimensional reconstruction based on RGB-D is used to obtain the three-dimensional key point coordinate information representing the flag leaf angle and leaf length, and the flag leaf angle and length data of oats are calculated to accelerate the superior breeding of the flag leaf angle and flag leaf length of oats. In order to improve the calculation accuracy of the flag leaf phenotype in this paper, a black background board is used to assist in photographing oat images. In addition, to overcome the influence of strong light irradiation and high ultraviolet intensity in the plateau environment on the lack of depth information of the RGB-D camera, the object detection frame and key points on the flag leaf are used as prompts, and the SAM model is used to segment the flag leaf area, avoiding introducing depth information other than the target area for depth information compensation. By calculating the angle and gradient direction of the depth information in the non-missing area and using the central depth value of the depth image as a constraint, the missing depth information is compensated.
[0092] Inspired by key point detection in the calculation of other crop phenotypes, the present invention proposes a method for calculating oat flag leaf phenotype parameters by integrating key point detection and image segmentation. First, RGB and depth images of oat flag leaves are collected, and then the collected RGB images are corrected; the preprocessed images are fed into a key point detection model to detect the key points representing the oat flag leaf phenotype. Then, the key point detection results are used as the prompt input of the SegmentAnything Model (SAM) to accurately segment the flag leaf. Finally, the respective output results are fed into the phenotype calculation module, and the phenotype parameters such as the leaf angle and leaf length of the oat flag leaf are calculated according to the pinhole imaging principle in combination with the depth information image.
[0093] Experimental results
[0094] Key point detection evaluation
[0095] Experiments were conducted using the dataset collected in this paper. The dataset follows the standard COCO format, and the training, test, and validation datasets are set in a ratio of 8:1:1. The experimental results are shown in Table 1. The model in this paper achieved a mAP (mean Average Precision) of 93.5%, showing a certain improvement compared to the 93.05% of the baseline model RTMPose.
[0096] Table 1 Flag leaf key point detection results based on the dataset of this paper
[0097]
[0098] Extraction of Flag Leaf Phenotypic Features
[0099] The Pearson correlation coefficient (r), coefficient of determination (R 2 ) and mean absolute error (MAE) were used as statistical indicators to evaluate the performance of the system in estimating flag leaf angle and flag leaf length.
[0100]
[0101] The flag leaf angle data of different dimensions calculated by the method in this paper were compared with the true values of the flag leaf angles measured manually on site. The results showed that there was a high correlation between the angles formed by the key points representing the flag leaf angle detected on the two-dimensional color image calculated directly and the true measurement values, where r > 0.95 and MAE < 6 for the two. The r between the flag leaf angle in the three-dimensional space calculated by combining the depth image and the true value measured manually was > 0.95 and MAE < 5. Compared with the two-dimensional calculation results, the result of three-dimensional reconstruction was better, and the flag leaf angle had a strong correlation with the true measured angle. In terms of flag leaf length calculation, common two-dimensional image-based methods could not directly obtain the true value, but only obtained the result expressed in pixels. The method in this paper combined with the depth image could directly calculate the true leaf length, and also showed a high correlation with the manual measurement value, r = 0.827, MAE = 1.121. In addition, this paper used two-dimensional and three-dimensional methods to calculate the flag leaf angle of images of the same oat plant at different angles. The results showed that the error range of the angle obtained by the three-dimensional reconstruction method was smaller than that of the two-dimensional key point calculation result compared with the true measured angle, which could reduce the dependence on the camera shooting angle and showed better robustness.
[0102] Results and Discussion: We compared the calculation results of flag leaf phenotypes in three-dimensional and two-dimensional spaces. The experimental results showed that there was a strong correlation (R > 0.95) between the flag leaf angles calculated by two-dimensional calculation and those calculated by three-dimensional reconstruction method compared with the manual measurement results. Compared with the two-dimensional calculation results, the three-dimensional reconstruction method was more capable of overcoming the influence brought by different shooting angles (MAE < 5), and the flag leaf length (R = 0.8 - 0.85, MAE = 1.121 cm). The method of the present invention could directly calculate the phenotypes such as the flag leaf angle and length of the flag leaf more accurately.
[0103] The present invention first uses an existing detector to obtain the flag leaf region of the oat color image acquired by an RGB-D camera, then detects the key points representing the flag leaf, and the detection results are used to directly calculate the flag leaf angle in two-dimensional space. At the same time, the SAM model is used as a prompt to segment the flag leaf to improve the accuracy of depth information interpolation and for the segmentation of the flag leaf point cloud. Finally, the key points are back-projected into the oat point cloud generated using depth information to calculate the three-dimensional phenotypic traits such as the flag leaf angle and flag leaf length of a single oat plant.
[0104] Figure 5 This is a calculation system for oat flag leaf phenotypes provided by an embodiment of the present invention. As Figure 5 shown, the calculation system for oat flag leaf phenotypes based on an RGB-D camera includes an image acquisition module, an image preprocessing module, a flag leaf key point detection module, a flag leaf segmentation module, and a flag leaf phenotype calculation module.
[0105] The image acquisition module uses an RGB-D camera to acquire RGB images and depth images of oat flag leaves;
[0106] The image preprocessing module performs image preprocessing on the RGB images of the oat flag leaves;
[0107] The flag leaf key point detection module inputs the preprocessed RGB images into a key point detection model to perform flag leaf key point detection and obtains oat flag leaf key point information;
[0108] The flag leaf segmentation module performs flag leaf segmentation based on the key point information and an image segmentation model to obtain a flag leaf segmentation result;
[0109] The flag leaf phenotype calculation module performs flag leaf phenotype calculation based on the depth image, the key point information, and the flag leaf segmentation result. The flag leaf phenotypes include flag leaf angle and flag leaf length.
[0110] The above-mentioned calculation system for oat flag leaf phenotypes based on an RGB-D camera can be implemented in the form of a computer program, and this computer program can run on a computer device.
[0111] The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a non-volatile storage medium and an internal memory.
[0112] The non-volatile storage medium can store an operating system and a computer program. This computer program includes program instructions. When the program instructions are executed, the processor can execute a calculation method for oat flag leaf phenotypes based on an RGB-D camera.
[0113] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.
[0114] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute a calculation method for oat flag leaf phenotypes based on an RGB-D camera.
[0115] The network interface is used for network communication with other devices. Those skilled in the art can understand that the above computer device structure is only a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0116] Wherein, the processor is used to run a computer program stored in the memory, and this program implements the oat flag leaf phenotype calculation method based on the RGB-D camera described in the first embodiment.
[0117] It should be understood that in the embodiments of this application, the processor may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0118] Those of ordinary skill in the art can understand that all or part of the processes in the methods of implementing the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program includes program instructions, and the computer program can be stored in a storage medium, and this storage medium is a computer-readable storage medium. These program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0119] The present invention also provides a storage medium. This storage medium may be a computer-readable storage medium. This storage medium stores a computer program, wherein when this computer program is executed by a processor, the processor executes an oat flag leaf phenotype calculation method based on an RGB-D camera described in the first embodiment.
[0120] The storage medium may be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc, etc., all kinds of computer-readable storage media that can store program codes.
[0121] Those of ordinary skill in the art will appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described in terms of function in the above description. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0122] In several embodiments provided by the present invention, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0123] The steps in the method embodiments of the present invention can be adjusted, combined, and deleted according to actual needs. The units in the apparatus embodiments of the present invention can be combined, divided, and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0124] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0125] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A method for calculating oat flag leaf phenotype based on RGB-D camera, characterized in that: Includes steps: S1. Use an RGB-D camera to collect RGB images and depth images of oat flag leaves; S2, performing image preprocessing on the RGB image of the oat flag leaf; S3, inputting the preprocessed RGB image into the key point detection model to perform flag leaf key point detection to obtain oat flag leaf key point information; The key point detection model is an improved RTMPose model, which is used to detect the key points of flag leaves; the improved RTMPose model is specifically: StarNet is used as the backbone network, and two-dimensional selective scanning SS2D is introduced to reconstruct the Head part; S4, performing flag leaf segmentation based on the key point information and an image segmentation model to obtain a flag leaf segmentation result; S5. Calculate a flag leaf phenotype based on the depth image, key point information, and flag leaf segmentation result, where the flag leaf phenotype includes a flag leaf angle and a flag leaf length.
2. The method according to claim 1, characterized in that In step S2, the image preprocessing includes distortion correction and stereo correction.
3. The method according to claim 1, characterized in that In step S3, the improved RTMPose model is used to detect the key points of flag leaves, specifically including: taking the RGB image of the oat flag leaf as input, using the backbone network StarNet for feature extraction, outputting a multi-scale feature map, inputting the feature map of the last layer into the Head module, first passing through a DwConv layer, which can reduce the amount of calculation and effectively extract high-level features, and then linearly segmenting the output feature map, the feature map of the first branch is input into the SS2D module, which scans the local area of the feature image horizontally, vertically and in reverse through the Cross-Scan Module, uses the selective state space module Selective State SpaceBlock to process the scanning result, and combines the SiLU activation function to increase the nonlinear expression ability of the model; the second branch passes through a SiLU activation layer and the result output by the first branch is combined with the features, and finally outputs the classified feature maps of the X-axis and Y-axis for representing the key point coordinate prediction.
4. The method according to claim 1, characterized in that: In step S4, the image segmentation model is SegmentAnything Model.
5. The method according to claim 4, characterized in that After step S4, it includes: based on the flag leaf segmentation result, interpolating the depth information and filtering to obtain the oat flag leaf point cloud; the interpolating the depth information includes: using the oat plant mask obtained by image segmentation to extract the non-zero depth information in the corresponding depth image area, which is used to calculate the gradient and angle of the depth information, evaluate the interpolation contribution of the known depth value point to the point with zero depth information, and construct a zero-based gradient direction weight ω angle,i and spatial distance weight ω dinstance,i ;ω angle,i Represents the similarity between the gradient direction of the known depth point and the gradient direction of the target point, ω dinstance,i Indicates the influence of the spatial distance between the known depth point and the target point on the interpolation contribution; then the depth is predicted for the points with zero depth information in the mask area, and finally bilateral filtering is used to make the predicted depth information smoother and more natural; Among them, d i Represents the target point (x t ,y t ) and the known point (x i ,y i ), θ t and θ i Represent the gradient directions of the target point and the known point respectively; ω combined,i The larger the value, the greater the interpolation contribution of the known point to the target point. t as follows: where w i is the weight of the i-th known point to the interpolation point, z i is the depth value of the i-th known point.
6. The method according to claim 1, characterized in that In step S5, the flag leaf angle calculation step includes: based on the depth image and key point information, obtaining the depth information of each key point, and using the camera parameters to convert the two-dimensional key point coordinates and depth information into three-dimensional coordinates (x, y, z) through the camera's projection model, as follows: Among them, (c x ,c y ) is the principal point of the camera, (f x ,f y ) is the focal length of the camera, d is the depth value obtained from the depth map, and (u,v) is the key point coordinates; The formula for calculating the flag leaf angle is as follows: Among them, point C is the point on the leaf sheath and the vertex of the flag leaf angle. The point on the flag leaf closest to the leaf sheath is point B, and point A is located on the stem above the leaf sheath.
7. The method according to claim 5, characterized in that In step S5, the flag leaf length calculation step includes: obtaining the three-dimensional coordinates of the key points by means of three-dimensional reconstruction using the interpolated depth image, extracting the three-dimensional coordinates of all the key points representing the flag leaf length, and calculating the distance between the key points using the Euclidean distance formula in three-dimensional space; for any two three-dimensional coordinate points P i (X i , Y i , Z i ) and P i+1 (X i+1 , Y i+1 , Z i+1 ), the three-dimensional distance d between them i,i+1 It can be calculated by the following formula: Along the key points between the leaf tip and the leaf sheath, the distances between the key points are summed up in turn to obtain the true leaf length L leaf , 8. An oat flag leaf phenotype calculation system based on RGB-D camera, characterized in that: The oat flag leaf phenotype calculation system executes the oat flag leaf phenotype calculation method based on the RGB-D camera as claimed in claim 1, comprising: an image acquisition module, an image preprocessing module, a flag leaf key point detection module, a flag leaf segmentation module, and a flag leaf phenotype calculation module; The image acquisition module uses an RGB-D camera to collect RGB images and depth images of oat flag leaves; An image preprocessing module performs image preprocessing on the RGB image of the oat flag leaf; The flag leaf key point detection module inputs the preprocessed RGB image into the key point detection model to perform flag leaf key point detection and obtain the oat flag leaf key point information; A flag leaf segmentation module performs flag leaf segmentation based on the key point information and an image segmentation model to obtain a flag leaf segmentation result; The flag leaf phenotype calculation module calculates the flag leaf phenotype based on the depth image, key point information and flag leaf segmentation result, and the flag leaf phenotype includes the flag leaf angle and the flag leaf length.
9. A computer device, characterized in that: The device comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Phenotype data measurement method and system in semi-automatic field scene
CN121053548A