Soybean growth monitoring method based on depth image perception segmentation

By using depth image perception segmentation technology, soybean plant images and depth images are acquired, depth constraint cue points are generated, individual plant masks are segmented, and plant height and growth period are calculated. This solves the problems of low efficiency and poor accuracy in soybean growth monitoring and achieves rapid and stable growth monitoring.

CN121962886APending Publication Date: 2026-05-01INST OF PLANT PROTECTION CHINESE ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF PLANT PROTECTION CHINESE ACAD OF AGRI SCI
Filing Date
2025-12-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies for soybean growth monitoring suffer from low measurement efficiency, high labor intensity, and large subjective errors. Furthermore, they cannot simultaneously acquire multi-dimensional phenotypic information such as canopy width, leaf color, and pod quantity, leading to inaccurate monitoring results.

Method used

A depth-based image perception segmentation method is adopted. By acquiring soybean plant images and associated depth images, depth constraint cue points are generated, individual plant masks are segmented, and plant height is calculated using the depth median and the geometric relationship of the camera field of view. Growth features are extracted to determine the growth stage.

Benefits of technology

It achieves rapid, stable, and accurate soybean growth monitoring results, is suitable for greenhouse and close-range measurements, and improves the efficiency and accuracy of obtaining individual plant phenotypic indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962886A_ABST
    Figure CN121962886A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a soybean growth monitoring method based on depth image perception segmentation, and relates to the fields of plant phenotype analysis, biological breeding crop safety evaluation and the like. A soybean plant image and an associated depth image are obtained, a depth constraint prompt point is generated based on a depth value in a region of interest and a foreground depth critical value, and a single-plant mask is segmented from the soybean plant image according to the depth constraint prompt point. And converting the pixel height of the plant into the actual height of the plant by using the geometrical relationship between the depth median in the mask of the single plant and the view field of the camera. And then, extracting growth characteristics from the soybean plant image, and calculating a leaf yellowing ratio and a pod number according to the growth characteristics so as to judge a growth period level. According to the method, on the basis of the RGB-D image, the depth information is utilized to guide the image segmentation model to extract the single-plant soybean area, height measurement and growth period judgment are carried out, rapid and stable monitoring of soybean single-plant phenotypic indexes is achieved, and the accuracy of soybean growth monitoring results is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to technical fields such as plant phenotypic analysis and safety evaluation of bio-bred crops, and in particular to a soybean growth monitoring method based on depth image perception segmentation. Background Technology

[0002] Plant height, growth stage, and other crop phenotypic traits are important indicators for evaluating variety characteristics and growth performance. With the development of breeding technologies such as transgenics and gene editing, breeding efficiency has been significantly improved. However, the technology for rapid and accurate phenotypic measurement is relatively lagging behind, becoming a major bottleneck restricting the safety evaluation of new varieties' phenotypes and other traits, and also a significant constraint on promoting intelligent pest and disease control in production.

[0003] Crop phenotypic indicators such as plant height and growth stage generally rely on manual surveys and records, with periodic measurements used to determine the crop's growth period. However, manual surveys and records suffer from low measurement efficiency, high labor intensity, and significant subjective errors, making them unsuitable for large-scale, multi-time-point, high-throughput breeding trials and safety evaluations. For monitoring crop phenotypic traits and other growth characteristics, devices such as ultrasonic sensors, laser displacement meters, and rotary encoders can be used to measure the distance from the top of the plant to a reference plane.

[0004] For soybeans, an important economic crop, while the aforementioned growth monitoring methods offer high single-point measurement accuracy, the acquired results are mostly representative of local heights and cannot simultaneously obtain more phenotypic information such as canopy width, leaf color, and pod quantity, leading to inaccurate monitoring results. Furthermore, using laser sensors or laser scanners to construct plant point clouds to estimate three-dimensional structure suffers from problems such as large point cloud computation, inaccurate results, and long processing times, hindering its widespread adoption in high-throughput scenarios. Summary of the Invention

[0005] In view of this, embodiments of this application provide a soybean growth monitoring method based on depth image perception segmentation to solve the problem of inaccurate soybean growth monitoring results.

[0006] According to a first aspect of this application, a soybean growth monitoring method based on depth image perception segmentation is provided, the method comprising: Acquire soybean image data, which includes soybean plant images and associated depth images; the soybean plant images include regions of interest, which are image regions covering the soybean plant under test and the root growth baseline of the soybean plant under test; Depth constraint cue points are generated based on regions in the region of interest whose depth values ​​are less than the foreground depth threshold. The foreground depth threshold is set according to the distribution range of effective depth values ​​in the soybean image data. The soybean plant image is segmented according to the depth constraint cue points to obtain a single plant mask; Using the median depth within the single-plant mask and the geometric relationship of the camera's field of view, the pixel height of the soybean plant in the image is converted into the actual height of the plant. Growth features are extracted from the soybean plant image, and the proportion of yellowing leaves and the number of pods are calculated based on the growth features; the growth features include leaf color and pod target. The growth stage level is generated based on the actual height of the plant, the proportion of yellowing leaves, and the number of pods.

[0007] According to a second aspect of this application, a soybean growth monitoring system based on depth image perception segmentation is provided, the system comprising: The data acquisition module is used to acquire soybean image data, which includes soybean plant images and associated depth images; the soybean plant images include regions of interest, which are image regions covering the soybean plant under test and the root growth baseline of the soybean plant under test; The image preprocessing and alignment module is used to generate depth constraint cue points based on regions in the region of interest whose depth values ​​are less than the foreground depth threshold, wherein the foreground depth threshold is set according to the distribution range of effective depth values ​​in the soybean image data. A depth-guided segmentation module is used to segment the soybean plant image according to the depth constraint prompts to obtain a single plant mask; The plant height calculation module is used to convert the plant pixel height in the soybean plant image into the actual plant height by using the median depth within the single plant mask and the geometric relationship of the camera field of view. The color and pod feature extraction module is used to extract growth features from the soybean plant image and calculate the leaf yellowing ratio and pod number based on the growth features; the growth features include leaf color and pod target. The growth period determination module is used to generate a growth period level based on the actual height of the plant, the proportion of yellowing leaves, and the number of pods.

[0008] According to a third aspect of this application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described soybean growth monitoring method based on depth image perception segmentation.

[0009] According to a fourth aspect of this application, a storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described soybean growth monitoring method based on depth image perception segmentation.

[0010] By employing the above technical solution, this application provides a soybean growth monitoring method based on depth image perception segmentation. The method acquires soybean plant images and associated depth images, generates depth constraint cue points based on depth values ​​in the region of interest and foreground depth thresholds, and segments individual plant masks from the soybean plant image according to these cue points. Then, using the median depth within the individual plant mask and the geometric relationship of the camera's field of view, the plant pixel height is converted to the actual plant height. Next, growth features are extracted from the soybean plant image, and the proportion of yellowing leaves and the number of pods are calculated based on these features to determine the growth stage. This method can extract individual soybean plant regions using depth-guided image segmentation models based on RGB-D images, and perform height measurement and growth stage determination, achieving rapid and stable monitoring of individual soybean plant phenotypic indicators and improving the accuracy of soybean growth monitoring results.

[0011] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0012] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic flowchart of a soybean growth monitoring method based on depth image perception segmentation provided in an embodiment of this application; Figure 2 This is a schematic diagram of the overall process for soybean growth monitoring provided in the embodiments of this application; Figure 3 This is a schematic diagram of the prompt point filtering results provided in the embodiments of this application; Figure 4 This is a schematic diagram of the reproductive period determination process provided in the embodiments of this application; Figure 5 A scatter plot comparing the measurement results provided in the embodiments of this application; Figure 6 The measurement error distribution histogram provided in the embodiments of this application; Figure 7 A schematic diagram of the structure of a soybean growth monitoring system based on depth image perception segmentation provided in an embodiment of this application. Detailed Implementation

[0013] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0014] In this embodiment, the soybean growth monitoring method based on depth image perception segmentation can be applied to plant phenotypic analysis and crop growth monitoring, using color depth (RGB-D) images to measure the height of individual soybean plants and intelligently determine their growth stage. It should be noted that this method is not only applicable to soybean growth monitoring but can also be extended to other potted, greenhouse, or some field-grown crops.

[0015] Crop phenotypic traits are important indicators for evaluating variety characteristics and growth performance. Therefore, in some embodiments of this application, crop phenotypic analysis refers to measuring the phenotypic traits of crop plants through automated analysis and measurement methods, in order to promote the development of modern breeding technology and its safety evaluation technology.

[0016] In some embodiments, for measuring key phenotypic parameters such as crop height, devices such as ultrasonic sensors, laser displacement meters, and rotary encoders can be used to measure the distance from the top of the plant to a reference plane. These devices offer high single-point measurement accuracy, but they primarily acquire locally representative heights and cannot simultaneously obtain more phenotypic information such as crown width, leaf color, and pod number.

[0017] For example, plant point clouds can be constructed using laser sensors or laser scanners, and the 3D structure can be estimated based on these point clouds. However, this measurement method suffers from problems such as large point cloud computation, inaccurate results, and long processing times, hindering its adoption in high-throughput scenarios.

[0018] To improve the accuracy of monitoring results, in some embodiments, depth cameras can be used to collect depth information of crop plants and estimate crop phenotypic traits based on the depth information. However, due to problems such as insufficient resolution, significant interference from natural light, and poor robustness, depth cameras require additional shading structures when measuring crop phenotypic traits, making them unsuitable for complex field environments.

[0019] Therefore, in some embodiments, a color depth (RGB-D) camera can be used to simultaneously acquire depth maps and high-resolution color images. RGB-D cameras have advantages such as low cost, high resolution, and good environmental adaptability, and are therefore used in crop phenotyping.

[0020] Before using RGB-D images for height measurement and phenotypic analysis, accurate segmentation of plant regions is required. In some embodiments, image segmentation methods can be based on algorithms such as color thresholding, vegetation index, background modeling, or Otsu's thresholding, combined with morphological operations to distinguish crops from the background. These image segmentation methods perform well in simple backgrounds and with stable lighting, but they are prone to omissions or missegments in complex scenes with shadows, reflections, wall textures, or interference from multiple plants, leading to poor stability in subsequent height calculations.

[0021] Therefore, deep learning techniques can be combined to achieve agricultural image segmentation tasks. Semantic segmentation networks such as U-Net and DeepLab can achieve high accuracy in leaf segmentation and contour extraction. However, these models rely on a large amount of labeled data, resulting in high training costs, and their transferability across crops or scenes is limited. Although depth maps can be used as additional input or clustering can be performed at the point cloud level to enhance the robustness of segmentation, a dedicated training process is still required, leading to insufficient versatility of deep learning techniques.

[0022] Based on the soybean growth monitoring methods shown in the above embodiments, it can be seen that in greenhouse or close-range environments, the measurement of height and analysis of growth period of individual soybean plants still suffer from low efficiency and high labor intensity due to manual and contact methods, making it difficult to meet the needs of high-throughput experiments; image segmentation has poor robustness under complex backgrounds and light variation conditions, resulting in unstable height estimation; and the separate processes for height measurement and growth period determination, lacking a unified multimodal analysis framework, lead to inaccurate soybean growth monitoring results.

[0023] To address the issue of inaccurate soybean growth monitoring results, this application provides a soybean growth monitoring method based on depth image perception segmentation in some embodiments. This method automatically guides an image segmentation model to extract individual soybean plant regions using depth information from RGB-D images, and then achieves plant height measurement and intelligent determination of growth stage. This method eliminates the need for complex 3D point cloud reconstruction or extensive model training, enabling rapid and stable acquisition of individual soybean plant phenotypic indicators. It is suitable for various application scenarios, including greenhouse and close-range measurement.

[0024] The method can be applied to electronic devices with data processing capabilities. These electronic devices include, but are not limited to, computers, servers, mobile terminals, smart wearable devices, and industrial control machines. For ease of description, this application embodiment uses an electronic device as the execution subject of the method. It should be understood that the method can also be applied to other types of execution subjects, which are not illustrated in this application embodiment. Figure 1 As shown, the method includes: S101. Obtain soybean image data.

[0025] When monitoring soybean growth, soybean image data is required. This image data is obtained by capturing images of the soybean plant under test using an RGB-D sensor. Therefore, the soybean image data includes soybean plant images and associated depth images. These soybean plant images and associated depth images are correlated. The soybean plant images and associated depth images can be obtained by simultaneously capturing images of the same soybean plant under test at the same location using an RGB-D sensor, or a color camera combined with a depth camera.

[0026] For example, the electronic device can be configured with a data acquisition module, using an Intel Realsense D455 RGB-D camera. This camera can simultaneously output color and depth images, with a color image resolution of at least 1280×720 pixels and a depth image resolution of at least 640×480 pixels, at a frame rate of 30 frames per second. The camera is mounted on an adjustable-height tripod, with its optical axis essentially perpendicular to the potted soybean plant, at a working distance of approximately 0.8-1.0 meters to ensure that each soybean plant falls completely within the field of view and maintains good depth accuracy.

[0027] To facilitate image analysis of the soybean plants under test, the soybean plant image can include a Region of Interest (ROI). The ROI is an image region covering the soybean plant under test and its root growth baseline. Taking potted soybean monitoring as an example, the ROI can be a rectangular area marked on the soybean plant image, covering the soybean plant under test and the pot rim or soil baseline near its roots.

[0028] In some embodiments, to acquire soybean image data, the electronic device can first acquire the original color image and original depth image of the soybean plant to be tested using an RGB-D camera, and delineate the region of interest in the original color image to obtain an image of the soybean plant. The region of interest can be manually set by the user or determined by target recognition of the soybean plant to be tested in the original color image.

[0029] For example, electronic devices can run a soybean growth monitoring application to display a monitoring interface. Users can interactively control the interface to load a raw color image and its corresponding raw depth image. Then, using mouse operations, a rectangular area can be drawn on the raw color image, covering the individual soybean plant under test and the pot rim or soil reference near its roots, as the region of interest.

[0030] After defining the region of interest (ROI), the original depth image can be cropped into depth sub-images based on the ROI. Each depth sub-image includes the depth values ​​associated with pixels within the ROI. Then, by converting the units of the depth values ​​in the depth sub-images, an associated depth image is obtained.

[0031] For example, after drawing a rectangular region of interest in the original color image, the electronic device can crop out the depth sub-image corresponding to the rectangular region of interest from the original depth image, and convert the depth values ​​in the depth sub-image to represent the distance in millimeters, and remove pixels that are zero or invalid, to obtain an associated depth image that can be used for subsequent analysis.

[0032] S102. Generate depth constraint prompts based on regions in the region of interest whose depth values ​​are less than the foreground depth threshold.

[0033] After acquiring soybean image data, depth constraint cue points can be generated within the region of interest. For example... Figure 2 As shown, depth constraint cue points can be determined using the foreground depth threshold determined by Otsu's method, or obtained by adaptively adjusting density sampling based on plant morphology density, or by combining color or vegetation index to obtain dense sampling within the candidate region.

[0034] Taking the Otsu method as an example, the foreground depth threshold is set according to the distribution range of effective depth values ​​in soybean image data. Therefore, when generating depth constraint cue points, depth constraint cue points can be generated based on regions in the region of interest where the depth value is less than the foreground depth threshold.

[0035] In some embodiments, to generate depth constraint cue points, the electronic device can obtain the distribution range of effective depth values ​​by traversing the depth values ​​of each pixel in the associated depth image. Then, a foreground depth threshold is set based on this distribution range.

[0036] Specifically, when the distribution range is less than the range threshold, the foreground depth threshold is set based on the maximum depth value in the associated depth image; when the distribution range is greater than or equal to the range threshold, the depth value is linearly normalized to the grayscale range, and the foreground depth threshold is set based on the histogram of the grayscale range.

[0037] For example, for a depth submap, the distribution range of effective depth values ​​can be statistically analyzed. This distribution range can be determined by the maximum and minimum depth values, representing the distribution of all effective depth values. After obtaining the distribution range, it can be compared with a preset range threshold. If the distribution range is smaller than the range threshold, it can be determined that the depth distribution range is narrow. Therefore, the value near the maximum depth can be directly taken as the upper limit of the foreground depth, i.e., the foreground depth critical value.

[0038] When the distribution range is greater than or equal to the range threshold, it can be determined that the current depth range is wide enough, and the depth can be linearly normalized to the grayscale range. For example, assuming the minimum effective depth in the depth sub-image is Dmin and the maximum is Dmax, then for any depth value D(x, y), the normalized grayscale value G(x, y) can be calculated according to the following formula:

[0039] Based on Histogram calculation of Otsu threshold T otsu Then, it is mapped back to the true depth threshold, i.e., the foreground depth critical value D. thresh It can be calculated using the following formula:

[0040] in, This is an empirical margin constant, such as 50mm, to ensure that the plant edge is included within the foreground area.

[0041] It should be noted that the calculation of the foreground depth threshold (depth threshold) does not have to be limited to the Otsu method. It can also be determined by adaptive clustering (such as K-means), histogram peak detection, etc., which will not be shown one by one.

[0042] After setting the foreground depth threshold, the region of interest can be uniformly divided into grids according to a preset grid size, and the center pixel of each grid can be taken as a candidate cue point. Then, the associated depth value of the candidate cue point is queried, and the obtained associated depth value is compared with the foreground depth threshold. If the associated depth value is greater than zero and less than the foreground depth threshold, the candidate cue point corresponding to the associated depth value can be added to the cue set. The cue set is a set of points composed of depth-constrained cue points.

[0043] For example, after determining the foreground depth range, the electronic device can uniformly divide the rectangular area into grids according to a preset grid size, and take the center pixel of each grid as a candidate cue point. For each candidate cue point, its corresponding associated depth value is queried. If the associated depth value is greater than zero and less than the foreground depth threshold... If the desired depth value is found, the candidate cue point corresponding to that depth value is added to the cue set; otherwise, the candidate cue point corresponding to that depth value is discarded.

[0044] If the final set is empty, it means that the area lacks effective depth information or does not contain the target plant. In this case, the electronic device can prompt the user to reselect the area of ​​interest or reacquire images through the monitoring interface.

[0045] S103. Segment the soybean plant image according to the depth constraint cue points to obtain a single plant mask.

[0046] After generating depth constraint cue points, single soybean plant segmentation can be performed based on depth guidance. This involves segmenting the soybean plant image according to the depth constraint cue points to obtain a single plant mask. Electronic devices can then use image segmentation models to segment soybean plant images.

[0047] Image segmentation models are neural network models capable of generating high-quality object masks from input cues such as points, bounding boxes, and text descriptions. For example, general image segmentation models such as the Segment Anything Model (SAM) can be used. SAM can achieve zero-sample or few-sample segmentation using cue points or cue boxes, and can be applied to agricultural scenarios such as leaf segmentation, plot boundary extraction, and canopy structure analysis.

[0048] In order to segment individual plant masks, in some embodiments, when performing segmentation of soybean plant images according to depth constraint cue points to obtain individual plant masks, the soybean plant images can be input into the image segmentation model, and the depth constraint cue points can be mapped from the image coordinate system to the model's internal coordinate system using the coordinate transformation function provided by the image segmentation model to form the prediction input.

[0049] Based on the predicted input, an image segmentation model is used to perform image segmentation on the soybean plant image to obtain several candidate masks corresponding to the depth constraint cue points and the predicted cross-union ratios corresponding to the candidate masks. The candidate mask with the highest predicted cross-union ratio is selected as the preferred mask.

[0050] For example, when the cue set is not empty, the electronic device can send a complete color image to the SAM and, using the coordinate transformation function provided by the model, map the cue points from image coordinates to the model's internal coordinate system to form the prediction input. To improve segmentation performance and stability, the model can be called at once for all depth-constrained cue points, and the multi-mask output option can be enabled, so that each depth-constrained cue point corresponds to several candidate masks and their predicted Intersection over Union (IoU) values. For each depth-constrained cue point, the candidate mask with the highest predicted IoU value is selected as the preferred mask for that cue.

[0051] The preferred mask can be used for soybean growth monitoring. To obtain more accurate monitoring results, the preferred mask can be further filtered. In some embodiments, when segmenting the soybean plant image according to depth constraint cue points to obtain a single-plant mask, the non-zero depth values ​​within the preferred mask can be counted on the associated depth image, and the median depth can be calculated based on these non-zero depth values. Then, based on the median depth and the foreground depth threshold, non-target regions in the soybean plant image are identified. These non-target regions are mask regions with fewer than a pixel count threshold for effective depth, and / or mask regions with a median depth greater than the foreground depth threshold.

[0052] Next, non-target regions are removed from the soybean plant image, and pixel-level union operations are performed on the soybean plant image after removing non-target regions to obtain a preliminary foreground mask. Then, using the region of interest as a constraint, the preliminary foreground mask is restricted to a specified range to obtain a restricted mask. Finally, connected component analysis is performed on the restricted mask to calculate the connected region with the largest area, which is used as the single-plant mask.

[0053] For example, on the associated depth map, preferred masks can be filtered one by one. For each preferred mask, the non-zero depth values ​​within it are counted, and the median depth is calculated. If the number of effective depth pixels within the mask is too small, or the median depth is greater than the foreground depth threshold, then the preferred mask is considered to correspond to the background or a non-target area and needs to be removed. Then, a pixel-level union operation is performed on the remaining masks to obtain a preliminary foreground mask. Finally, a rectangular region is used as a hard constraint to restrict the mask to a user-specified range.

[0054] To remove scattered noise and multiple simultaneously existing connected regions, connected component analysis can be performed on the restricted mask to calculate the area of ​​each connected region. The connected region with the largest area is retained as the final single-plant soybean mask, i.e., the single-plant mask. This single-plant mask visually corresponds to the outline of an entire soybean plant and can be used for subsequent geometric and color feature calculations.

[0055] like Figure 3 As shown, the automatic selection mechanism for depth constraint cue points based on depth guidance can use user-specified or system-initialized bounding boxes as regions of interest (ROIs). Within the ROI, grid sampling points are generated, and these points are classified according to a depth threshold. Circles represent points with depth values ​​less than the foreground threshold, which will be used as positive samples to prompt the input image segmentation model; crosses represent points with depth values ​​greater than the foreground threshold (background) or invalid points, which will be automatically discarded. Therefore, the automatic selection mechanism for depth constraint cue points can effectively suppress interference from background walls and the ground in segmentation.

[0056] It should be noted that the selection of a single plant mask is not limited to using the maximum connected component rule; shape features, such as aspect ratio and compactness, can also be incorporated. For scenarios with multiple adjacent plants, watersheds or instance segmentation algorithms can be introduced onto the mask to achieve effective separation of closely distributed plants.

[0057] S104. Using the median depth within a single plant mask and the geometric relationship of the camera's field of view, the pixel height of the soybean plant in the image is converted into the actual height of the plant.

[0058] After obtaining the single-plant mask, the plant height can be estimated based on the single-plant mask. That is, by using the median depth within the single-plant mask and the geometric relationship of the camera's field of view, the pixel height of the soybean plant in the image can be converted into the actual height of the plant.

[0059] In some embodiments, when estimating plant height, the row coordinates of foreground pixels in a single plant mask can be statistically analyzed. The plant top position and the plant root position can then be extracted from the soybean plant image based on these row coordinates. The plant top position is the pixel position corresponding to the lowest row coordinate or the lowest quantile in the foreground pixel row coordinates. The plant root position is the pixel position corresponding to the lower boundary of the root growth reference in the soybean plant image. The plant pixel height is then calculated as the difference between the row coordinates corresponding to the plant top position and the row coordinates corresponding to the plant root position.

[0060] For example, after obtaining a single-plant mask, the row coordinates of the foreground pixels in the mask can be statistically analyzed, and the lowest row coordinate or a low quantile (such as the 0.5% quantile) can be taken as the top position of the plant to reduce the influence of occasional noise. Simultaneously, for the pot rim or soil baseline, the row coordinates corresponding to the lower boundary of a user-drawn rectangular area can represent the position of the plant roots; the difference between these two coordinates is the pixel height of the soybean plant in the soybean plant image. .

[0061] After calculating the plant pixel height, the camera's field of view geometry can be obtained, and the actual plant height can be calculated based on the camera's field of view geometry and the plant pixel height. The camera's field of view geometry can include the camera's vertical field of view angle and pixel vertical resolution. Therefore, when calculating the actual plant height, all non-zero depth values ​​can be extracted from the associated depth image using a single-plant mask. The camera representative distance is then calculated based on these non-zero depth values. This camera representative distance characterizes the distance between the entire soybean plant and the camera; correspondingly, the camera representative distance is the median of the non-zero depth values.

[0062] Based on the acquired vertical field of view and pixel vertical resolution of the camera, a vertical length coefficient is calculated according to the vertical field of view, pixel vertical resolution, and the distance represented by the camera. This vertical length coefficient represents the actual length represented by a unit pixel in the vertical direction of the soybean plant image.

[0063] For example, in a correlated depth image, all non-zero depth values ​​can be extracted using a single-plant mask, and the median value can be calculated as the representative distance from the entire plant to the camera. Combining this with the camera's vertical field of view and the image's vertical resolution, the actual length represented by a single pixel in the vertical direction of the image can then be converted using a vertical length coefficient, which can be calculated using the following formula:

[0064] Where S is the vertical length coefficient, in mm / pixel; The distance represented by the camera is the median of all non-zero depth values ​​in a single mask, expressed in millimeters (mm). This refers to the vertical field of view. This indicates the vertical resolution of the image.

[0065] After calculating the vertical length coefficient, a calibration correction coefficient can be set to calculate the actual plant height based on the calibration correction coefficient and the vertical length coefficient. The calibration correction coefficient is used to correct linear errors caused by lens distortion or depth sensor system deviations; it is calculated using the least squares method based on a known height plant image; the actual plant height is the product of the plant pixel height, the vertical length coefficient, and the calibration correction coefficient.

[0066] For example, regarding the plant pixel height measured in a soybean plant image... Uncalibrated physical height can be calculated based on the vertical length factor. That is, unmarked physical height It can be calculated using the following formula:

[0067] Combined with calibration correction factors λ Calculate the final estimated plant height, that is, calculate the actual plant height according to the following formula. :

[0068] Among them, calibration correction coefficient λ This is used to correct linear errors caused by lens distortion or depth sensor system deviations. Calibration correction coefficients. λ This can be obtained by selecting several plant images with known heights and analyzing them using the least squares method.

[0069] S105. Extract growth features from soybean plant images, and calculate the leaf yellowing rate and pod number based on the growth features.

[0070] After estimating the actual height of the plant, the electronic device can continue to determine the growth stage of the soybean plant under test. Before determining the growth stage, growth characteristics such as color and pods can be extracted. That is, growth characteristics are extracted from the soybean plant image, and the proportion of yellowing leaves and the number of pods are calculated based on the growth characteristics. Among them, the growth characteristics include leaf color and pod count.

[0071] In order to calculate the leaf yellowing rate, in some embodiments, a plant sub-image can be extracted from the soybean plant image within the area covered by the single-plant mask, and a color space conversion can be performed on the plant sub-image to convert it from the RGB color space to the HSV color space.

[0072] Next, transformation determination parameters are obtained, which are used to characterize the transformation of soybean leaves from green to yellow; the transformation determination parameters include hue range, saturation threshold, and brightness threshold. Then, a yellow pixel mask is constructed based on the transformation determination parameters, and the total number of pixels in a single plant mask and the number of pixels in the yellow mask are counted. Finally, the leaf yellowing ratio is obtained by calculating the ratio of the number of yellow mask pixels to the total number of pixels.

[0073] For example, within the area covered by a single plant mask, a corresponding colored plant sub-image can be extracted from a colored soybean plant image. This plant sub-image is then converted from the Red Green Blue (RGB) color space to the Hue Saturation Value (HSV) color space.

[0074] Based on the hue range of soybean leaves transitioning from green to yellow, and appropriate saturation and brightness thresholds, a yellow pixel mask is constructed. Then, by counting the total number of pixels within a single plant mask and the number of pixels belonging to the yellow mask, the ratio of the two is the leaf yellowing ratio.

[0075] For ease of explanation, the color status can be divided into several levels of leaf yellowing, such as "green", "beginning to yellow", and "obvious yellowing", based on the specific numerical value of the leaf yellowing ratio, and displayed in text form on the monitoring interface of the electronic device.

[0076] To calculate the number of soybean pods, in some embodiments, when calculating the number of pods based on growth characteristics, a local color image can be cropped from the soybean plant image using a single-plant mask. This local color image is then input into an object detection network to output candidate pod boxes and their confidence scores. The object detection network is a neural network used to extract objects of specific shapes and colors from a color image. For example, the object detection network could be a You Only Look Once (YOLO) network, which treats object detection as a single regression problem, directly mapping image pixels to bounding box coordinates and class probabilities.

[0077] After the object detection network outputs pod candidate boxes and their confidence scores, the center points of the pod candidate boxes can be obtained. Based on the center points and confidence scores, the target pods are then determined. Specifically, a target pod is a set of pixels covered by a pod candidate box whose center point is located within a single plant mask and whose confidence score is higher than a preset confidence threshold. The number of pods is then determined by counting the number of target pods.

[0078] For example, the growth characteristics corresponding to the number of pods are called pod features. Pod feature extraction can be achieved using an object detection network. A local color image is cropped near the bounding rectangle of a single plant mask and used as input to the object detection network. Through internal classification calculations, the object detection network can output several candidate pod boxes and their confidence scores.

[0079] To suppress the interference of background regions in soybean plant images on the pod feature extraction process, after the target detection network outputs pod candidate boxes and their confidence scores, only detection boxes whose center points fall within a single plant mask and whose confidence scores are higher than a set threshold are retained. The number of retained detection boxes is recorded as the number of pods. Simultaneously, the pod detection boxes and their confidence scores are overlaid on the resulting image for user verification.

[0080] It should be noted that the target detection network for pod detection in the above embodiments is not limited to using the YOLO series algorithms; other lightweight detection networks or instance segmentation networks can also be used. Furthermore, the leaf color yellowing analysis can incorporate the Normalized Difference Vegetation Index (NDVI) and color histogram statistics to enhance robustness to changes in light intensity.

[0081] S106. Generate growth stage levels based on the actual height of the plant, the proportion of yellowing leaves, and the number of pods.

[0082] After calculating the leaf yellowing ratio and the number of pods based on growth characteristics, the growth period can be determined by combining the actual plant height determined during the soybean plant height measurement process. That is, the growth period level is generated based on the actual plant height, leaf yellowing ratio, and number of pods.

[0083] Based on the agronomical classification of soybean growth processes, the growth process can be divided into several stages, including the vegetative growth stage, flowering stage, mid-reproductive stage, and maturity stage. Combining three characteristics—plant height, yellowing rate, and pod number—a multimodal growth stage determination rule is constructed.

[0084] To generate reproductive age levels, electronic devices can also determine reproductive age based on a decision tree-based, multimodal classifier. That is, for example... Figure 4 As shown, in some embodiments, the rules for determining the growth period may include, but are not limited to: when the proportion of yellowing leaves is higher than a preset threshold and a certain number of pods are detected, the plant can be considered to have entered a stage close to maturity; when pods are detected and the leaves are still predominantly green, and the plant height is close to the general height range of the variety, it is determined to be the mid-reproductive stage; when no pods are detected and the plant height is low, it is determined to be the vegetative growth stage; when no pods are detected and the plant height is high and remains green, it can be determined to be the flowering stage. The above rules can be adjusted according to different varieties and management conditions, or they can be replaced by a classification model based on the same set of features within this framework.

[0085] It should be noted that when sufficient labeled data can be obtained during soybean growth monitoring, the current decision tree can be expanded into a multimodal classifier such as a random forest, gradient boosting tree, or lightweight neural network to further improve the accuracy and adaptability of growth period determination.

[0086] In some embodiments, the method can also display the monitoring results in a graphical interface. For example, in offline dataset mode, the user specifies the directory containing the color image and depth map. The electronic device can automatically match the file names, grouping multiple frames of the same plant together and listing the plant number and corresponding frame number on the left side of the interface. After the user selects the target plant and a frame, the image display area presents the corresponding color image. After the user interactively selects a rectangular area on the image, depth-guided segmentation, plant height calculation, and color and pod feature extraction are automatically performed. The single plant mask, baseline, center line, and pod detection box are overlaid on the image, and the height, yellowing ratio, number of pods, and growth stage determination results are output in the text area.

[0087] By applying the technical solutions of the above embodiments, the soybean growth monitoring method based on depth image perception segmentation in the above embodiments can improve the accuracy of segmentation and geometric measurement, enhance the multimodal growth stage discrimination performance, and has high application value. Since the experimental data are mainly concentrated in a few growth stages, the following experimental results are also presented using these stages as examples.

[0088] To verify the effectiveness of the depth-guided segmentation process, a comparative experiment was constructed on the same set of RGB-D images of potted soybeans. First, the regions of individual soybean plants were manually annotated to generate ground truth masks. Then, rectangular regions of interest were manually selected near the plant's circumscribed region, serving as the unified input for the three segmentation methods. Evaluation metrics included Intersection over Union (IoU), Dice coefficient, Precision, and Recall, reflecting the degree of overlap between the segmentation results and the ground truth, overall consistency, and the number of missed and false detections, respectively. The IoU was calculated as follows:

[0089] The method for calculating the Dice coefficient is as follows:

[0090] Precision is calculated as follows:

[0091] The recall rate is calculated as follows:

[0092] Where TP represents the number of true positive pixels, FP represents the number of false positive pixels, and FN represents the number of false negative pixels.

[0093] In contrast, Method 1 is a traditional color or threshold segmentation method. Within a manually selected rectangular region, the image is converted to HSV or vegetation index space, and a fixed or empirical threshold is used to distinguish plants from the background. Morphological operations such as opening and closing operations are then used to remove isolated noise, resulting in single-plant candidate regions. Method 2 is the SegmentAnything segmentation method based solely on RGB images and bounding box cues. It uses the center and four corners of a bounding box as positive cue points on the color image and calls a general SAM model to generate a mask without introducing any depth information. Method 3 is the depth-guided SAM segmentation method proposed in this application. Within the same rectangular region, Otsu threshold segmentation is performed based on the depth histogram to estimate the foreground depth range. Positive cue points are generated by grid sampling within regions that meet depth constraints. The SAM model is called all at once to obtain candidate masks, and the median depth and connected component area within the masks are used for filtering to obtain the final single-plant mask.

[0094] Table 1 shows the statistical results of the average splitting evaluation metrics for the three methods on the same dataset: Table 1. Comparison of accuracy results of the three methods;

[0095] As can be seen, Method 1 (traditional color thresholding segmentation) achieved an average IoU of 0.276, a Dice coefficient of 0.419, a precision of 0.951, and a recall of 0.284 on the test image. Therefore, this method exhibits strong suppression of background regions and fewer false positives, resulting in high precision. However, it suffers from significant missed detections in slender branches and edge regions, leading to lower recall and IoU. Furthermore, the overall mask is too thin and cannot completely cover the plant outline.

[0096] Method 2 (SAM segmentation using only RGB and rectangular cues) achieved an average IoU of 0.034, a Dice coefficient of 0.062, a precision of 0.035, and a recall of 0.297 on the same data. However, because this method relies entirely on the generalization ability of SAM on RGB texture and shape features without being optimized for the current potted soybean scene, it is prone to mistaking background structures for foreground structures or only segmenting local branches and leaves under complex background and lighting conditions, resulting in significantly insufficient overall segmentation quality.

[0097] Method 3 (Depth-guided SAM segmentation method) achieved an average IoU of 0.821, a Dice coefficient of 0.899, a precision of 0.898, and a recall of 0.913 on the same data. Compared to Method 1, the IoU increased by more than three times, significantly improving recall while maintaining high precision. The single-plant mask can cover the outline of branches and leaves relatively completely. Compared to Method 2, there were orders of magnitude improvements in IoU, Dice, precision, and recall, demonstrating that the cue point generation and mask selection mechanism after introducing depth constraints can effectively suppress background interference, enabling the general SAM to obtain stable and reliable results in soybean single-plant segmentation tasks. Therefore, the soybean growth monitoring method based on depth image perception segmentation described in the above embodiments can maintain high segmentation accuracy and robustness under different growth stages and background conditions, providing a reliable mask basis for subsequent plant height estimation, color analysis, and pod detection.

[0098] To verify the accuracy of plant height measurement, several plants at different growth stages were selected in a potted soybean experiment. The plant height measured vertically with a measuring tape was used as the control value, while the plant height obtained by depth-guided segmentation and geometric calculation methods was used as the evaluation value. In the experiment, the measured height of individual soybean plants was roughly distributed in the range of about 5 to 38 cm, covering the typical plant morphology from the seedling stage to the reproductive stage.

[0099] The coefficient of determination (R²), mean absolute error (MAE), and root mean square error (RMSE) were used as statistical indicators to evaluate the performance of the method in measuring plant height.

[0100] The coefficient of determination (R²) can be calculated using the following formula:

[0101] The mean absolute error (MAE) can be calculated using the following formula:

[0102] The root mean square error (RMSE) can be calculated using the following formula:

[0103] in, Indicates the first i The actual value of each observation; Indicates the first i The predicted value of an observation, that is, the output value predicted by the model based on the input variables; This represents the average of the observed values; n This represents the total number of observations.

[0104] Statistical analysis of the sample after removing obvious outliers yielded a mean absolute error (MAE) of approximately 0.87 cm, a root mean square error (RMSE) of approximately 1.13 cm, a mean relative error of approximately 4.9%, and a mean deviation of approximately [missing value]. The value is 0.19 cm, and the coefficient of determination R² is approximately 0.983. Figure 5 A scatter plot comparing the measurement results, such as Figure 5 As shown, most of the sample points in the method are closely distributed around the ideal straight line y=x, with only slight deviations at a few tall plant samples. Figure 6 For the measurement error distribution histogram, such as Figure 6 As shown, the errors of the method are mainly concentrated in The error range is within 2cm to 2cm, and the overall value is slightly negative, indicating that while ensuring high linear consistency, this method has a slight tendency to underestimate the true plant height. The overall error level meets the application requirements for monitoring the growth of potted soybeans and determining the growth period.

[0105] Regarding the logic and accuracy of growth stage determination, the method uses three characteristics as inputs: plant height, leaf yellowing rate, and pod quantity. It then uses rule-based decision-making to determine the growth stage level. In the experiment, several potted soybeans in the vegetative growth stage, flowering stage, mid-reproductive stage, and near-maturity stage were manually labeled by technicians familiar with soybean cultivation based on field observations and agronomic stage classification standards, serving as a control. On the same plant, the system automatically performed single-plant segmentation, plant height calculation, color analysis, and pod detection, outputting the corresponding growth stage determination level.

[0106] On samples where pod detection is relatively stable—that is, when the target detection network can correctly mark multiple pod candidate boxes near a single plant mask—the judgment results of this application are highly consistent with manual annotation. For example, on plants with obvious pods detected and leaves still predominantly green, the determination of mid-reproductive stage is consistently given; on plants with extensive yellowing leaves and detected mature pods, the determination of near-maturity stage is consistent with manual judgment. For early plants that have not yet formed pods or have very few pods, the distinction is mainly based on plant height morphology thresholds. In the absence of significant organ features, the seedling stage and flowering stage can be distinguished by whether the plant height has been successfully shaped, providing a reasonable stage division.

[0107] It should be noted that the pod detection model used in the current experiment is YOLOv8, which still exhibits missed detections and low confidence levels in images with limited complex lighting or severe occlusion. Even so, the multi-feature decision rule constructed in this application demonstrates significant robustness: as long as the main pod structure can be identified within the segmentation mask, the key information of "whether pods appear" can be used to reliably transition the plant from the vegetative growth stage to the reproductive growth stage. Furthermore, this application does not limit the specific network structure for pod detection; YOLOv8 can be incrementally trained as needed or directly replaced with other high-precision object detection models. As the accuracy of the detection model improves, the stability of the entire growth stage discrimination module will improve synchronously without altering the overall system framework.

[0108] In terms of application effectiveness and ease of use, the method integrates various functional modules through a graphical interface, significantly improving the efficiency of phenotypic monitoring. For offline dataset mode: users only need to select the plant number and image frame to be analyzed, and circle the area of ​​a single plant on the image; the system can automatically complete depth-guided segmentation, plant height calculation, color and pod feature extraction, and growth stage determination. The interface will overlay a single plant mask, baseline, centerline, and pod candidate boxes, and simultaneously provide height values, yellowing rate, pod quantity, and textual descriptions of the growth stage, achieving one-click generation of richly illustrated analysis results.

[0109] For online camera mode: After connecting the electronic device to the RGB-D camera, the captured image can be displayed in real time. The operator can freeze the current frame at the appropriate time and delineate the area of ​​a single plant to obtain immediate information on plant height and growth stage on-site. This "what you see is what you get" measurement method avoids the tedious process of repeatedly bending over to measure and manually recording data, significantly reducing the labor intensity of field or greenhouse monitoring.

[0110] Based on the above experimental results and practical user experience, the proposed method not only meets the application requirements in terms of segmentation and geometric measurement accuracy, but also provides an intuitive and easy-to-use operating interface under conventional hardware conditions. It offers a technical means with promotional value for soybean pot experiments and growth monitoring of greenhouse breeding materials.

[0111] In some embodiments, as a specific implementation of the soybean growth monitoring method based on depth image perception segmentation described in the above embodiments, some embodiments of this application also provide a soybean growth monitoring system based on depth image perception segmentation, such as... Figure 7 As shown, the system includes: The data acquisition module is used to acquire soybean image data, which includes soybean plant images and associated depth images; the soybean plant images include regions of interest, which are image regions covering the soybean plant under test and the root growth baseline of the soybean plant under test; The image preprocessing and alignment module is used to generate depth constraint cue points based on regions in the region of interest whose depth values ​​are less than the foreground depth threshold, wherein the foreground depth threshold is set according to the distribution range of effective depth values ​​in the soybean image data. A depth-guided segmentation module is used to segment the soybean plant image according to the depth constraint prompts to obtain a single plant mask; The plant height calculation module is used to convert the plant pixel height in the soybean plant image into the actual plant height by using the median depth within the single plant mask and the geometric relationship of the camera field of view. The color and pod feature extraction module is used to extract growth features from the soybean plant image and calculate the leaf yellowing ratio and pod number based on the growth features; the growth features include leaf color and pod target. The growth period determination module is used to generate a growth period level based on the actual height of the plant, the proportion of yellowing leaves, and the number of pods.

[0112] By applying the technical solutions of the above embodiments, the soybean growth monitoring system based on depth image perception segmentation described in the above embodiments can acquire soybean plant images and associated depth images through the data acquisition module, and generate depth constraint cue points based on the depth values ​​in the region of interest and the foreground depth threshold by the image preprocessing and alignment module. The depth-guided segmentation module then segments individual plant masks from the soybean plant images according to the depth constraint cue points, so that the plant height calculation module can use the median depth within the individual plant mask and the geometric relationship of the camera's field of view to convert the plant pixel height into the actual plant height. The color and pod feature extraction module extracts growth features from the soybean plant images, and calculates the leaf yellowing ratio and pod number based on the growth features, so that the growth period determination module can determine the growth period level. The system can extract individual soybean regions based on RGB-D images using a depth-guided image segmentation model, and perform height measurement and growth period determination, achieving rapid and stable monitoring of soybean individual plant phenotypic indicators and improving the accuracy of soybean growth monitoring results.

[0113] It should be noted that other corresponding descriptions of the functional units involved in the soybean growth monitoring system based on depth image perception segmentation provided in the embodiments of this application can be found in the corresponding descriptions in the soybean growth monitoring method based on depth image perception segmentation provided in the above embodiments, and will not be repeated here.

[0114] This application also provides a computer device, specifically a personal computer, server, network device, etc. The computer device includes a bus, processor, memory, and communication interface, and may also include input / output interfaces and a display device. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores location information. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the various method embodiments.

[0115] Those skilled in the art will understand that the structure of the computer device described above is only a partial structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. A specific computer device may include more or fewer components, or combine certain components, or have different component arrangements.

[0116] In one embodiment, a computer-readable storage medium is also provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0117] In one embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0118] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0119] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.

[0120] Any references to memory, database, or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc.

[0121] Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can take many forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0122] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchain. The processors involved in the embodiments provided in this application may be, but are not limited to, general-purpose processors, graphics processors, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc.

[0123] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0124] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A soybean growth monitoring method based on depth image perception segmentation, characterized in that, The method includes: Acquire soybean image data, which includes soybean plant images and associated depth images; the soybean plant images include regions of interest, which are image regions covering the soybean plant under test and the root growth baseline of the soybean plant under test; Depth constraint cue points are generated based on regions in the region of interest whose depth values ​​are less than the foreground depth threshold. The foreground depth threshold is set according to the distribution range of effective depth values ​​in the soybean image data. The soybean plant image is segmented according to the depth constraint cue points to obtain a single plant mask; Using the median depth within the single-plant mask and the geometric relationship of the camera's field of view, the pixel height of the soybean plant in the image is converted into the actual height of the plant. Growth features are extracted from the soybean plant image, and the proportion of yellowing leaves and the number of pods are calculated based on the growth features; the growth features include leaf color and pod target. The growth stage level is generated based on the actual height of the plant, the proportion of yellowing leaves, and the number of pods.

2. The method according to claim 1, characterized in that, Acquire soybean image data, including: Acquire the original color image and original depth image of the soybean plant to be tested; The region of interest is delineated in the original color image to obtain an image of the soybean plant; Based on the region of interest, the original depth image is cropped into a depth sub-image, the depth sub-image including the depth values ​​associated with the pixels in the region of interest; The depth values ​​in the depth sub-map are converted to units to obtain the associated depth image.

3. The method according to claim 1, characterized in that, Generate depth constraint cue points based on regions in the region of interest whose depth values ​​are less than the foreground depth threshold, including: The distribution range of effective depth values ​​is obtained by traversing the depth values ​​of each pixel in the associated depth image; The foreground depth threshold is set according to the distribution range; wherein, when the distribution range is less than the range threshold, the foreground depth threshold is set according to the maximum depth value in the associated depth image; when the distribution range is greater than or equal to the range threshold, the foreground depth threshold is set by linearly normalizing the depth value to the grayscale range and based on the histogram of the grayscale range. The region of interest is divided into grids according to a preset grid size, and the center pixel of each grid is taken as a candidate cue point. Query the association depth value of the candidate suggestion points; If the association depth value is greater than zero and the association depth value is less than the foreground depth threshold, the candidate cue point corresponding to the association depth value is added to the cue set, and the cue set is a set of points composed of the depth constraint cue points.

4. The method according to claim 1, characterized in that, The soybean plant image is segmented according to the depth constraint cue points to obtain a single plant mask, including: Input the soybean plant image into the image segmentation model; By combining the coordinate transformation function provided by the image segmentation model, the depth constraint cue points are mapped from the image coordinate system to the model's internal coordinate system to form the prediction input; Based on the predicted input, the image segmentation model is used to perform image segmentation on the soybean plant image to obtain several candidate masks corresponding to the depth constraint cue points and the predicted intersection-union ratio corresponding to the candidate masks; The candidate mask with the highest predicted crossover / union ratio is selected as the preferred mask.

5. The method according to claim 4, characterized in that, The soybean plant image is segmented according to the depth constraint cue points to obtain a single plant mask, and the process further includes: On the associated depth image, count the non-zero depth values ​​inside the preferred mask; Calculate the median depth based on the statistically significant non-zero depth values; Based on the median depth and the foreground depth threshold, non-target regions in the soybean plant image are found. The non-target regions are masked regions where the number of effective depth pixels is less than the pixel number threshold, and / or, masked regions where the median depth is greater than the foreground depth threshold. A pixel-level union operation is performed on the soybean plant image after removing the non-target regions to obtain a preliminary foreground mask; Using the region of interest as a constraint, the initial foreground mask is limited to a specified range to obtain a restricted mask; Perform connected component analysis on the restricted mask to calculate the connected region with the largest area, which is then used as the single-plant mask.

6. The method according to claim 1, characterized in that, Using the median depth within the single-plant mask and the geometric relationship of the camera's field of view, the pixel height of the soybean plant in the image is converted to the actual plant height, including: The row coordinates of the foreground pixels in the single-plant mask are statistically analyzed; The plant top position is determined based on the foreground pixel row coordinates, where the plant top position is the pixel position corresponding to the minimum row coordinate or the low quantile of the foreground pixel row coordinates. Extract the location of the plant root from the soybean plant image. The location of the plant root is the pixel position corresponding to the lower boundary of the root growth reference in the soybean plant image. Calculate the plant pixel height, which is the difference between the row coordinates corresponding to the top position of the plant and the row coordinates corresponding to the root position of the plant. The camera field of view geometry is obtained, and the actual height of the plant is calculated based on the camera field of view geometry and the plant pixel height.

7. The method according to claim 6, characterized in that, The camera field of view geometry includes the camera's vertical field of view angle and pixel vertical resolution; Obtaining the camera's field-of-view geometry and calculating the actual height of the plant based on the camera's field-of-view geometry and the plant's pixel height includes: In the associated depth image, all non-zero depth values ​​are extracted using the single-plant mask; The camera representative distance is calculated based on the non-zero depth value, and the camera representative distance is used to characterize the distance between the entire soybean plant under test and the camera; the camera representative distance is the median of the non-zero depth values. Obtain the vertical field of view and pixel vertical resolution of the camera; The vertical length coefficient is calculated based on the vertical field of view, the pixel vertical resolution, and the camera representative distance. The vertical length coefficient is used to represent the actual length represented by a unit pixel in the vertical direction of the soybean plant image. A calibration correction factor is set, which is used to correct linear errors caused by lens distortion or depth sensor system deviation; the calibration correction factor is calculated based on a known height plant image using the least squares method; The actual height of the plant is calculated as the product of the plant pixel height, the vertical length coefficient, and the calibration correction coefficient.

8. The method according to claim 1, characterized in that, Extracting growth features from the soybean plant image, and calculating the leaf yellowing rate and pod number based on the growth features, includes: Within the area covered by the single-plant mask, extract plant sub-images from the soybean plant image; Perform a color space conversion on the plant sub-image to convert it from the RGB color space to the HSV color space; Obtain transformation determination parameters, which are used to characterize the transformation of soybean leaves from green to yellow; the transformation determination parameters include hue range, saturation threshold, and brightness threshold; A yellow pixel mask is constructed based on the transformation determination parameters, and the total number of pixels in a single mask and the number of pixels in the yellow mask are counted. The ratio of the number of yellow mask pixels to the total number of pixels is calculated to obtain the yellowing ratio of the leaves.

9. The method according to claim 1, characterized in that, Extracting growth features from the soybean plant image, and calculating the leaf yellowing rate and pod number based on the growth features, includes: Based on the single-plant mask, a local color image is extracted from the soybean plant image; The local color image is input into the target detection network to output pod candidate boxes and the confidence scores of the pod candidate boxes. Obtain the center point of the candidate pod frame; The pod target is determined based on the center point and the confidence level. The pod target is the set of pixels covered by the pod candidate box whose center point is located within the single plant mask and whose confidence level is higher than a preset confidence threshold. The number of the target pods is counted to obtain the total number of pods.

10. A soybean growth monitoring system based on depth image perception segmentation, characterized in that, The system includes: The data acquisition module is used to acquire soybean image data, which includes soybean plant images and associated depth images; the soybean plant images include regions of interest, which are image regions covering the soybean plant under test and the root growth baseline of the soybean plant under test; The image preprocessing and alignment module is used to generate depth constraint cue points based on regions in the region of interest whose depth values ​​are less than the foreground depth threshold, wherein the foreground depth threshold is set according to the distribution range of effective depth values ​​in the soybean image data. A depth-guided segmentation module is used to segment the soybean plant image according to the depth constraint prompts to obtain a single plant mask; The plant height calculation module is used to convert the plant pixel height in the soybean plant image into the actual plant height by using the median depth within the single plant mask and the geometric relationship of the camera field of view. The color and pod feature extraction module is used to extract growth features from the soybean plant image and calculate the leaf yellowing ratio and pod number based on the growth features; the growth features include leaf color and pod target. The growth period determination module is used to generate a growth period level based on the actual height of the plant, the proportion of yellowing leaves, and the number of pods.