Plant three-dimensional Gaussian phenotype extraction method based on improved parameter adaptive rapid pose optimization algorithm

Through the global adaptive fast pose optimization algorithm (GAPose-GS) and Gaussian Splating, the pose estimation error and phenotypic parameter extraction accuracy of plant three-dimensional reconstruction in agricultural scenarios are solved, and efficient and stable 3D modeling and phenotypic analysis are achieved.

CN120580355APending Publication Date: 2025-09-02NORTHWEST A & F UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510673735.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The three-dimensional reconstruction of plants in the prior art has problems such as accumulation of pose estimation errors, low reconstruction efficiency, and insufficient phenotypic parameter extraction accuracy in the prior art, which is difficult to meet the needs of high-throughput phenotypic analysis.

Method used

The global adaptive fast pose optimization algorithm (GAPose-GS) was used and combined with the Gaussian Splatting method, and the global pose optimization was fused with Gaussian characterization to generate dense three-dimensional point clouds and extract plant phenotypic features.

Benefits of technology

Effectively eliminate pose estimation errors in agricultural scenarios, improve the quality and efficiency of point cloud reconstruction, and provide high-precision phenotypic parameter support, suitable for three-dimensional modeling and phenotypic analysis of agricultural crops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580355A_ABST
    Figure CN120580355A_ABST
Patent Text Reader

Abstract

The invention discloses a plant three-dimensional Gaussian phenotype extraction method based on an improved parameter adaptive fast pose optimization algorithm, and the method comprises the steps: obtaining a multi-view image sequence of a plant through a camera system, estimating the pose of a camera through a global adaptive fast pose optimization algorithm, and generating a sparse three-dimensional point cloud of a scene, pose estimation errors in an agricultural scene can be effectively eliminated, the problem of error accumulation in a traditional method is avoided, and high-precision camera pose estimation is guaranteed; secondly, taking the camera pose and the sparse three-dimensional point cloud as input, converting an object in a scene into a plurality of three-dimensional Gaussian points, performing three-dimensional reconstruction based on a Gaussian representation method, and generating a dense three-dimensional point cloud, thereby solving common shielding and dynamic change problems in an agricultural scene, and improving the quality and efficiency of point cloud reconstruction; and finally, based on the obtained dense three-dimensional point cloud, phenotypic characteristics of the plants are extracted, and more accurate data support is provided for crop phenotypic analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of plant phenotyping and computer vision, and particularly relates to a method for extracting three-dimensional Gaussian phenotypes of plants based on an improved parameter-adaptive fast pose optimization algorithm, which is mainly used in three-dimensional modeling and phenotypic analysis of agricultural crops. Background Art

[0002] Three-dimensional reconstruction technology has been maturely applied in fields such as industrial inspection and autonomous driving. However, the complexity of agricultural scenes (such as cross-occlusion of plant organs and unstructured canopy morphology) poses unique challenges to traditional algorithms. Although multi-view image-based methods such as NeRF and 3D Gaussian Splatting have reduced hardware costs, their core reliance on COLMAP for camera pose estimation has significant flaws: COLMAP uses an incremental structure-from-motion (SFM) strategy. In agricultural scenes, due to problems such as uneven lighting and dynamic leaf interference, pose estimation requires repeated iterations, and the PSNR value is less than 25dB. More seriously, the flexible deformation characteristics of plants can cause temporal pose jitter, further reducing the extraction accuracy of phenotypic parameters (such as petiole angle), making it difficult to meet the needs of high-throughput phenotyping analysis.

[0003] In fields such as autonomous driving and urban modeling, global pose optimization technology has achieved rapid reconstruction of large scenes (36k images takes 5.5 hours) by jointly optimizing camera parameters through multi-viewpoint geometric constraints. However, the unique characteristics of agricultural scenes make the transplantation of such technologies difficult: 1) the diversity of plant morphology (such as corn tillers and wheat ear layers) undermines the geometric consistency of traditional feature matching; 2) dynamic growth characteristics require algorithms with temporal smoothness, but existing methods do not incorporate plant physiological priors. Therefore, adapting the global optimization framework to agricultural needs has become the key to breaking through the bottleneck of complex plant 3D reconstruction. Summary of the Invention

[0004] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide an efficient, stable, and dynamically adaptive three-dimensional Gaussian phenotype extraction method for plants. Based on the three-dimensional reconstruction algorithm of Gaussian Splatting and global adaptive fast pose optimization (Global Adaptive Pose Optimization with Gaussian Splatter Algorithm, GAPose-GS), by fusing global pose optimization with Gaussian representation, it solves the problems of camera pose estimation error accumulation, low reconstruction efficiency, and insufficient phenotypic parameter extraction accuracy in complex agricultural scenes, providing a high-precision, low-cost universal technical solution for crop precision breeding and phenotypic genomics research.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A method for extracting three-dimensional Gaussian phenotypes of plants based on an improved parameter-adaptive fast pose optimization algorithm comprises:

[0007] Step 1: Use the camera system to acquire a multi-view image sequence of the plant, use the global adaptive fast pose optimization algorithm to estimate the camera pose, and then generate a sparse 3D point cloud of the scene;

[0008] Step 2: Using the camera pose and sparse 3D point cloud as input, convert the objects in the scene into multiple 3D Gaussian points, perform 3D reconstruction based on the Gaussian representation method, and generate a dense 3D point cloud;

[0009] Step 3: Extract the phenotypic characteristics of the plant based on the obtained dense three-dimensional point cloud.

[0010] In one embodiment, the camera system uses a stepper motor to drive a circular track, the camera is installed on an adjustable shooting rod connected to the circular track, the plant is placed on a rack in the center of the circular track, and video is captured during rotation. The shooting angles include overhead, horizontal, and overhead shots. During the data acquisition process, the plant height, canopy width, and petiole angle are manually measured, and the obtained video is decomposed into frames to obtain the multi-view image sequence.

[0011] In one embodiment, the camera pose is estimated using a global adaptive fast pose optimization algorithm, and the implementation method is as follows:

[0012] Extract local features from the input image;

[0013] Match input images pairwise using descriptor matching;

[0014] A global optimization strategy is used to estimate the relative poses of all cameras at once and obtain a sparse 3D point cloud of the scene.

[0015] In one embodiment, in step 2, first, the sparse three-dimensional point cloud is converted into a dynamically optimized three-dimensional Gaussian ellipsoid as a three-dimensional Gaussian point, each Gaussian ellipsoid is defined by a mean vector μ, a covariance matrix ∑, an opacity α, and spherical harmonic coefficients, wherein the covariance matrix ∑ is decomposed into:

[0016] ∑=SS T R T

[0017] The initial scaling radius is calculated based on the K-nearest neighbor algorithm, so that the Gaussian ellipsoid can cover the local geometric structure. The color properties are modeled by spherical harmonics, and the order of the spherical harmonics determines the ability to express the view-dependent illumination changes.

[0018] Then, the three-dimensional Gaussian ellipsoid is projected onto a two-dimensional image space ellipse for rendering to generate an image. Given the viewing transformation W, the covariance matrix ∑′ in the camera coordinate is as follows:

[0019] ∑′=JW∑W T J T

[0020] where J is the Jacobian matrix of the affine approximation of the projective transformation.

[0021] In one embodiment, step 2 further includes:

[0022] Differentiable optimization is performed on the 3D reconstruction, and the position, shape and opacity of the 3D Gaussian points are optimized by back propagation.

[0023] In one embodiment, the differentiable optimization is a continuous iterative process based on rendering, in which the generated image is compared with the real data in the training view. During the backpropagation process, the gradient information optimizes the Gaussian parameters and drives the dynamic density adjustment mechanism. In this process, for areas where the gradient amplitude is greater than a threshold, the Gaussian density is increased through cloning or splitting operations to enhance details; and Gaussians with transparency below a threshold are pruned to reduce computational overhead.

[0024] In one embodiment, in step 3, the reconstructed dense 3D point cloud is preprocessed as follows:

[0025] First, straight-through filtering is used to remove invalid background point clouds;

[0026] Secondly, SOR filtering is used to remove noise.

[0027] In one embodiment, the straight-through filtering first specifies a dimension and a range of values ​​for that dimension, then iterates over each point in the point cloud, checking whether the value of that point in the specified dimension is within the range. Points that are not within the range are deleted. After the traversal is completed, the remaining points constitute the filtered point cloud.

[0028] The SOR filter sets a density threshold. When the point cloud density in a certain area is lower than the threshold, the point cloud is considered invalid.

[0029] In one embodiment, the step 3, extracting the phenotypic characteristics of the plant, is implemented as follows:

[0030] The AlphaShape algorithm was used to obtain the edge contours of the plant's vertical direction and canopy from the 3D reconstructed model. Based on the extracted contours, in a rectangular coordinate system, the direction from the xz plane and the direction perpendicular to the horizontal ground were used as the y-axis. By obtaining the lowest and highest points of the central axis, the vertical distance between the two planes, i.e., the plant height, was calculated based on the lowest and highest points of the plant height.

[0031] The canopy width is expressed by the average value of the maximum canopy width in the two horizontal directions. The vertical distance of the farthest horizontal point cloud of the canopy projected onto the ground is defined as the maximum length of the canopy. The corresponding longitudinal distance is defined as the maximum width of the canopy. The average of the sum of the maximum length and maximum width of the canopy is the canopy width.

[0032] The stem-leaf angle is calculated by the angle between two spatial vectors: the leaf's growth direction at birth and the local stem growth direction. Specifically, any point A on the plant's main stem is connected to the petiole and any endpoint B above point A to form a line segment. Vertex A is connected to endpoint C on the petiole to form a segment. The petiole angle is calculated using the coordinates of points A, B, and C and the lengths of line segments AB and AC.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] First, based on the global pose optimization technology, the pose estimation error in agricultural scenes can be effectively eliminated, the error accumulation problem in traditional methods can be avoided, and high-precision camera pose estimation can be ensured.

[0035] Secondly, three-dimensional reconstruction using Gaussian characterization technology solves the common occlusion and dynamic change problems in agricultural scenes, and improves the quality and efficiency of point cloud reconstruction.

[0036] Finally, phenotypic feature extraction is performed based on high-precision point cloud files, providing more accurate data support for crop phenotypic analysis.

[0037] Overall, the present invention provides a high-precision, low-cost technical solution suitable for high-throughput phenotyping in agriculture, which is widely applicable to the fields of three-dimensional modeling, precision breeding and phenotypic research of agricultural crops. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 The figure shows the overall process flow of the present invention. (a) shows the multi-view image acquisition process. (b) shows the 3D reconstruction of a pepper plant based on GAPose-GS, generating the final point cloud data. (c) shows the point cloud data preprocessing process. (d) shows the point cloud-based phenotypic extraction method. (e) shows the pepper phenotypic data obtained using the above method.

[0039] Figure 2 It is the pepper cultivation environment and multi-view data acquisition platform in the embodiment of the present invention.

[0040] Figure 3 These are partial images captured from three viewing angles in an embodiment of the present invention.

[0041] Figure 4Schematic diagram of the GAPose-GS algorithm.

[0042] Figure 5 Schematic diagram of the AlphaShape method principle and contour extraction results.

[0043] Figure 6 Schematic diagram of phenotypic extraction in an embodiment of the present invention, wherein (a) is plant height, (b) is canopy width, and (c) is petiole angle.

[0044] Figure 7 Comparison of NeRF, Instant-NGP, and GAPose-GS reconstruction and rendering results for the same pepper plant in an embodiment of the present invention. The blue dashed box shows the Instant-NGP reconstruction result, the green dashed box shows the GAPose-GS reconstruction result, and the pink dashed box shows the NeRF reconstruction result.

[0045] Figure 8 The results of the straight-through filtering in the embodiment of the present invention are shown in FIG. (a) is the original point cloud data generated; (b) is the result after the straight-through filtering.

[0046] Figure 9 This is the SOR filtering result in the embodiment of the present invention. Qian Zhongshu's green point cloud represents the filtered plant point cloud. The red point cloud represents the removed outliers.

[0047] Figure 10 1 is a box plot of the measured plant height values ​​and the correlation between the measured values ​​and the calculated values ​​in an embodiment of the present invention.

[0048] Figure 11 1 is a box plot of the measured values ​​of canopy width and the correlation between the measured values ​​and the calculated values ​​in an embodiment of the present invention.

[0049] Figure 12 1 is a box plot of the measured values ​​of the petiole angle and the correlation between the measured values ​​and the calculated values ​​in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The embodiments of the present invention are described in detail below with reference to the accompanying drawings and examples.

[0051] The present invention is a method for extracting plant three-dimensional Gaussian phenotypes based on an improved parameter adaptive fast pose optimization algorithm. In the embodiment, pepper plants are used as an example for illustration. The overall flow chart is as follows: Figure 1 shown.

[0052] First, a camera system is used to capture multi-view video data. The camera pose is estimated using an improved parameter-adaptive fast pose optimization algorithm, resulting in a preliminary sparse 3D point cloud. Global pose optimization techniques are then used to jointly optimize the geometric constraints in the multi-view images, accurately estimating the camera pose for each frame. This step effectively addresses pose estimation errors in agricultural scenes caused by factors such as uneven lighting and leaf dynamics, avoiding the inefficiencies and error accumulation associated with repeated iterations.

[0053] Next, the optimized camera pose and the initial sparse 3D point cloud are used as input for 3D reconstruction using the Gaussian characterization method. This technique effectively fuses the multi-viewpoint point cloud information extracted from the image to produce an accurate 3D model of the plant.

[0054] Finally, the phenotypic characteristics of the plant are extracted based on the resulting point cloud file. Various phenotypic characteristics, such as petiole angle, stem length, and canopy area, are calculated from the 3D reconstructed point cloud. These characteristics can accurately reflect the growth status of the plant and provide reliable data support for high-throughput phenotyping analysis of agricultural crops, precision breeding, and plant growth analysis. Therefore, the method of the present invention can be widely applied to fields such as high-throughput phenotyping of crops, automated agricultural monitoring, intelligent agricultural machinery, and plant growth analysis, providing a high-precision, low-cost solution.

[0055] Specifically, the experiment of the embodiment of the present invention was carried out in August 2024 at the Key Laboratory of Agricultural Internet of Things, Ministry of Agriculture and Rural Affairs, Northwest A&F University, Xianyang City, Shaanxi Province (34°07'39"N, 107°59'50"E, 648m above sea level), and the experimental material was pepper seedlings (Latianxia No. 18). During the experiment, Pindstrup substrate (Denmark) was used for cultivation, with an organic matter mass fraction of more than 50%, a substrate porosity of about 60%, and a pH of 6-7. The test sample germinated and grew in a 72-hole seed tray. When it grew to 2 leaves and 1 heart, it was transplanted into a seedling cup with a diameter of 20 cm and a height of 16 cm, and cultured in a CO2 artificial climate chamber (RGL-P500D-CO2, Darth Carter, China). Figure 2 As shown, the incubator is set to have a daytime of 12 hours and a temperature of 25°C, a nighttime of 12 hours and a temperature of 20°C. The humidity is 50% and the CO2 concentration is 400 μmol·mol -1 .

[0056] Depending on whether the plant occupies a fixed position during measurement and the overall design of the device, plant multi-view image acquisition methods can be divided into two modes: "plant-to-camera" and "camera-to-plant." In the "plant-to-camera" mode, the target plant is placed on a rotating turntable and rotated, with the camera position fixed, to obtain multi-view images of the plant. In the "camera-to-plant" mode, the target plant is maintained in a static position, and multi-view images are acquired by rotating one or more cameras around the plant. However, tall plants or plants with flexible organs tend to wobble when rotating on the turntable, resulting in poorly reconstructed point clouds, such as blurred leaf edges and thicker stems. Therefore, under the same image acquisition environment, the "camera-to-plant" mode achieves higher accuracy and robustness in 3D reconstruction than the "plant-to-camera" mode. Therefore, the data acquisition method of the present invention adopts the "camera-to-plant" mode to construct a 3D data acquisition platform for pepper plants, namely the camera system of the present invention.

[0057] To ensure uniformity in the acquisition position and image overlap, a stepper motor drives a circular track. The camera is mounted on an adjustable shooting rod connected to the track. The pepper plant is placed on a shelf in the center of the track, and video footage is captured during rotation. Because different rotation speeds and the number of cameras significantly affect reconstruction quality, the present invention incorporates three adjustable camera mounts on the adjustable shooting rod, allowing one to three cameras to capture images simultaneously. The camera position can be adjusted to ensure that the pepper plant is within the camera's imaging range and captures three different angles: overhead, horizontal, and overhead. By sending pulse signals, the stepper motor speed can be controlled to customize the rotation speed. This embodiment of the present invention customizes the rotation speed in three gears.

[0058] Specifically, during the data acquisition process, the camera positions of the pepper plants at three viewing angles, namely horizontal, vertical, and vertical, are used, and the camera bracket rotates counterclockwise along the circular track. In this invention, the circular video data of the pepper in its early flowering period is obtained using the acquisition platform. The partial images of the pepper plants at three viewing angles after 2 frames per second are shown as follows: Figure 3 shown.

[0059] In order to obtain the phenotypic traits of pepper plants and verify the effectiveness and accuracy of the reconstruction algorithm, the plant height and canopy width of pepper plants in the early flowering stage were manually measured using a ruler during the video data acquisition process. The petiole angle of the pepper plants was manually measured using a double-arm protractor. To reduce measurement errors, the same pepper plant was measured from four different angles in the horizontal direction, and the average value was taken.

[0060] The three-dimensional reconstruction process of the present invention can be specifically described as follows:

[0061] In recent years, radiation field technology has revolutionized the field of new view synthesis of scenes based on multi-view images or videos. This type of method can achieve rendering effects close to real scenes through implicit neural representations (such as NeRF), but its core defect is that the realization of high-quality output depends on the computationally intensive neural network training process, and existing acceleration solutions (such as Instant-NGP) often require efficiency improvements at the expense of reduced visual quality. For example, although the hash grid-based optimization method can shorten the training time, it is limited by the design of the structured grid, and it is difficult to effectively handle complex airspaces and lead to loss of details. In order to solve these problems, the present invention adopts a global adaptive pose optimization Gaussian splashing algorithm (GAPose-GS). As Figure 4 As shown in the figure, GAPose-GS is a 3D scene representation method based on 3D Gaussian distribution. By converting objects in the scene into multiple 3D Gaussian points, each containing information such as position, covariance matrix, and opacity, it can accurately express the geometric shape and lighting characteristics of complex scenes. GAPose-GS first obtains the precise camera pose through a global adaptive pose optimization algorithm, then generates a sparse 3D point cloud using SfM (Structure From Motion) technology. Then, based on this sparse 3D point cloud, it initializes a set of 3D Gaussian points and optimizes their position, shape, and opacity through backpropagation.

[0062] To generate dense 3D point clouds, GAPose-GS employs an adaptive density control strategy, removing unimportant Gaussian points after each backpropagation step and splitting or cloning Gaussian points as needed to ensure the accuracy and detail of scene reconstruction. Furthermore, GAPose-GS combines texture-aware densification and geometry-aware splitting strategies to generate a denser distribution of Gaussian points in texture-rich areas while maintaining sparsity in low-texture areas, resulting in high-quality, lighting-realistic dense point clouds.

[0063] The present invention uses GAPose-GS to extract training data from multi-view videos, and needs to segment the collected videos through FFmpeg to form a multi-view image sequence. Fast Forward Moving Picture Experts Group (FFmpeg) is an open source multimedia framework that supports functions such as recording, transcoding, and streaming media transmission. Its core codec library libavcodec provides efficient audio and video processing capabilities. In 3D Gaussian Splatting, FFmpeg is mainly used to convert panoramic videos that cannot be directly processed by 3D Gaussian Splatting into single-frame image sequences, providing standardized input data for subsequent SFM algorithms, thereby supporting high-precision 3D scene reconstruction. Therefore, the present invention uses FFmpeg to decompose the input video into frames, and the setting is 2 frames per second.

[0064] After acquiring a multi-view image sequence, GLOMAP, combined with structure from motion (SfM) and multi-view stereo (MVS) techniques, can reconstruct a sparse 3D point cloud of the scene and the camera pose from either an unordered or ordered image set. Unlike traditional methods, GLOMAP estimates all camera poses at once, thus avoiding the error accumulation and computational redundancy inherent in the iterative estimation process. First, GLOMAP extracts local features (such as SIFT features) from the input images. These features are typically easily identifiable and discriminative points in the image, providing the basis for subsequent feature matching. Next, the input images are matched pairwise using descriptor matching. These matched point pairs form the basis for subsequent estimation of camera poses and 3D structure. After obtaining sufficient feature matching pairs, GLOMAP employs a global optimization strategy to estimate the relative poses of all cameras at once. Unlike traditional step-by-step pose estimation methods, GLOMAP achieves simultaneous optimization of camera poses through global constraints, typically using essential or fundamental matrices to calculate rotations and translations between cameras, thereby obtaining the accurate poses of all cameras and the sparse 3D points of the scene.

[0065] After obtaining the camera parameters and the initial sparse 3D point cloud, it is necessary to initialize the 3D Gaussian distribution and convert the sparse point cloud generated by GLOMAP into a 3D Gaussian ellipsoid that can be dynamically optimized. Each Gaussian ellipsoid is defined by the mean vector μ (position), the covariance matrix ∑ (shape), the opacity α, and the spherical harmonic coefficients (color). The covariance matrix ∑ is decomposed into the following through the rotation matrix R (quaternion representation) and the scaling matrix S (scaling coefficients for each axis):

[0066] ∑=RSS T R T

[0067] This decomposition ensures the semi-positive definiteness of the covariance matrix while allowing for optimization of the rotation and scaling parameters via gradient descent. The initial scaling radius is typically computed using a K-nearest neighbor algorithm, such that a Gaussian ellipsoid covers the local geometry. Color properties are modeled using spherical harmonics (SH), whose order (e.g., 3rd order corresponds to 48 coefficients) determines the ability to express view-dependent illumination variations.

[0068] After that, the three-dimensional Gaussian ellipsoid is projected into the two-dimensional image space ellipse for rendering to generate the image. Given the viewing transformation W, the covariance matrix ∑′ in the camera coordinate is as follows:

[0069] ∑′=JW∑W T J T

[0070] where J is the Jacobian matrix of the affine approximation of the projective transformation.

[0071] The present invention optimizes the position, shape and opacity of three-dimensional Gaussian points through differentiable optimization back propagation. Differentiable optimization is a continuous iterative process based on rendering, and the generated image is compared with the real data in the training view. During the back propagation process, the gradient information not only optimizes the Gaussian parameters, but also drives the dynamic density adjustment mechanism. For areas where the gradient amplitude is greater than the threshold (such as the edge of the object or the area with complex texture), the Gaussian density is increased by cloning or splitting operations to enhance the details; Gaussians with transparency below the threshold (such as α < 0.0002) are pruned to reduce computational overhead. This adaptive density control enables 3D Gaussian Splatting to achieve efficient rendering while maintaining high fidelity.

[0072] The point cloud data generated by the GAPose-GS algorithm contains redundant background points and a certain amount of noise, which has a certain impact on the interpretation of the pepper plant's 3D structure and phenotypic traits. Therefore, the present invention preprocesses the reconstructed dense 3D point cloud to preserve the plant's 3D spatial range and reduce background noise. This involves the following two steps:

[0073] First, considering the presence of redundant background points in the original point cloud, we use straight-through filtering to remove invalid background points. This filtering process first specifies a dimension and a range for that dimension. It then iterates over each point in the point cloud, checking whether its value for the specified dimension falls within the range. Points outside the range are deleted. After the traversal is complete, the remaining points constitute the filtered point cloud. This method uses straight-through filtering to constrain point cloud coordinates, as shown in the following equation.

[0074]

[0075] Where (X d ,X u ), (Y d ,Yu) and (Z d ,Z u ) represents a finite range of three-dimensional Cartesian coordinates.

[0076] Due to environmental interference during data collection, the generated point cloud still contains many close-in noise points and outliers. To address the noise and outlier issues in the pepper plant point cloud, this paper uses the Statistical Outlier Removal Filter (SOR) algorithm to remove noise, thereby improving the accuracy of phenotyping and the speed of feature recognition.

[0077] SOR filtering is based on statistical methods and removes outliers by analyzing the local density of the point cloud. Its characteristic is that the neighborhood distance distribution of normal points conforms to the Gaussian distribution, while the distance of outliers deviates significantly from this distribution. Based on this characteristic of outliers, a density threshold can be set. When the point cloud density in a certain area is lower than this threshold, the point cloud is considered invalid. The core steps are as follows: For each point p in the point cloud, i , calculate the average distance between it and the nearest k points, and calculate the average Euclidean distance from the point to the neighboring points as follows:

[0078]

[0079] Here, ||·|| represents the Euclidean distance, and k is the preset number of neighborhood points.

[0080] Calculate the average distance μ and its standard deviation σ of all statistical points as follows:

[0081]

[0082] Where N is the total number of point clouds. Set the outlier discrimination threshold τ=μ+ασ, α is the adjustment factor. If a point satisfies μ i >τ, it is considered as an outlier and removed.

[0083] The pepper plant point cloud is extracted through straight-through filtering and SOR filtering to make the structure clearer and more complete, thereby improving the efficiency and accuracy of subsequent pepper plant phenotype extraction.

[0084] Based on the above-mentioned straight-through filtering and statistical filtering preprocessing, the present invention removes background noise and redundant data, and obtains point cloud data of a single pepper plant. In order to extract accurate plant contours, the influence of noise points on the extraction of phenotypic data is reduced. The present invention uses the AlphaShape algorithm to obtain the edge contours of the vertical direction and canopy of the pepper from the three-dimensional model. The AlphaShape algorithm is a point cloud boundary extraction method based on computational geometry, which extracts contour lines from point cloud data by dynamically adjusting the radius of a circle or sphere. The principle diagram is shown in FIG. Figure 5 As shown in (a), Figure 5 (b) shows the boundary extraction of triangular point cloud using Alphashape.

[0085] The AlphaShape algorithm is used to extract the contours of the pepper plant point cloud in the xy plane and xz plane. Figure 5 As shown in (c), the contour extraction results of the xy plane and xz plane are shown respectively. The red circle is the contour circle. From the extraction results, it can be seen that AlphaShape can extract the contour of the pepper very well.

[0086] Plant height is one of the most basic indicators in plant morphology research, and it is also one of the important traits that affect pepper plant type and yield. And the initial flowering period is a critical period for peppers. The change in plant height at this stage can reflect the growth status of the plant, thereby helping to predict the final yield. The plant height of peppers in the initial flowering period is defined as the distance measured from the base of the plant to the tip of the top leaf. According to the contour extracted by the AlphaShape algorithm, in the rectangular coordinate system, the present invention selects the direction starting from the xz plane and the direction perpendicular to the horizontal ground as the y-axis. By obtaining the lowest and highest points of the central axis, according to the lowest and highest points of the pepper plant height, that is, the pepper plant height H, the vertical distance between the two planes is calculated, such as Figure 6 As shown in (a).

[0087] Canopy width not only reflects the growth status and spatial distribution of the plant, but is also closely related to photosynthetic efficiency, nutrient absorption capacity and yield potential. The measurement of canopy width is particularly important during the initial flowering period. During this stage, the canopy structure of the pepper expands rapidly, and the number and area of ​​leaves increase significantly, providing sufficient photosynthetic products for flowering and fruit development. Therefore, changes in canopy width directly reflect the health status and growth potential of the plant. Based on the AlphaShape algorithm, the average value of the maximum canopy width of the pepper canopy in the two horizontal directions is used to represent the canopy width. The vertical distance from the farthest horizontal point cloud of the canopy to the ground is defined as the maximum length of the canopy, and the corresponding longitudinal distance is defined as the maximum width of the canopy. The average of the sum of the maximum length and maximum width of the canopy is the canopy width, such as Figure 6 As shown in (b).

[0088] The petiole angle is a key metric in pepper breeding. Selecting varieties with ideal petiole angles can lead to new varieties with higher yields and enhanced disease and pest resistance. Measuring the petiole angle of peppers can optimize planting density and row spacing, improving farmland utilization efficiency and output returns. Based on the petiole angle definition and manual measurement methods, the stem-leaf angle can be calculated by measuring the angle between two spatial vectors: the leaf's growth direction at its origin and the local stem growth direction. Figure 6 Figure (c) shows the result of segmenting the pepper stem and petiole skeletons using the Density-Based Spatial Clustering of Point Clouds with Geometric Feature Fusion (DBSCAN) algorithm, followed by contour extraction using the AlphaShape algorithm. A line segment is formed by connecting any point A on the pepper plant's main stem to the petiole and any endpoint B above point A. Vertex A is connected to endpoint C on the petiole to form a segment. The petiole angle θ is then calculated using the coordinates of points A, B, and C and the lengths of line segments AB and AC.

[0089] To verify the accuracy and feasibility of the method of the present invention in pepper phenotyping research, the performance of the method of the present invention was compared with NeRF, Instant-NGP and 3D Gaussian Splatting methods, and the correlation analysis between the calculated values ​​and the measured values ​​was performed.

[0090] The present invention uses training time and Peak Signal-to-Noise Ratio (PSNR) as evaluation indicators in the evaluation of the 3D reconstruction method.

[0091] This experiment uses an NVIDIA RTX4060 GPU and records the total time from the start of training to the convergence of the loss function for each model under the PyTorch framework. It also compares the training efficiency of GAPose-GS with 3D Gaussian Splatting, NeRF, and Instant-NGP.

[0092] Peak signal-to-noise ratio (PSNR) is widely considered an important quantitative indicator of image quality. It is often used to evaluate the difference between compressed and denoised images and their original counterparts. However, its application has now expanded to emerging fields such as computer vision. PSNR is the mean squared difference between the pixel values ​​of a reference image and the image being tested. PSNR is defined based on the mean squared error (MSE). Given an original image X of size m*n and a noisy image Y after adding noise to it, the MSE can be defined as:

[0093]

[0094] The calculation method of PSNR is:

[0095]

[0096] n is the number of bits per pixel, typically 8, meaning the number of pixel grayscales is 256. PSNR is measured in dB, with higher values ​​indicating lower distortion. Generally, values ​​below 20dB indicate unacceptable image quality, 20dB-30dB indicate poor image quality, 30dB-40dB generally indicate good image quality (i.e., distortion is noticeable but acceptable), and values ​​above 40dB indicate excellent image quality (i.e., very close to the original image). Therefore, higher PSNR values ​​indicate better image fidelity, indicating a closer match to the reference.

[0097] In the evaluation of the acquisition method, the present invention not only uses the training time and the Peak Signal-to-Noise Ratio, but also adds the Structural Similarity Index and the data acquisition time as evaluation indicators.

[0098] Data Acquisition Time (DAT) is defined as the total time from device startup to complete acquisition of the data required for the target scene. This metric directly reflects system deployment efficiency. Short acquisition time significantly reduces operating costs and improves practicality. It verifies the algorithm's ability to maintain reconstruction accuracy within a limited time, demonstrating a balance between efficiency and quality.

[0099] The SSIM (structural similarity) formula is used to calculate the difference between the rendered image and the real image as part of the loss function to optimize the quality of scene reconstruction. SSIM can be defined as:

[0100]

[0101] Among them, μ x ,μ y is the mean (brightness) of the two images; σ x ,σ y is the standard deviation (contrast); σ xy is the covariance (structural similarity); C1=(K1L) 2 and C2=(K2L) 2 is a constant to prevent the denominator from being 0. Usually K1=0.01, K2=0.03, and L is the dynamic range of the pixel value.

[0102] The present invention uses the coefficient of determination (R 2 ), root mean square error (RMSE) and mean absolute error (MSA) are used as evaluation indicators.

[0103] R 2 It is used to measure the degree of fit of the model to the data, and its value range is usually [0,1]. Its formula is:

[0104]

[0105] Among them, SSE is the residual sum of squares, SST is the total sum of squares, and y i is the actual measured value, Calculate values ​​for the model, is the mean of the actual measured values.

[0106] RMSE is the square root of the mean square of the error between the predicted value and the actual value, reflecting the overall size of the prediction error. The smaller the value, the more accurate the model prediction. Its formula is:

[0107]

[0108] MAE is the average of the absolute errors between the predicted value and the actual value, which directly reflects the average deviation between the predicted value and the actual value. The formula is:

[0109]

[0110] By R 2 The correlation between calculated and measured values ​​is measured using the RMSE and MAE metrics, respectively. The RMSE emphasizes the impact of large errors, while the MAE more evenly reflects all errors. These metrics validate the effectiveness and reliability of this method for 3D reconstruction and phenotypic trait extraction.

[0111] In this verification, the same dataset, using the same data acquisition method, was reconstructed and rendered using four algorithms: GAPose-GS, 3D Gaussian Splatting, NeRF360, and Instant-NGP. All video data was recorded using a mobile phone (1080p resolution, 60fps). The initial point cloud count was required to be greater than 2000. The quality of the reconstruction was evaluated using training time and PSNR.

[0112] Table 1 shows the training time of the complete dataset and the PSNR of each scene for multiple 3D reconstruction methods in different data acquisition methods. The leftmost column in the table is the data acquisition method, expressed as rotation speed (seconds / turn) / number of shooting angles / number of rotation turns.

[0113] Experimental data shows that GAPose-GS significantly outperforms NeRF and Instant-NGP in reconstruction quality and demonstrates a clear advantage in training efficiency. GAPose-GS achieves an average PSNR of 36.75dB across all data collection methods, with a minimum of 34.28dB. This represents improvements of 69.75%, 39.20%, and 3.38% over NeRF (average 21.65dB), Instant-NGP (average 26.40dB), and 3D Gaussian Splatting (average 35.55dB), respectively. In terms of training efficiency, GAPose-GS takes an average of 41 minutes, significantly faster than NeRF, demonstrating an outstanding balance between efficiency and reconstruction quality.

[0114] This invention uses the mesh model generated by Reality capture as a reference for the geometry extraction results. Figure 7Comparing the reconstruction and rendering results of the same pepper plant using three algorithms, GAPose-GS significantly outperforms in detail restoration and visual fidelity. Leaf veins are clearly visible, stem texture is distinct, the pot edge and background are unblurred, and leaf color transitions are natural. In contrast, the pepper plant leaves reconstructed using NeRF exhibit surface smoothing, with noticeable loss of detail, blurred veins, and blurred leaf overlap. Furthermore, the pepper plant reconstructed using Instant-NGP exhibits issues such as partially broken petioles and excessive background noise, leading to significant errors in plant phenotypic extraction.

[0115] From the data and reconstruction results, it can be concluded that GAPose-GS achieves better reconstruction quality than NeRF360, Instant-NGP, and 3D GaussianSplatting under the condition of ensuring real-time rendering, and reduces the training time, improves the reconstruction efficiency, and has less impact on phenotypic extraction due to reconstruction errors. Although Instant-NGP has an advantage in speed, its PSNR is significantly lower than that of GAPose-GS, and there are problems such as local deletions that make it unsuitable for subsequent phenotypic extraction. Therefore, the present invention uses GAPose-GS to perform multi-viewing of pepper plants. Figure 3 Reconstruction.

[0116] Table 1 Comparison of 3D reconstruction methods

[0117]

[0118]

[0119] Different rotation speeds significantly affect image quality. Furthermore, the number of shooting angles and the number of camera rotations also significantly impact reconstruction quality. Based on GAPose-GS, this paper evaluates reconstruction quality based on camera rotation speed, the number of shooting angles, and the number of rotations, determining an efficient and accurate data acquisition method.

[0120] Table 2 provides a comparison of training time, data acquisition time, PSNR, and SSIM for the same pepper plant using different acquisition methods. The leftmost column in the table shows the data acquisition method, expressed as rotation speed (seconds / turn) / number of shooting angles / number of rotation turns.

[0121] The 17 / 3 / 1 acquisition strategy demonstrated the best overall performance, achieving a PSNR of 37.529dB and an SSIM of 0.9710, both ranking among the top of all methods and demonstrating its superior balance between efficiency and quality. Compared to the 17 / 3 / 2 acquisition strategy in the same group, the 17 / 3 / 1 strategy achieved shorter acquisition times and higher SSIM, highlighting the effectiveness of single-loop acquisition strategies. This strategy improves data coverage density while avoiding the training burden caused by redundant data from multiple loops. Furthermore, the 17 / 3 / 1 acquisition strategy achieved significantly higher SSIM values ​​than other acquisition methods, demonstrating that it more accurately reproduces plant phenotypic details, such as leaf texture and stem geometry, reducing errors in subsequent parameter extraction.

[0122] Therefore, choosing the right acquisition method is crucial for ensuring data quality. In this study, we chose a single rotation, a rotation speed of 17s / rev, and three camera angles to perform multi-view 3D reconstruction of pepper plants using GAPose-GS. This proved to be the optimal choice for balancing efficiency and accuracy in subsequent phenotypic extraction studies.

[0123] Table 2 Comparison of data collection methods

[0124]

[0125] This paper uses GAPose-GS to perform multi-view 3D reconstruction of pepper plants using a data set of 30 pepper plants, each captured from three angles with a rotation speed of 17s / rev, and one rotation. The resulting point cloud data contains redundant background points and noise near the plant point clouds, which can hinder phenotypic extraction. Therefore, the present invention preprocesses the reconstructed 3D point cloud to preserve the 3D spatial extent of the plant and reduce background noise.

[0126] Figure 8 Shows the result of using pass-through filtering on the raw point cloud data of a pepper plant. Figure 8 The area surrounded by the red rectangle in (a) is the pepper plant required for research, and the other points are background information unrelated to the plant. The reconstructed point cloud is filtered to obtain the filtered point cloud as shown below: Figure 8 This method can accurately segment the required area in three-dimensional space, but it is difficult to handle the noise points at the edge and inside of the target. Statistical Outlier RemovalFilter is needed to filter out the noise points. The result of SOR filtering is shown in Figure 9 As shown in the figure, the green point cloud represents the plant point cloud after filtering by the SOR algorithm, and the red points represent noise points. Therefore, SOR filtering can effectively remove the complex distribution of outliers and noise points in the specified three-dimensional canopy space of pepper plants, while preserving the details of leaf tips, edge features, and stems.

[0127] Therefore, by removing redundant background point clouds and close noise points, accurate calculation of pepper plant phenotype extraction is ensured and the computational complexity of subsequent processing is reduced.

[0128] To verify the effectiveness of this method, the plant height of the reconstructed three-dimensional canopy was calculated and the correlation between the calculated value and the measured value was analyzed. To further evaluate the effectiveness of this method, the determination coefficient R 2 , root mean square error RMSE and mean absolute error MSA are used as evaluation indicators.

[0129] A linear correlation is established between the calculated values ​​and the measured values, and a box plot of the measured values ​​is drawn to show their distribution, such as Figure 10 、 Figure 11 、 Figure 12 As shown. In terms of plant height, R 2 The value is 0.997, the RMSE is 0.236, and the MAE is 0.209. In terms of canopy width, R 2 The value is 0.954, RMSE is 1.082, and MAE is 0.880. 2 The value was 0.978, the RMSE was 2.344, and the MAE was 1.965. The data showed that the phenotypic traits calculated based on the reconstructed pepper plant model were highly correlated with the measured values, further verifying the effectiveness of the GAPose-GS algorithm.

Claims

1. A method for extracting plant three-dimensional Gaussian phenotypes based on an improved parameter-adaptive fast pose optimization algorithm, characterized in that: include: Step 1: Use the camera system to acquire a multi-view image sequence of the plant, use the global adaptive fast pose optimization algorithm to estimate the camera pose, and then generate a sparse 3D point cloud of the scene; Step 2: Using the camera pose and sparse 3D point cloud as input, convert the objects in the scene into multiple 3D Gaussian points, perform 3D reconstruction based on the Gaussian representation method, and generate a dense 3D point cloud; Step 3: Extract the phenotypic characteristics of the plant based on the obtained dense three-dimensional point cloud.

2. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 1, characterized in that: The camera system uses a stepper motor to drive a circular track, and the camera is installed on an adjustable shooting rod connected to the circular track. The plant is placed on a shelf in the center of the circular track, and video is captured during rotation. The shooting angles include overhead, horizontal, and overhead shots. During the data acquisition process, the plant height, canopy width, and petiole angle are manually measured, and the obtained video is decomposed into frames to obtain the multi-view image sequence.

3. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 1, characterized in that: The camera pose is estimated using a global adaptive fast pose optimization algorithm, and the implementation method is as follows: Extract local features from the input image; Match input images pairwise using descriptor matching; A global optimization strategy is used to estimate the relative poses of all cameras at once and obtain a sparse 3D point cloud of the scene.

4. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 1, characterized in that: In step 2, first, the sparse three-dimensional point cloud is converted into a dynamically optimized three-dimensional Gaussian ellipsoid as a three-dimensional Gaussian point, each Gaussian ellipsoid is defined by a mean vector μ, a covariance matrix ∑, an opacity α, and spherical harmonic coefficients, wherein the covariance matrix ∑ is decomposed into: ∑=SS T R T The initial scaling radius is calculated based on the K-nearest neighbor algorithm, so that the Gaussian ellipsoid can cover the local geometric structure. The color properties are modeled by spherical harmonics, and the order of the spherical harmonics determines the ability to express the view-dependent illumination changes. Then, the three-dimensional Gaussian ellipsoid is projected onto the two-dimensional image space ellipse for rendering to generate an image. Given the viewing transformation W, the covariance matrix ∑ ′ as follows: ∑ ′ =JW∑W T J T where J is the Jacobian matrix of the affine approximation of the projective transformation.

5. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 4, characterized in that: The step 2 further comprises: Differentiable optimization is performed on the 3D reconstruction, and the position, shape and opacity of the 3D Gaussian points are optimized by back propagation.

6. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 5, characterized in that: The differentiable optimization is a continuous iterative process based on rendering. The generated image is compared with the real data in the training view. During the backpropagation process, the gradient information optimizes the Gaussian parameters and drives the dynamic density adjustment mechanism. In particular, for areas where the gradient amplitude is greater than a threshold, the Gaussian density is increased through cloning or splitting operations to improve details. Gaussians with transparency below the threshold are pruned to reduce computational overhead.

7. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 1, characterized in that: In step 3, the reconstructed dense 3D point cloud is preprocessed as follows: First, straight-through filtering is used to remove invalid background point clouds; Secondly, SOR filtering is used to remove noise.

8. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter-adaptive fast pose optimization algorithm according to claim 7, characterized in that: The straight-through filtering method first specifies a dimension and a range for that dimension, then iterates over each point in the point cloud, checking whether the value of that point in the specified dimension is within the range. Points that are not within the range are deleted. After the traversal is completed, the remaining points constitute the filtered point cloud. The SOR filter sets a density threshold. When the point cloud density in a certain area is lower than the threshold, the point cloud is considered invalid.

9. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 1, characterized in that: The step 3 is to extract the phenotypic characteristics of the plant, and the implementation method is as follows: The AlphaShape algorithm was used to obtain the edge contours of the plant's vertical direction and canopy from the 3D reconstructed model. Based on the extracted contours, in a rectangular coordinate system, the direction from the xz plane and the direction perpendicular to the horizontal ground were used as the y-axis. By obtaining the lowest and highest points of the central axis, the vertical distance between the two planes, i.e., the plant height, was calculated based on the lowest and highest points of the plant height. The canopy width is expressed by the average value of the maximum canopy width in the two horizontal directions. The vertical distance of the farthest horizontal point cloud of the canopy projected onto the ground is defined as the maximum length of the canopy. The corresponding longitudinal distance is defined as the maximum width of the canopy. The average of the sum of the maximum length and maximum width of the canopy is the canopy width. The stem-leaf angle is calculated by the angle between the two space vectors: the growth direction of the leaf at birth and the local growth direction of the stem.

10. The method for extracting plant three-dimensional Gaussian phenotypes based on the improved parameter adaptive fast pose optimization algorithm according to claim 9, characterized in that: Connect any point A on the main stem of the plant with the petiole and any end point B above point A to form a line segment; connect vertex A with end point C on the petiole to form a segment; calculate the petiole angle using the coordinate information of points A, B, and C and the lengths of line segments AB and AC.

Citation Information

Cited By

  • Crop phenotype in-situ analysis method and device, electronic equipment and storage medium

    CN121214383A

  • Paddy rice fine phenotype extraction method based on three-dimensional vision, terminal and storage medium

    CN121725184A

  • Rice fine phenotype extraction method based on three-dimensional vision, terminal and storage medium

    CN121725184B