Plant dense 3D point cloud reconstruction method and system
By using consumer-grade devices such as smartphones and a graded exposure control strategy, combined with geometrically constrained 3D Gaussian sputtering rendering and adaptive dense sampling, low-cost and highly applicable dense 3D point cloud reconstruction of plants was achieved. This solved the problems of complex operation and long time consumption in existing technologies, and generated high-precision and complete 3D point clouds of plants.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING RES CENT FOR INFORMATION TECH & AGRI
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies struggle to achieve low-cost, highly applicable 3D point cloud reconstruction of plants in the agricultural sector, especially in outdoor environments with strong sunlight where transparent or reflective features cannot be handled. Furthermore, traditional methods are complex and time-consuming, failing to meet the actual needs of agricultural personnel.
Consumer-grade image/video acquisition devices such as smartphones are used. Image exposure compensation is performed through a graded exposure control strategy. Sparse point cloud reconstruction is performed by combining a 3D reconstruction network. Dense 3D point cloud of plants is generated by using geometrically constrained 3D Gaussian sputtering rendering and adaptive densification sampling.
It significantly improves the accuracy, completeness, and practicality of 3D plant reconstruction, generates dense 3D point clouds with rich details and complete structure, solves the problem of inconsistent image quality caused by changes in lighting, and simplifies the operation process.
Smart Images

Figure CN122115705A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plant 3D phenotyping technology, and in particular to a method and system for reconstructing dense 3D point clouds of plants. Background Technology
[0002] Plant 3D data acquisition primarily relies on specialized 3D imaging equipment, such as terrestrial LiDAR, structured light scanners, and time-of-flight cameras. Researchers have explored alternatives based on consumer-grade devices, but none have successfully transitioned from small-scale laboratory validation to practical application by agricultural workers in the field. Early multi-view reconstruction based on consumer-grade cameras suffered from severe geometric distortions in the reconstructed models due to unstable device intrinsic parameters and lens distortion. While depth sensor-based solutions can quickly acquire 3D data, they fail in strong outdoor light conditions and cannot handle the transparent or reflective features of plant surfaces.
[0003] Traditional Structure from Motion (SfM) and Multi-View Stereo (MVS) methods can theoretically use ordinary cameras, but in practice, they require precise camera calibration, a stable shooting environment, and high image overlap. Processing times can reach several hours, and operators must possess certain professional knowledge to adjust parameters and handle failures. While these techniques can obtain high-quality 3D plant point clouds under specific conditions, they have consistently failed to solve the fundamental problems of practicality and robustness in 3D plant data acquisition. Summary of the Invention
[0004] This invention provides a method and system for reconstructing dense 3D point clouds of plants, which addresses the shortcomings of existing methods for generating 3D point clouds of plants in terms of practicality. It enables the reconstruction of dense 3D point clouds of plants using multi-view video / image data acquired by consumer-grade image / video acquisition devices such as smartphones, significantly improving the quality of the reconstructed point clouds.
[0005] This invention provides a method for reconstructing dense 3D point clouds of plants, comprising the following steps: acquiring a multi-view image sequence, wherein the multi-view image sequence is a sequence of plant videos or images acquired by an image acquisition device; performing image exposure compensation preprocessing based on the multi-view image sequence to obtain an exposure-compensated image sequence, wherein the image exposure compensation preprocessing adopts a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequence to a unified target range; reconstructing a sparse point cloud based on the exposure-compensated image sequence to obtain a sparse point cloud and a camera trajectory; performing geometrically constrained 3D Gaussian sputtering rendering based on the sparse point cloud and the camera trajectory to obtain an optimized 3D Gaussian model; and performing adaptive densification sampling on the optimized 3D Gaussian model to generate a dense 3D point cloud of plants.
[0006] According to the present invention, a method for reconstructing dense 3D point clouds of plants includes reconstructing sparse point clouds based on the exposure-compensated image sequence to obtain sparse point clouds and camera trajectories. The method comprises: inputting the exposure-compensated image sequence into a 3D reconstruction network and extracting multi-view image features through an encoder; inputting the multi-view image features into a fusion transformer module and fusing cross-view features through a self-attention mechanism and a cross-attention mechanism to obtain fused features; simultaneously inputting the fused features into parallel-connected global and local branches; performing global geometric structure understanding and camera pose estimation on the fused features through the global branch to obtain global feature information; performing local feature matching and depth information prediction on the fused features through the local branch to obtain local feature information; and reconstructing sparse point clouds based on the global and local feature information to obtain sparse point clouds and camera trajectories.
[0007] According to the present invention, a method for reconstructing dense 3D point clouds of vegetation includes a geometrically constrained 3D Gaussian sputtering rendering based on the sparse point cloud and the camera trajectory to obtain an optimized 3D Gaussian model. The method comprises: initializing a 3D Gaussian model based on the sparse point cloud, wherein the 3D Gaussian model is represented by multiple Gaussian primitives, each containing a center position, covariance matrix, opacity, and color attribute; performing differentiable rendering on the 3D Gaussian model based on the camera trajectory to obtain a predicted color map and a predicted depth map for each view; calculating the color rendering loss of the 3D Gaussian model based on the predicted color map and the corresponding real image; and performing single-view rendering based on the predicted depth map. View geometry regularization is performed, and single-view geometry consistency loss is calculated, wherein the single-view geometry regularization includes depth smoothing constraints and surface normal vector consistency constraints. Based on the predicted depth map and the camera trajectory, multi-view geometry regularization is performed, and multi-view geometry consistency loss is calculated, wherein the multi-view geometry regularization includes depth consistency constraints, epipolar geometry constraints, and cross-view normal vector consistency constraints. Combining the color rendering loss of the 3D Gaussian model, the single-view geometry consistency loss, and the multi-view geometry consistency loss, a total loss function is constructed. By minimizing the total loss function, the parameters of the 3D Gaussian model are iteratively optimized to obtain the optimized 3D Gaussian model.
[0008] According to the present invention, a method for reconstructing dense 3D point clouds of plants is provided, wherein the depth smoothing constraint is expressed by the following formula: in, This indicates a depth smoothing constraint. Represents the image domain. Indicates pixel electricity The set of neighboring pixels, and Representing pixels With pixels The depth value, Represents the weighting function; The surface normal vector consistency constraint is expressed by the following formula: in, and These represent the pixels on the predicted depth map of a single view. With pixels The surface normal vector at that location.
[0009] According to the method for reconstructing dense 3D point clouds of plants provided by the present invention, the depth consistency constraint is expressed by the following formula: in, Represents the collection of all view pairs. For view pairs The set of corresponding pixels between them This indicates that the same 3D point is in the view. and view The pixel projection position in the image. Indicates in view medium pixel The predicted depth value, Indicates from view Reprojection to view The depth value after that, It is the Huber robust loss function; The epipolar geometric constraint is expressed by the following formula: in, and They are pixels With pixels homogeneous coordinates Indicates transpose. Represents view pairs The fundamental matrix between them; The cross-view normal vector consistency constraint is expressed by the following formula: in, Indicates in view medium pixel Surface normal vector, Indicates in view medium pixel Surface normal vector, Represents a view To view The relative rotation matrix.
[0010] According to the present invention, a method for reconstructing dense 3D point clouds of plants is provided. The method involves adaptively dense sampling of the optimized 3D Gaussian model to generate a dense 3D point cloud of plants, comprising: for each Gaussian element in the optimized 3D Gaussian model, calculating the ellipsoidal volume to shape ratio based on the scale parameter corresponding to its covariance matrix; determining the basic number of sampling points based on the ellipsoidal volume, and calculating an adaptive sampling density adjustment factor based on the shape ratio; determining the final number of sampling points for each Gaussian element according to the basic number of sampling points and the adaptive sampling density adjustment factor; and performing point sampling in the ellipsoidal space of each Gaussian element according to the final number of sampling points based on Mahalanobis distance constraints to generate a dense 3D point cloud of plants.
[0011] This invention also provides a dense 3D point cloud reconstruction system for plants, comprising the following modules: an acquisition module for acquiring multi-view image sequences, wherein the multi-view image sequences are plant videos or RGB image sequences acquired by an image acquisition device; an exposure compensation module for performing image exposure compensation preprocessing based on the multi-view image sequences to obtain an exposure-compensated image sequence, wherein the image exposure compensation preprocessing adopts a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequences to a unified target range; a sparse point cloud reconstruction module for performing sparse point cloud reconstruction based on the exposure-compensated image sequence to obtain a sparse point cloud and a camera trajectory; a 3D Gaussian sputtering module for performing geometrically constrained 3D Gaussian sputtering rendering based on the sparse point cloud and the camera trajectory to obtain an optimized 3D Gaussian model; and a dense point cloud generation module for performing adaptive dense sampling on the optimized 3D Gaussian model to generate a dense 3D point cloud of plants.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the plant dense 3D point cloud reconstruction method as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the plant dense 3D point cloud reconstruction method as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the plant dense 3D point cloud reconstruction method as described above.
[0015] The method and system for reconstructing dense 3D point clouds of plants provided by this invention effectively overcome the problem of inconsistent image quality caused by changes in lighting in plant scenes by performing exposure compensation preprocessing on multi-view image sequences through a hierarchical exposure control strategy, providing a reliable input for exposure standardization for subsequent reconstruction. Furthermore, 3D Gaussian sputtering rendering based on geometric constraints can make full use of the structural information contained in the sparse point cloud and camera trajectory to achieve high-fidelity modeling of the complex geometry and appearance of plants. Finally, through adaptive dense sampling, a dense 3D point cloud of plants with rich details and complete structure is generated on the optimized 3D Gaussian model, significantly improving the accuracy, completeness and practicality of plant 3D reconstruction. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the plant dense 3D point cloud reconstruction method provided by the present invention.
[0018] Figure 2 This is a schematic diagram of the modules of the plant dense 3D point cloud reconstruction system provided by the present invention.
[0019] Figure 3 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0021] The emergence of 3D Gaussian Splatting (3DGS) technology has provided new possibilities for low-cost acquisition of 3D plant data. Compared to Neural Radiance Fields (NeRF), which requires long training times, 3DGS can complete high-quality 3D reconstruction within minutes. Previous research has explored the application potential of 3DGS in agriculture, with PlantGaussian demonstrating its potential for 3D visualization and structural analysis of plants at different growth stages and across different scenes. The P3DFusion framework combines a visual base model with 3DGS technology to achieve high-fidelity 3D reconstruction of plants across scenes, significantly improving reconstruction quality in complex backgrounds. Point clouds, as a standard 3D representation for plant phenotypic analysis, contain rich geometric information, facilitating subsequent analysis, processing, and parameter extraction.
[0022] The main drawbacks of the related technologies are as follows: Traditional methods for acquiring 3D point cloud data of plants suffer from problems such as high cost, large site requirements, complex operation, and heavy workload of data preprocessing, making them unfriendly to researchers with an agricultural background and small teams.
[0023] Existing technologies position 3DGS as a visual rendering tool rather than a geometric measurement tool, resulting in reconstruction results that, while visually appealing, are difficult to use directly for precise phenotypic parameter extraction. Furthermore, 3DGS essentially represents scenes using numerous three-dimensional Gaussian ellipsoids, and its optimization process primarily relies on photometric consistency loss, lacking explicit geometric constraints, leading to geometric inconsistencies in the reconstruction results.
[0024] Existing methods for converting 3DGS models to point clouds simply extract the Gaussian center points, ignoring the shape, scale, and orientation information of the ellipsoid. The resulting sparse point clouds are far below the density standard of laser scanning. This series of technical challenges constitutes the core obstacle to achieving low-cost, low-barrier acquisition of plant 3D phenotypic data.
[0025] This invention addresses the lack of low-cost and highly applicable plant 3D data acquisition methods for agricultural professionals in plant 3D phenotyping research. It proposes a low-cost method for reconstructing dense 3D point clouds of plants based on geometrically constrained 3DGS. This method is applicable to the reconstruction of dense 3D point clouds of plants from videos captured by consumer-grade devices such as smartphones, action cameras, and panoramic cameras, providing technical support for plant 3D data acquisition, plant 3D phenotyping research, and applications.
[0026] Figure 1 This is a flowchart illustrating the plant dense 3D point cloud reconstruction method provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps.
[0027] Step 101: Obtain a multi-view image sequence, wherein the multi-view image sequence is a sequence of plant videos or images acquired by an image acquisition device.
[0028] In this embodiment of the invention, the reconstruction of dense 3D point clouds of plants using videos collected by consumer devices such as mobile phones and algorithms such as 3DGS specifically includes three core steps: sparse point cloud reconstruction, geometrically constrained 3DGS rendering, and 3DGS adaptive densification sampling.
[0029] In some embodiments, a multi-view image sequence covering the overall structure of a plant is obtained by using consumer-grade image acquisition devices such as smartphones, action cameras, or panoramic cameras to capture videos or take multiple photos from different angles.
[0030] For example, operators can hold the device and take photos at a constant speed around the plant to be reconstructed to ensure the acquisition of a continuous image sequence with sufficient visual overlap; or, from a fixed shooting point, they can adjust the device's pitch angle to acquire images from multiple heights. The resulting multi-view image sequence will serve as the raw data input for all subsequent processing stages.
[0031] Step 102: Perform image exposure compensation preprocessing based on the multi-view image sequence to obtain the exposure-compensated image sequence.
[0032] Among them, the image exposure compensation preprocessing adopts a hierarchical exposure control strategy to standardize the exposure state of multi-view image sequences to a unified target range.
[0033] In this embodiment of the invention, keyframe images are extracted from multi-view plant videos, or RGB multi-view image sequences are used as input. Exposure compensation preprocessing is performed on the input multi-view image sequences, specifically employing a hierarchical exposure control strategy. This involves an intelligent combination of adaptive histogram equalization (CLAHE), gamma correction, linear brightness adjustment, and contrast enhancement techniques. The most suitable processing mode is automatically selected based on the image exposure state, standardizing the exposure state of all images to a unified target range. Exposure differences between images are automatically detected and corrected to ensure the consistency and accuracy of subsequent feature extraction.
[0034] In some embodiments, the acquired multi-view image sequence undergoes image exposure compensation preprocessing to obtain an exposure-compensated image sequence. Specifically, this preprocessing employs a hierarchical exposure control strategy. First, the exposure state of each image in the sequence is automatically analyzed, for example, by calculating the overall brightness histogram of the image. Then, based on the evaluation results, the images are categorized into different exposure levels, such as underexposed, normally exposed, or overexposed. For images of different levels, a series of image enhancement algorithms are adaptively selected and combined, including adaptive histogram equalization, gamma correction, linear brightness adjustment, and contrast enhancement techniques. Through hierarchical processing, the visual performance (e.g., brightness and contrast) of all input images is adjusted and standardized to a preset unified target range, thereby effectively correcting exposure differences between images caused by variations in lighting conditions during shooting, and providing an exposure-compensated image sequence with consistent lighting conditions for subsequent processing steps.
[0035] Step 103: Sparse point cloud reconstruction is performed based on the image sequence after exposure compensation to obtain sparse point cloud and camera trajectory.
[0036] In this embodiment of the invention, the image sequence, after exposure compensation preprocessing, is input into an end-to-end 3D reconstruction network based on the Fast3R framework for processing. In the 3D reconstruction network, depth features are extracted from each image via an encoder. Subsequently, the depth features are fed into a Fusion Transformer module, which utilizes self-attention and cross-attention mechanisms to achieve feature fusion and geometric consistency constraints across images from different viewpoints, thereby establishing robust inter-image correspondences. The output of the Fusion Transformer is connected in parallel to two branches: a Global Head and a Local Head. The Global Head is responsible for parsing the global geometry of the scene and estimating the approximate camera trajectory (i.e., camera pose); the Local Head focuses on fine-grained matching of local features and accurate prediction of depth information.
[0037] Through the design of co-processing global and local information, the network can simultaneously output a sparse point cloud representing the spatial structure of the plant and a camera trajectory describing the changes in the spatial position and orientation of the image acquisition device during the shooting process.
[0038] Step 104: Perform geometrically constrained 3D Gaussian sputtering rendering based on sparse point cloud and camera trajectory to obtain an optimized 3D Gaussian model.
[0039] In this embodiment of the invention, a sparse point cloud is used as the initial point set, and a three-dimensional Gaussian primitive is instantiated for each point to construct an initial 3D Gaussian model. Each Gaussian primitive is defined by parameters such as its center position, covariance matrix, opacity, and view-dependent color.
[0040] In the process of using camera trajectory for differentiable rendering and model optimization, multi-level geometric constraints are introduced, specifically including single-view geometric regularization and multi-view geometric regularization.
[0041] Single-view geometric regularization: On the rendering result of a single view, a depth map is calculated and an adaptive depth smoothing loss is introduced. This loss can adjust the constraint strength according to the image color edge information to maintain the sharpness of object edges while promoting the smoothness of flat areas. At the same time, surface normal vectors are derived from the depth map and normal vector consistency loss is applied to enhance the geometric rationality of local surfaces.
[0042] Multi-view geometric regularization: Apply reprojection depth consistency loss between different views to ensure that the depth observation value of the same 3D point is consistent in multiple views; introduce epipolar geometric constraints to make the matched pixel pairs satisfy the fundamental matrix relationship; and calculate cross-view normal vector consistency loss to require that the normal vector direction of the same surface point remains consistent in different views after coordinate system transformation.
[0043] The final optimization objective function is composed of the original color difference-based rendering loss, combined with the weighted sum of the single-view geometric constraint loss and the multi-view geometric constraint loss. By iteratively optimizing the parameters of the 3D Gaussian model, an optimized 3D Gaussian model with high consistency and accuracy in both photometric and geometric aspects is finally obtained.
[0044] Step 105: Adaptive dense sampling is performed on the optimized 3D Gaussian model to generate a dense 3D point cloud of plants.
[0045] In this embodiment of the invention, the following operations are performed for each three-dimensional Gaussian element in the optimized 3D Gaussian model.
[0046] First, the volume and shape ratio of the Gaussian primitive are calculated. Then, the number of sampling points for the primitive is dynamically determined based on the calculated shape ratio; the larger the shape ratio, the more sampling points are allocated. Finally, based on Mahalanobis distance constraints, a corresponding number of 3D sampling points are generated within the ellipsoid defined by the Gaussian primitive. By traversing and processing all Gaussian primitives in the model and merging the 3D point sets generated by each primitive, a dense 3D point cloud of vegetation is finally output.
[0047] Through the embodiments of this invention, exposure compensation preprocessing of multi-view image sequences is performed using a graded exposure control strategy, effectively overcoming the problem of inconsistent image quality caused by changes in illumination in plant scenes, and providing reliable input for exposure standardization for subsequent reconstruction. Furthermore, 3D Gaussian sputtering rendering based on geometric constraints can fully utilize the structural information contained in sparse point clouds and camera trajectories to achieve high-fidelity modeling of the complex geometry and appearance of plants. Finally, through adaptive dense sampling, a dense 3D point cloud of plants with rich details and complete structure is generated on the optimized 3D Gaussian model, significantly improving the accuracy, completeness and practicality of plant 3D reconstruction.
[0048] According to the present invention, a method for reconstructing dense 3D point clouds of plants is provided, which reconstructs sparse point clouds based on an image sequence after exposure compensation to obtain sparse point clouds and camera trajectories, including: The exposure-compensated image sequence is input into a 3D reconstruction network, and multi-view image features are extracted by an encoder. Multi-view image features are input into the fusion transformer module, and cross-view feature fusion is performed through self-attention mechanism and cross-attention mechanism to obtain fused features; The fused features are simultaneously input into the global and local branches of the parallel connection; Global feature information is obtained by performing global geometric structure understanding and camera pose estimation on the fused features through global branching; Local feature information is obtained by performing local feature matching and depth information prediction on the fused features through local branches; Sparse point cloud reconstruction is performed based on global and local feature information to obtain sparse point cloud and camera trajectory.
[0049] In this embodiment of the invention, the exposure-compensated image sequence is input into the complete network architecture of Fast3R for end-to-end 3D reconstruction processing. The features output by the encoder then enter the core Fusion Transformer module, which achieves cross-view feature fusion and geometric consistency constraints through advanced self-attention and cross-attention mechanisms, establishing a more robust matching effect than traditional handcrafted features. The output of the Fusion Transformer is simultaneously connected to two parallel branches: GlobalHead and LocalHead. The GlobalHead is responsible for understanding the global geometric structure and coarsely estimating the camera pose, while the LocalHead focuses on fine-grained matching of local features and accurate prediction of depth information. Through this global-local collaborative design, the system can output high-quality sparse point cloud structures and accurate camera motion trajectories, which can be directly used as input for 3DGS training.
[0050] For example, an exposure-compensated image sequence is input into a 3D reconstruction network. The network's encoder extracts features from the input images, obtaining multi-view image features. These multi-view image features are then fed into a fusion transformer module. This module utilizes a self-attention mechanism to handle feature correlations within a single viewpoint and establishes feature correspondences between different viewpoints through a cross-attention mechanism, thereby achieving cross-view feature fusion and outputting the fused features.
[0051] The fused features are simultaneously fed into two parallel branches: a global branch and a local branch. In the global branch, the fused features are used to understand the global geometry and to perform a coarse estimation of the camera pose, thus outputting global feature information containing overall scene layout and camera motion information. In the local branch, the focus is on fine-grained matching of local features and accurate prediction of depth information, outputting local feature information.
[0052] By integrating global and local feature information, joint optimization and 3D reconstruction calculations are performed to generate a sparse point cloud representing the spatial structure of the plant and a camera trajectory describing the camera's spatial position and orientation during image acquisition.
[0053] Through the embodiments of the present invention, by fusing global feature information and local feature information, and combining self-attention and cross-attention mechanisms, the matching accuracy and geometric consistency of multi-view image features are effectively improved, thereby achieving more accurate and robust sparse point cloud reconstruction and camera trajectory estimation in complex plant scenes.
[0054] According to the present invention, a method for reconstructing dense 3D point clouds of plants is provided, which performs geometrically constrained 3D Gaussian sputtering rendering based on sparse point clouds and camera trajectories to obtain an optimized 3D Gaussian model, including: A 3D Gaussian model is initialized based on sparse point cloud. The 3D Gaussian model is represented by multiple Gaussian elements, each of which contains the center position, covariance matrix, opacity and color attribute. Based on the camera trajectory, a differentiable rendering of the 3D Gaussian model is performed to obtain the predicted color map and predicted depth map of each view. Based on the predicted color map and the corresponding real image, calculate the color rendering loss of the 3D Gaussian model; Based on the predicted depth map, single-view geometric regularization is performed, and single-view geometric consistency loss is calculated. The single-view geometric regularization includes depth smoothing constraint and surface normal vector consistency constraint. Based on the predicted depth map and camera trajectory, multi-view geometric regularization is performed, and multi-view geometric consistency loss is calculated. The multi-view geometric regularization includes depth consistency constraint, epipolar geometry constraint and cross-view normal vector consistency constraint. A total loss function is constructed by combining the color rendering loss of the 3D Gaussian model, the single-view geometric consistency loss, and the multi-view geometric consistency loss. By minimizing the total loss function, the parameters of the 3D Gaussian model are iteratively optimized to obtain the optimized 3D Gaussian model.
[0055] The single-view geometric regularization and multi-view geometric regularization in the embodiments of the present invention are described below.
[0056] In single-view geometric regularization, for each (3D) Gaussian element... ,in As the central location, Let covariance matrix be the variance matrix. For opacity, For color, pixels Depth value at and weight function The calculation method is shown in formulas (1) and (2): (1) (2) in, Indicates the first A Gaussian unit at a pixel Opacity at the location, Indicates ranking in The first before Gaussian Yuan A Gaussian unit at a pixel Opacity at the location, The depth distance from the center of the Gaussian unit to the camera is represented by the weight function. Strictly adhere to alpha blending rules to ensure consistency between depth rendering and color rendering.
[0057] According to the method for reconstructing dense 3D point clouds of plants provided by the present invention, the depth smoothing constraint is expressed by the following formula: in, This indicates a depth smoothing constraint. Represents the image domain. Indicates pixel electricity The set of neighboring pixels, and Representing pixels With pixels The depth value, Represents the weighting function; Surface normal vector consistency constraints are expressed by the following formula: in, and These represent the pixels on the predicted depth map of a single view. With pixels The surface normal vector at that location.
[0058] In this embodiment of the invention, in order to ensure the spatial continuity of the predicted depth map (rendered depth map) and avoid unreasonable depth jumps, an adaptive depth smoothing constraint is introduced, as shown in formula (3). This constraint, while maintaining depth continuity, can adaptively adjust the constraint strength according to the image content: (3) in Represents the image domain. Let p be the neighborhood of pixel p. The weighting function combines spatial distance and color similarity: (4) in, This represents the parameter used to control the rate of decay of spatial distance weights. This represents a parameter used to control the rate at which color difference weights decay. and Representing pixels With pixels The color value at that location.
[0059] Based on the predicted depth map obtained from the rendering, surface normal vectors are calculated to further constrain the geometry. The surface normal vectors are calculated using gradient information from the depth map: (5) in and These represent the gradients of the predicted depth map in the x and y directions, respectively. To ensure local smoothness of the surface, a normal vector consistency loss is introduced, requiring the normal vector directions of adjacent pixels to be coordinated. (6) The final total loss function for single-view geometric consistency optimization is: (7) in and These are the weights of the smoothing constraint and the normal vector constraint, respectively.
[0060] Through the embodiments of the present invention, the above constraints effectively suppress noise and discontinuities in the depth map, improve the smoothness of depth prediction and surface geometric consistency; the weight function adaptively adjusts the neighborhood influence, preserves the details of plant edges while avoiding excessive smoothing, and enhances the structural integrity and realism of the reconstructed point cloud.
[0061] According to the method for reconstructing dense 3D point clouds of plants provided by the present invention, the depth consistency constraint is expressed by the following formula: in, Represents the collection of all view pairs. For view pairs The set of corresponding pixels between them This indicates that the same 3D point is in the view. and view The pixel projection position in the image. Indicates in view medium pixel The predicted depth value, Indicates from view Reprojection to view The depth value after that, It is the Huber robust loss function; The polar geometry constraint is expressed by the following formula: in, and They are pixels With pixels homogeneous coordinates Indicates transpose. Represents view pairs The fundamental matrix between them; The cross-view normal vector consistency constraint is expressed by the following formula: in, Indicates in view medium pixel Surface normal vector, Indicates in view medium pixel Surface normal vector, Represents a view To view The relative rotation matrix.
[0062] In multi-view geometry regularization, given a pair of views ( ) and the corresponding camera parameters, can view pixels in Reprojection to view In the middle, the corresponding pixel position is obtained. The multi-view depth consistency loss is defined as shown in Equation (8): (8) in Represents the collection of all view pairs. For view pairs The set of corresponding pixels between them Indicates from view Reprojection to view The depth value after that, It is the Huber robust loss function, which can effectively handle outlier occlusion.
[0063] To further enhance the reliability of geometric constraints, classic epipolar geometric constraints are introduced. For the corresponding pixel pairs... They should satisfy the epipolar geometry relation: (9) in and These are the homogeneous coordinates of the pixels. It is a view pair The fundamental matrix between them. This constraint ensures that corresponding point pairs satisfy basic geometric relationships, improving the accuracy of geometric reconstruction.
[0064] Considering the differences in the representation of normal vectors across different coordinate systems, appropriate coordinate transformations are required. Cross-view normal vector consistency constraints are implemented using the following loss function: (10) in It is the relative rotation matrix between views. and Let represent the rotation matrices of the two views, respectively. This constraint ensures that the normal vector of the same surface point has a consistent direction in different views. The multi-view geometric consistency loss is: (11) Final training loss Loss from original 3DGS reconstruction Single-view loss Multi-view loss composition.
[0065] (12) Through the embodiments of the present invention, the above-mentioned multi-view constraints effectively improve the consistency of depth and normal vector prediction: depth consistency combined with reprojection and robust loss suppresses mismatches; epipolar geometry constraints enhance the accuracy of camera geometric relationships; cross-view normal vector consistency utilizes rotation alignment to ensure surface orientation continuity, jointly improving the accuracy and robustness of plant complex structure reconstruction.
[0066] According to the present invention, a method for reconstructing dense 3D point clouds of plants is provided, which involves adaptively densifying the sampling of an optimized 3D Gaussian model to generate a dense 3D point cloud of plants, including: For each Gaussian element in the optimized 3D Gaussian model, the ratio of ellipsoidal volume to shape is calculated based on the scale parameter corresponding to its covariance matrix. The number of basic sampling points is determined based on the ellipsoidal volume, and the adaptive sampling density adjustment factor is calculated based on the shape ratio. The final number of sampling points for each Gaussian unit is determined based on the base number of sampling points and the adaptive sampling density adjustment factor. Based on Mahalanobis distance constraints, point sampling is performed in the ellipsoidal space of each Gaussian element according to the final number of sampling points to generate a dense 3D point cloud of vegetation.
[0067] In this embodiment of the invention, the number and distribution strategy of sampling points are dynamically determined based on the volume and shape characteristics of each Gaussian ellipsoid, thereby maximizing the preservation of geometric information while maintaining computational efficiency. An ellipsoid shape ratio is introduced as an adaptive factor to differentiate different geometric features, as detailed below.
[0068] For slender ellipsoids with obvious directionality, increase the sampling density to capture surface details; for ellipsoids that are close to spherical, maintain moderate sampling to avoid redundancy. The specific calculation process is shown in formulas (12) to (14): (12) (13) (14) in, Represents the approximate volume of a single (3D) Gaussian element. This indicates the shape ratio of the (3D) Gaussian element, used to measure its degree of anisotropy or stretching of shape. This is represented by the adaptive sampling threshold dynamically calculated by the Gaussian element. These represent the scale (or radius) of the Gaussian element (ellipsoid) along the x, y, and z axes, respectively. Indicates the basic sampling threshold. This represents the control coefficient. The natural logarithm of the shape ratio AR.
[0069] The algorithm first calculates the precise volume and shape features of each Gaussian ellipsoid, and then dynamically adjusts the sampling threshold based on this information to ensure that appropriate point cloud density can be obtained in Gaussian regions of different scales and shapes. Subsequently, a sampling strategy constrained by Mahalanobis distance is used to generate dense point clouds that conform to the distribution characteristics of each Gaussian ellipsoid, thereby preserving the geometric structure and detail information of the original 3DGS.
[0070] For each Gaussian element in the optimized 3D Gaussian model, scale parameters corresponding to the three principal axes of its covariance matrix are extracted. Based on these scale parameters, the volume-to-shape ratio of the ellipsoid represented by the Gaussian element is calculated. The volume is calculated as the product of the three scale parameters, while the shape ratio is obtained by calculating the ratio of the largest to the smallest scale parameter, which reflects the degree of anisotropy of the Gaussian element.
[0071] The number of basic sampling points is determined based on the ellipsoidal volume. Larger Gaussian elements represent a larger spatial range, thus requiring a corresponding increase in the number of basic sampling points. Simultaneously, an adaptive sampling density adjustment factor is calculated based on the shape ratio. When the shape ratio is close to 1, it indicates that the Gaussian element is nearly spherical, and the adjustment factor is relatively small; when the shape ratio is significantly greater than 1, it indicates that the Gaussian element is elongated or flat, and the adjustment factor is correspondingly increased to ensure denser sampling of regions with significant directional characteristics.
[0072] Based on the base number of sampling points and the adaptive sampling density adjustment factor, the final number of sampling points for each Gaussian element is determined through weighted or multiplicative operations. This step enables the adaptive adjustment of the sampling density according to the geometric characteristics of the Gaussian elements.
[0073] Based on Mahalanobis distance constraints, probabilistic sampling is performed within the ellipsoidal space of each Gaussian primitive according to the final number of sampling points. The Mahalanobis distance constraint ensures that the sampling points are generated within the ellipsoidal probability distribution defined by the Gaussian primitive, making the sampling point distribution conform to the characteristics of a Gaussian distribution. By traversing and processing all Gaussian primitives and merging the 3D point sets generated by each primitive sampling, a dense 3D point cloud of plants is finally output. This method, through adaptive adjustment of the sampling strategy, significantly improves the density and fidelity of the point cloud at plant detail features while maintaining overall computational efficiency.
[0074] Through the above embodiments of the present invention, a Fast3R sparse reconstruction method with integrated exposure compensation is proposed. This method solves the stability problem of consumer devices under varying lighting conditions through adaptive exposure optimization, significantly shortens processing time, and provides more reliable geometric initialization.
[0075] A geometrically constrained 3DGS rendering framework is constructed and an adaptive dense sampling algorithm is proposed. Multi-level geometric constraints are introduced at the single-view and multi-view levels to ensure the geometric accuracy of the reconstruction results. Furthermore, the sampling density is dynamically adjusted according to the geometric characteristics of the Gaussian ellipsoid to convert the 3DGS representation into a high-fidelity dense point cloud that meets the requirements of phenotypic analysis.
[0076] By utilizing multi-view video / image data acquired through consumer-grade image / video acquisition devices such as smartphones, dense 3D point cloud reconstruction of plants can be achieved, significantly improving the quality of the reconstructed point cloud. This results in a complete and clear point cloud structure with less environmental noise, sharp and continuous edge contours of organs such as leaves, and complete preservation of geometric details. It effectively solves common problems in traditional methods such as noise, voids, and geometric distortion, and can achieve the reconstruction of dense 3D point clouds of plants.
[0077] The following describes the plant dense 3D point cloud reconstruction system provided by the present invention. The plant dense 3D point cloud reconstruction system described below can be referred to in correspondence with the plant dense 3D point cloud reconstruction method described above.
[0078] refer to Figure 2 , Figure 2 This is a schematic diagram of the modules of the plant dense 3D point cloud reconstruction system provided by the present invention.
[0079] The acquisition module 201 is used to acquire multi-view image sequences, wherein the multi-view image sequences are plant videos or RGB image sequences acquired by an image acquisition device; The exposure compensation module 202 is used to perform image exposure compensation preprocessing based on multi-view image sequences to obtain an exposure-compensated image sequence. The image exposure compensation preprocessing adopts a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequence to a unified target range. The sparse point cloud reconstruction module 203 is used to reconstruct sparse point clouds based on the image sequence after exposure compensation, so as to obtain sparse point clouds and camera trajectories. The 3D Gaussian sputtering module 204 is used for geometrically constrained 3D Gaussian sputtering rendering based on sparse point clouds and camera trajectories to obtain an optimized 3D Gaussian model. The dense point cloud generation module 205 is used to perform adaptive dense sampling on the optimized 3D Gaussian model to generate a dense 3D point cloud of plants.
[0080] Specifically, the plant dense 3D point cloud reconstruction system provided by the present invention can realize all the method steps implemented in the above-mentioned plant dense 3D point cloud reconstruction method embodiment, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0081] Figure 3 This is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as... Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can call logic instructions in the memory 330 to execute a method for reconstructing dense 3D point clouds of plants. This method includes: acquiring a multi-view image sequence, wherein the multi-view image sequence is a sequence of plant videos or images acquired by an image acquisition device; performing image exposure compensation preprocessing based on the multi-view image sequence to obtain an exposure-compensated image sequence, wherein the image exposure compensation preprocessing employs a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequence to a unified target range; reconstructing sparse point clouds based on the exposure-compensated image sequence to obtain sparse point clouds and camera trajectories; performing geometrically constrained 3D Gaussian sputtering rendering based on the sparse point clouds and camera trajectories to obtain an optimized 3D Gaussian model; and performing adaptive densification sampling on the optimized 3D Gaussian model to generate a dense 3D point cloud of plants.
[0082] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0083] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the plant dense 3D point cloud reconstruction method provided by the above methods. The method includes: acquiring a multi-view image sequence, wherein the multi-view image sequence is a sequence of plant videos or images acquired by an image acquisition device; performing image exposure compensation preprocessing based on the multi-view image sequence to obtain an exposure-compensated image sequence, wherein the image exposure compensation preprocessing adopts a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequence to a unified target range; performing sparse point cloud reconstruction based on the exposure-compensated image sequence to obtain a sparse point cloud and a camera trajectory; performing geometrically constrained 3D Gaussian sputtering rendering based on the sparse point cloud and the camera trajectory to obtain an optimized 3D Gaussian model; and performing adaptive densification sampling on the optimized 3D Gaussian model to generate a plant dense 3D point cloud.
[0084] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the plant dense 3D point cloud reconstruction method provided by the above methods. This method includes: acquiring a multi-view image sequence, wherein the multi-view image sequence is a sequence of plant videos or images acquired by an image acquisition device; performing image exposure compensation preprocessing based on the multi-view image sequence to obtain an exposure-compensated image sequence, wherein the image exposure compensation preprocessing employs a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequence to a unified target range; reconstructing a sparse point cloud based on the exposure-compensated image sequence to obtain a sparse point cloud and a camera trajectory; performing geometrically constrained 3D Gaussian sputtering rendering based on the sparse point cloud and the camera trajectory to obtain an optimized 3D Gaussian model; and performing adaptive densification sampling on the optimized 3D Gaussian model to generate a plant dense 3D point cloud.
[0085] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0087] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for reconstructing dense 3D point clouds of plants, characterized in that, include: Acquire multi-view image sequences, wherein the multi-view image sequences are plant videos or image sequences acquired by an image acquisition device; Based on the multi-view image sequence, image exposure compensation preprocessing is performed to obtain an exposure-compensated image sequence. The image exposure compensation preprocessing adopts a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequence to a unified target range. Sparse point cloud reconstruction is performed on the image sequence after exposure compensation to obtain sparse point cloud and camera trajectory. Based on the sparse point cloud and the camera trajectory, geometrically constrained 3D Gaussian sputtering rendering is performed to obtain an optimized 3D Gaussian model. Adaptive dense sampling is performed on the optimized 3D Gaussian model to generate a dense 3D point cloud of plants.
2. The method for reconstructing dense 3D point clouds of plants according to claim 1, characterized in that, The sparse point cloud reconstruction based on the exposure-compensated image sequence to obtain the sparse point cloud and camera trajectory includes: The exposure-compensated image sequence is input into a 3D reconstruction network, and multi-view image features are extracted by an encoder. The multi-view image features are input into the fusion transformer module, and cross-view feature fusion is performed through self-attention mechanism and cross-attention mechanism to obtain the fused features. The fused features are simultaneously input into the global and local branches of the parallel connection; The global branch is used to perform global geometric structure understanding and camera pose estimation on the fused features to obtain global feature information; Local feature information is obtained by performing local feature matching and depth information prediction on the fused features through the local branches; Based on the global feature information and the local feature information, sparse point cloud reconstruction is performed to obtain sparse point cloud and camera trajectory.
3. The method for reconstructing dense 3D point clouds of plants according to claim 1, characterized in that, The 3D Gaussian sputtering rendering based on the sparse point cloud and the camera trajectory, with geometric constraints, yields an optimized 3D Gaussian model, including: A 3D Gaussian model is initialized based on the sparse point cloud, wherein the 3D Gaussian model is represented by multiple Gaussian elements, each of which includes a center position, a covariance matrix, opacity, and color attributes. Based on the camera trajectory, the 3D Gaussian model is rendered in a differentiable manner to obtain the predicted color map and predicted depth map of each view. Based on the predicted color map and the corresponding real image, calculate the color rendering loss of the 3D Gaussian model; Based on the predicted depth map, single-view geometric regularization is performed, and single-view geometric consistency loss is calculated. The single-view geometric regularization includes depth smoothing constraints and surface normal vector consistency constraints. Based on the predicted depth map and the camera trajectory, multi-view geometric regularization is performed, and multi-view geometric consistency loss is calculated. The multi-view geometric regularization includes depth consistency constraint, epipolar geometry constraint and cross-view normal vector consistency constraint. A total loss function is constructed by combining the color rendering loss of the 3D Gaussian model, the single-view geometric consistency loss, and the multi-view geometric consistency loss. By minimizing the total loss function, the parameters of the 3D Gaussian model are iteratively optimized to obtain the optimized 3D Gaussian model.
4. The method for reconstructing dense 3D point clouds of plants according to claim 3, characterized in that, The depth smoothing constraint is expressed by the following formula: in, This indicates a depth smoothing constraint. Represents the image domain. Indicates pixel electricity The set of neighboring pixels, and Representing pixels With pixels The depth value, Represents the weighting function; The surface normal vector consistency constraint is expressed by the following formula: in, and These represent the pixels on the predicted depth map of a single view. With pixels The surface normal vector at that location.
5. The method for reconstructing dense 3D point clouds of plants according to claim 3, characterized in that, The depth consistency constraint is expressed by the following formula: in, Represents the collection of all view pairs. For view pairs The set of corresponding pixels between them This indicates that the same 3D point is in the view. and view The pixel projection position in the image. Indicates in view medium pixel The predicted depth value, Indicates from view Reprojection to view The depth value after that, It is the Huber robust loss function; The epipolar geometric constraint is expressed by the following formula: in, and They are pixels With pixels homogeneous coordinates Indicates transpose. Represents view pairs The fundamental matrix between them; The cross-view normal vector consistency constraint is expressed by the following formula: in, Indicates in view medium pixel Surface normal vector, Indicates in view medium pixel Surface normal vector, Represents a view To view The relative rotation matrix.
6. The method for reconstructing dense 3D point clouds of plants according to claim 1, characterized in that, The step of adaptively densely sampling the optimized 3D Gaussian model to generate a dense 3D point cloud of plants includes: For each Gaussian element in the optimized 3D Gaussian model, the ratio of ellipsoidal volume to shape is calculated based on the scale parameter corresponding to its covariance matrix. The number of basic sampling points is determined based on the ellipsoidal volume, and an adaptive sampling density adjustment factor is calculated based on the shape ratio. The final number of sampling points for each Gaussian unit is determined based on the basic number of sampling points and the adaptive sampling density adjustment factor. Based on Mahalanobis distance constraints, point sampling is performed in the ellipsoidal space of each Gaussian element according to the final number of sampling points to generate a dense 3D point cloud of plants.
7. A dense 3D point cloud reconstruction system for plants, characterized in that, include: The acquisition module is used to acquire multi-view image sequences, wherein the multi-view image sequences are plant videos or RGB image sequences acquired by an image acquisition device; An exposure compensation module is used to perform image exposure compensation preprocessing based on the multi-view image sequence to obtain an exposure-compensated image sequence. The image exposure compensation preprocessing adopts a hierarchical exposure control strategy to standardize the exposure state of the multi-view image sequence to a unified target range. The sparse point cloud reconstruction module is used to reconstruct sparse point clouds based on the exposure-compensated image sequence to obtain sparse point clouds and camera trajectories. The 3D Gaussian sputtering module is used for geometrically constrained 3D Gaussian sputtering rendering based on the sparse point cloud and the camera trajectory to obtain an optimized 3D Gaussian model. The dense point cloud generation module is used to perform adaptive dense sampling on the optimized 3D Gaussian model to generate a dense 3D point cloud of plants.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the plant dense 3D point cloud reconstruction method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the plant dense 3D point cloud reconstruction method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the plant dense 3D point cloud reconstruction method as described in any one of claims 1 to 6.