A method for three-dimensional imaging of rice from multiple perspectives, balancing processing time and reconstruction accuracy.

A rice 3D imaging method based on multi-gradient segmentation and residual driving mechanism solves the problem of balancing efficiency and accuracy in rice 3D imaging, achieving efficient rice 3D reconstruction, which is suitable for real-time analysis in smart agriculture.

CN121392103BActive Publication Date: 2026-04-03YAZHOUWAN NATIONAL LABORATORY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing rice 3D imaging technology struggles to balance processing efficiency and accuracy. The image segmentation process is time-consuming and prone to losing key features. Point cloud fusion does not consider occlusion and matching quality, resulting in low accuracy of the reconstructed model, which cannot meet the needs of real-time field analysis.

Method used

A multi-gradient segmentation strategy and residual-driven mechanism are adopted. A multi-gradient segmentation dataset is constructed by acquiring RGB images from multiple angles. The improved Faster-RCNN model is used to output three masks in parallel. Combined with sparse-dense 3D reconstruction and reinforcement learning strategies, point clouds are aligned and fused, and finally a 3D model of rice is constructed.

Benefits of technology

It achieves a balance between segmentation accuracy and processing speed, improving segmentation inference speed by 30%, reducing point cloud fusion time by 20%-30%, and achieving a geometric error of ≤0.5mm for the fused model, thus meeting the real-time phenotypic analysis needs of smart agriculture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121392103B_ABST
    Figure CN121392103B_ABST
Patent Text Reader

Abstract

This invention discloses a method for 3D imaging of rice plants from multiple perspectives, balancing processing time and reconstruction accuracy. Belonging to the fields of agricultural image processing and computer vision, the method includes: acquiring and standardizing RGB images of individual rice plants from multiple angles to obtain a standardized image dataset; constructing a multi-gradient segmentation dataset and training a segmentation network with multi-gradient output capabilities to obtain a multi-gradient segmentation model; using the multi-gradient segmentation model to infer from the standardized image dataset, obtaining a set of differentiated feature points, and then performing sparse-dense 3D reconstruction to obtain three sets of point clouds corresponding to three different masks; aligning, fusing, and correcting the three sets of point clouds based on a residual-driven mechanism to obtain a fused point cloud; and constructing a 3D model of rice plants based on the fused point cloud and outputting the 3D imaging result. This invention achieves rapid construction and accuracy optimization of 3D models of rice plants by integrating deep learning segmentation and 3D reconstruction techniques.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of agricultural image processing and computer vision technology, and in particular relates to a method for three-dimensional imaging of rice multi-view images that balances processing time and reconstruction accuracy. Background Technology

[0002] As one of the world's most important food crops, rice's yield and quality directly impact food security. With the development of precision agriculture and intelligent breeding, higher demands are being placed on the accurate and efficient acquisition of rice phenotypic information. Three-dimensional quantitative analysis of rice phenotypic characteristics can comprehensively reflect the plant's growth status, providing crucial data support for the selection of superior varieties, monitoring of growth dynamics, and yield prediction in crop breeding. Specifically, three-dimensional imaging technology, by reconstructing the three-dimensional structure of the rice plant, lays a solid foundation for the accurate extraction of important phenotypic parameters such as plant height, leaf area index, tiller number, and leaf angle.

[0003] In recent years, significant progress has been made in computer vision-based 3D reconstruction technology. Among these, 3D reconstruction techniques based on multi-view RGB images, such as Structure of Motion (SfM) and Multi-View Stereo Matching (MVS), have been widely used in agricultural scenarios due to their advantages of low cost, ease of operation, and easy deployment. This technology can reconstruct the 3D structure of an object by processing multiple 2D images of the same object taken from different perspectives, without relying on expensive 3D scanning equipment, thus greatly lowering the barrier to entry for 3D imaging in agriculture.

[0004] In the field of 3D imaging of rice, existing technologies still face many challenges. Firstly, in the image segmentation stage, traditional methods mostly employ a single-precision segmentation mode. While fine segmentation can accurately preserve the detailed features of rice plants, providing high-quality input for subsequent 3D reconstruction, it requires processing a large amount of pixel information, resulting in a long segmentation process and severely impacting overall processing efficiency. On the other hand, coarse segmentation, although it can shorten processing time to some extent, easily loses key features such as rice leaf edges and tillers, leading to information gaps in the reconstructed 3D model and failing to meet the requirements of high-precision phenotypic analysis. Therefore, balancing segmentation accuracy and processing speed has become a significant bottleneck restricting the efficiency of 3D imaging of rice. During 3D reconstruction, rice plants are characterized by slender, numerous, and heavily occluded leaves, which greatly complicates point cloud fusion. Traditional point cloud fusion strategies mostly employ a fixed-weight approach, assigning the same weight to point clouds obtained from different viewpoints or with different segmentation accuracies for fusion. This method does not take into account the differences in occlusion degree, matching quality, and feature importance in different regions. This not only leads to a large amount of redundant computation and increases fusion time, but also easily causes the accumulation of errors, resulting in low accuracy of the final reconstructed 3D model.

[0005] Existing 3D modeling methods also have significant shortcomings when dealing with complex rice structures. For example, the mainstream Poisson reconstruction method often suffers from surface discontinuities and insufficient detail restoration when faced with the delicate structure and complex spatial distribution of rice leaves, making it difficult to realistically reflect the morphological characteristics of the rice plant. Furthermore, the entire 3D imaging process is time-consuming, failing to meet the practical needs of real-time monitoring and rapid analysis of rice phenotypes in smart agriculture, thus limiting its widespread application in the field. In conclusion, developing a 3D rice imaging method that balances processing efficiency and reconstruction accuracy, and optimizing the overall processing time through innovative segmentation strategies and fusion mechanisms, is of significant practical importance for promoting the development of smart agriculture. Summary of the Invention

[0006] To address the shortcomings of traditional single-segmentation methods in balancing the accuracy and processing speed of rice image segmentation, which leads to redundancy or missing information in 3D reconstruction input and affects subsequent efficiency; the use of fixed weight strategies in point cloud fusion during 3D reconstruction without considering the degree of occlusion and matching quality in different regions, resulting in long fusion times and large errors; and the technical problems of insufficient model detail and excessively long overall processing time when dealing with the complex structure of rice, making it difficult to meet the needs of real-time field analysis, this invention provides a multi-view 3D imaging method for rice that balances processing time and reconstruction accuracy, including:

[0007] Multi-angle RGB image acquisition and standardized preprocessing were performed on individual rice plants to obtain a standardized image dataset;

[0008] A multi-gradient segmentation dataset is constructed based on the standardized image dataset, and the multi-gradient segmentation dataset includes at least three segmentation masks with different background expansion ratios;

[0009] Based on the multi-gradient segmentation dataset, a segmentation network with multi-gradient output capability is trained to obtain a multi-gradient segmentation model that can output three masks in parallel.

[0010] The multi-gradient segmentation model is used to infer the standardized image dataset to obtain three corresponding masks and extract the set of differential feature points.

[0011] Based on the differentiated feature point set, sparse-dense 3D reconstruction is performed respectively to obtain three sets of point clouds corresponding to the three masks;

[0012] Based on the residual-driven mechanism, the three sets of point clouds are aligned, fused, and error-corrected to obtain a fused point cloud.

[0013] A three-dimensional model of rice is constructed based on the fused point cloud, and the three-dimensional imaging results are output.

[0014] Preferably, the process of performing multi-angle RGB image acquisition and standardization preprocessing on individual rice plants to obtain a standardized image dataset includes:

[0015] Centered on a single rice plant, a multi-angle shooting strategy was adopted, taking one image every 10 degrees horizontally and three layers vertically. The original image was obtained by combining an array camera with a turntable for rotating shooting.

[0016] Based on the original image, data cleaning, lens distortion correction, and color calibration are performed to obtain a standardized image dataset.

[0017] Preferably, the process of constructing a multi-gradient segmentation dataset based on the standardized image dataset includes:

[0018] The standardized image dataset is segmented and labeled at the pixel level to form g1 complete segmentation mask, g2 partial background mask and g3 large background mask, and the mask confidence, occlusion level, growth period label and variety number are recorded.

[0019] Preferably, the segmentation network with multi-gradient output capability is an improved Faster-RCNN, and the model structure includes:

[0020] The shared backbone network, RPN, RoI Head, and three gradient segmentation heads connected in parallel with the RoI Head;

[0021] The segmentation head dynamically selects the corresponding gradient mask for output based on the confidence level of the proposal output by the RPN.

[0022] Preferably, the process of training a segmentation network with multi-gradient output capability based on the multi-gradient segmentation dataset includes:

[0023] By introducing a processing time penalty term, Focal Loss, and a dynamic weight adjustment mechanism, and employing transfer learning initialization, mixed precision training, and cosine annealing learning rate scheduling, the multi-gradient segmentation model is obtained.

[0024] Preferably, the extraction process of the differentiated feature point set includes:

[0025] Morphological optimization, GrabCut post-processing, and edge sub-pixel thinning are performed on the three types of masks respectively. Then, high-response feature points in the RGB region are extracted according to the retention ratio to form a set of differentiated feature points corresponding to the g1 complete segmentation mask, the g2 partial background mask, and the g3 large background mask.

[0026] Preferably, the process of performing sparse-dense 3D reconstruction according to the differentiated feature point set to obtain three sets of point clouds corresponding to the three masks includes:

[0027] The differential feature point set was sequentially reconstructed using SFM sparse reconstruction and MVS dense reconstruction using COLMAP to obtain point clouds g1, g2 and g3 respectively.

[0028] Preferably, the process of aligning, fusing, and error-correcting the three sets of point clouds based on a residual-driven mechanism to obtain a fused point cloud includes:

[0029] The residual matrices of the three sets of point clouds are calculated, and the point cloud overlap and the mean of the residuals are used as the state space. Fusion weights are dynamically generated through a reinforcement learning strategy, and ICP registration and weighted fusion are performed on the three sets of point clouds to obtain the fused point cloud.

[0030] Preferably, the process of constructing a three-dimensional model of rice based on the fused point cloud includes:

[0031] Based on the fused point cloud and camera parameters, the 3D Gaussian Splatting method is used to assign position, covariance, color and transparency attributes to each three-dimensional point, and a three-dimensional rice model in PLY or STL format is generated through optimization iteration.

[0032] Compared with the prior art, the present invention has the following advantages and technical effects:

[0033] This invention balances segmentation accuracy and processing efficiency through a multi-gradient segmentation strategy—g1 ensures accuracy in key regions, while g2 and g3 reduce matching failures by expanding the background, thus solving the problem of "accuracy and speed cannot be achieved simultaneously" in a single segmentation mode.

[0034] The improved Faster-RCNN model of this invention introduces a processing time penalty term and dynamic weight adjustment, which improves the segmentation inference speed by more than 30%, while achieving a mask accuracy (mAP) of more than 92%.

[0035] This invention combines residual fusion and reinforcement learning strategies to reduce point cloud fusion time by 20%-30%, and the geometric error (mean Euclidean distance) of the fused model is ≤0.5mm;

[0036] This invention is based on a 3D Gaussian Splatting modeling method and system-level optimization (GPU acceleration, model lightweighting) to meet the needs of real-time phenotypic analysis in smart agriculture fields. Attached Figure Description

[0037] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0038] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention;

[0039] Figure 2 This is a structural diagram of the improved Faster-RCNN model according to an embodiment of the present invention. Detailed Implementation

[0040] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0041] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0042] like Figure 1 As shown, this embodiment provides a method for three-dimensional imaging of rice multi-view images that balances processing time and reconstruction accuracy, including:

[0043] Multi-angle RGB image acquisition and standardized preprocessing were performed on individual rice plants to obtain a standardized image dataset;

[0044] A multi-gradient segmentation dataset is constructed based on a standardized image dataset. The multi-gradient segmentation dataset includes at least three segmentation masks with different background expansion ratios.

[0045] A segmentation network with multi-gradient output capability is trained based on a multi-gradient segmentation dataset to obtain a multi-gradient segmentation model that can output three masks in parallel.

[0046] The multi-gradient segmentation model is used to infer the standardized image dataset, obtain the corresponding three masks, and extract the differential feature point set;

[0047] Based on the differential feature point set, sparse-dense 3D reconstruction is performed respectively to obtain three sets of point clouds corresponding to the three masks;

[0048] Based on the residual-driven mechanism, the three sets of point clouds are aligned, fused and error-corrected to obtain a fused point cloud;

[0049] A 3D model of rice is constructed based on the fused point cloud, and the 3D imaging results are output.

[0050] This embodiment uses multi-angle RGB images of rice as the sole data source. First, a standardized image dataset is constructed using a rotational shooting strategy. Second, experts with agronomic background complete pixel-level precise annotation of three gradients (g1 complete segmentation, g2 partial background, g3 large background preservation) on an interactive annotation platform, forming a multi-gradient segmentation dataset. Subsequently, an improved Faster-RCNN network is used to train the multi-gradient segmentation model, outputting three complementary masks. Then, a two-stage process of COLMAP + 3DGaussian Splatting is employed to complete sparse-dense 3D reconstruction. Finally, a residual-driven weight adaptive mechanism is used to align, fuse, and correct the three sets of point clouds, outputting a high-precision 3D rice model. This embodiment achieves a dynamic balance between speed and accuracy without adding any extra sensors or expensive hardware, making it suitable for multiple scenarios such as high-throughput phenotypic analysis in smart agriculture, precision breeding, and agricultural machinery automation, possessing both scientific research and industrialization value.

[0051] Furthermore, the process of acquiring and standardizing RGB images of individual rice plants from multiple angles to obtain a standardized image dataset includes:

[0052] Centered on a single rice plant, a multi-angle shooting strategy was adopted, taking one image every 10 degrees horizontally and three layers vertically. The original image was obtained by combining an array camera with a turntable for rotating shooting.

[0053] Based on the original images, data cleaning, lens distortion correction, and color calibration are performed to obtain a standardized image dataset.

[0054] Furthermore, the multi-angle RGB image acquisition and standardized preprocessing process of rice in this embodiment includes: using a 12-array camera combined with a turntable rotation shooting strategy, taking a single rice plant as the center, acquiring one image every 10° in a horizontal 360° range (36 images in total), and acquiring images in three layers in the vertical direction: bottom, middle, and top (432 images in total); performing data cleaning (screening and removing blurry, overexposed / underexposed images based on Sobel sharpness, brightness histogram exposure, texture entropy, and background complexity), lens distortion correction (based on camera intrinsic parameter matrix), and color calibration (unifying white balance and color gamut) on the original images to generate a standardized image dataset; the images are taken at a preset resolution that meets the accuracy of 3D reconstruction;

[0055] The rotating shooting strategy was implemented as follows: 12 array-type high-resolution RGB cameras were used to photograph potted rice plants located on a rotating platform. During shooting, automatic exposure and white balance were locked (exposure time 1 / 1000-1 / 500s), ISO was set to 200, and images were stored in JPEG lossless compression format. Image selection was based on a joint score of three indicators: Sobel sharpness (≥0.8), brightness histogram exposure (mean 50-200), and texture entropy background complexity (≥3.0), retaining the top 90% of the images.

[0056] Specifically, the imaging unit consists of a 12-camera array of Canon RGB cameras, each with a resolution of 6000×4000. An array arrangement combined with a rotating camera system enables automatic shooting: the 12 cameras are positioned around a rotating platform in a pre-defined pattern (4 at the top, 6 in the middle, and 2 at the bottom), with the potted rice plant placed in the center. Horizontally, the rotating platform rotates uniformly around the potted rice plant, triggering a shot every 10° of rotation. With the orderly operation of these 12 cameras, each camera captures 36 images per rice plant, comprehensively covering a 360° horizontal field of view, for a total of 432 images. Vertically, considering the growth characteristics of the potted rice and the layered distribution of the cameras, the top 4 cameras correspond to the 60-110cm height range, the middle 6 cameras to the 20-60cm height range, and the bottom 2 cameras to the base to the 20cm height range. This invention collected images of 60 rice varieties, with three replicates for each variety, totaling 180 potted plants. These images covered different growth stages (seedling, tillering, jointing, heading, and maturity) to ensure the diversity and representativeness of the dataset. The shooting distance was optimized based on the potted plant dimensions (30 cm height, 25 cm diameter) and the maximum plant height (1.1 m). The distance between the lenses of the 12 array cameras and the rice plants was 100 cm, ensuring clear capture of detailed features while fully covering the entire plant area. Shooting was conducted once a week, for a total of 18 times. Automatic exposure (1 / 1000-1 / 500s) and white balance were locked during shooting, with ISO set to 200. Images were stored in lossless JPEG format. In the preprocessing stage, the Sobel operator is used to calculate image sharpness (threshold ≥ 0.8), the exposure is analyzed by brightness histogram (mean 50-200), and the background complexity is evaluated based on texture entropy (threshold ≥ 3.0). The top 90% of high-quality images are selected to ensure input quality.

[0057] Furthermore, the process of constructing a multi-gradient segmentation dataset based on a standardized image dataset includes:

[0058] Pixel-level segmentation and annotation are performed on the standardized image dataset to form g1 complete segmentation mask, g2 partial background mask and g3 large background mask, and the mask confidence, occlusion level, growth period label and variety number are recorded.

[0059] Furthermore, the multi-gradient segmentation dataset construction process in this embodiment includes: on a standardized image, three gradient segmentation labels are applied to rice plants using a labeling tool: g1 (complete segmentation, containing only rice plant pixels), g2 (expanding by 5% background pixels based on g1), and g3 (expanding by 15% background pixels based on g1); simultaneously, the mask confidence level (0-100%), leaf occlusion level (1-5 levels, the higher the level, the more severe the occlusion), growth period label, variety number, and timestamp are recorded to form a labeled training dataset; the three gradient segmentations, by differentiating the background range, lay a data foundation for reducing invalid calculations in the subsequent process;

[0060] Specifically, using the Labelme annotation tool, 300 rice datasets at different growth stages were randomly selected. The annotators segmented the standardized images according to agronomic standards: g1 strictly selected rice plant pixels (without background), g2 expanded by 5% pixels outside the g1 boundary, and g3 expanded by 15% background pixels. The mask confidence (e.g., g1 confidence ≥ 95%, g3 confidence ≥ 80%), leaf occlusion level (level 1 no occlusion, level 5 severe occlusion), growth stage (tillering stage, heading stage, etc.) and variety number were recorded simultaneously. The annotation format followed the COCOpanoptic standard, and finally, a training dataset of 300 rice plants was generated (432 annotated images for each plant).

[0061] Furthermore, the segmentation network with multi-gradient output capability is an improved Faster-RCNN, and its model structure includes:

[0062] Shared backbone network, RPN, RoI Head, and three gradient segmentation heads connected in parallel with RoI Head;

[0063] The segmentation head dynamically selects the corresponding gradient mask for output based on the confidence level of the proposal output by the RPN.

[0064] Furthermore, such as Figure 2As shown, the training process of the improved Faster-RCNN model in this embodiment includes: adding three parallel multi-gradient segmentation heads (with ResNet-50 backbone, RPN, and RoI Head) to the traditional Faster-RCNN framework (including ResNet-50 backbone, RPN, and RoI Head). The head shares a feature map, and the output dimension is H×W×1, where H and W are the image height and width. After generating proposals through RPN, the background expansion ratio is dynamically adjusted based on the confidence of the proposal (g1 is used when the confidence is ≥80%, g2 is used when the confidence is 60%-80%, and g3 is used when the confidence is <60%), realizing parallel output of three masks: g1, g2, and g3. During training, a processing time penalty term (expressed as λtime・T, where T is the inference time per image (milliseconds), λtime=0.01), FocalLoss (adjusting the weights of easy and difficult samples), and a dynamic weight adjustment mechanism (dynamically allocating the proportion of classification loss and segmentation loss based on the number of iterations) are introduced. Transfer learning (initial weights are from the ImageNet pre-trained model), mixed precision (FP16+FP32), dynamic batch size (adjustable from 16-32), and cosine annealing scheduling (initial learning rate 1×10) are adopted. -3 (After 50 rounds of processing), a preliminary segmentation model M0 was obtained;

[0065] Specifically, the improved Faster-RCNN model in this embodiment is based on the traditional two-stage object detection framework. By introducing a multi-gradient segmentation head and optimizing the training strategy, it achieves a balance between accuracy and speed in rice image segmentation. The following is a detailed description of the model's specific structure, the connections between its parts, and their functions.

[0066] The model structure specifically includes:

[0067] Backbone Network (ResNet-50): ResNet-50 is used as the feature extraction network. A standardized RGB image of rice (resolution H×W×3) is input, and a high-dimensional feature map (dimensions H / 16×W / 16×2048) is generated through multiple convolutional layers and residual connections. This feature map captures the texture, edges, and semantic information of the rice plant, providing shared features for subsequent region proposal and segmentation tasks.

[0068] Region Proposal Network (RPN): Based on the feature maps of the backbone network, the RPN generates candidate regions (proposals) through 3×3 convolutions. Each proposal contains bounding box coordinates and a confidence score. The RPN outputs fixed-size proposals (128×128) and dynamically distributes them to subsequent segmentation heads based on the confidence scores.

[0069] RoI Head: Responsible for classification and bounding box regression tasks. The RoI Head performs the RoIAlign operation on the RPN proposal, extracts fixed-size features (7×7×2048), and outputs the classification results (rice plant / background) and bounding box coordinate correction values ​​through a fully connected layer.

[0070] Multi-gradient segmentation heads: Three parallel segmentation heads (g1, g2, g3) are added, each consisting of a three-layer convolutional network (3×3 kernels, stride 1, output dimension H×W×1). The segmentation heads share the feature maps of the RPN with the RoI Head, and dynamically select the output mask based on the proposal confidence: g1 (precise segmentation, containing only plant pixels) is output when the confidence is ≥80%, g2 (expanding background pixels by 5%) is output when the confidence is 60%–80%, and g3 (expanding background pixels by 15%) is output when the confidence is <60%.

[0071] The connections include:

[0072] (1) Input RGB image (H×W×3) and extract features through ResNet-50 to generate shared feature map (H / 16×W / 16×2048).

[0073] (2) Input the shared feature map into the RPN to generate multiple proposals (including bounding boxes and confidence scores).

[0074] (3) The proposal is passed to the RoI Head and the multi-gradient segmentation head. The RoI Head normalizes the proposal features to 7×7×2048 using RoI Align, which is used for classification (cross-entropy loss Lcls) and bounding box regression (SmoothL1 loss Lbox). The segmentation head dynamically selects g1, g2 or g3 segmentation heads according to the confidence threshold (≥80%, 60%−80%, <60%) to generate the corresponding mask (Dice loss Lmask).

[0075] (4) The mask (H×W×1) output by each segmentation head is used together with the classification and bounding box results of the RoI Head to optimize the model.

[0076] The functions of each part of the model include:

[0077] ResNet-50: Extracts multi-scale features through deep convolutions and residual connections, ensuring that the model can capture the fine structure of rice plants (such as leaf edges and tillers).

[0078] RPN generates high-quality candidate regions, reduces invalid computations, and improves segmentation efficiency by implementing dynamic mask allocation through confidence scoring.

[0079] RoI Head: Responsible for accurate target classification and localization, ensuring semantic consistency between segmented regions and rice plants.

[0080] Multi-gradient segmentation head: Balances accuracy and speed through three types of masks: g1, g2, and g3. g1 ensures details in high-precision areas, while g2 and g3 reduce matching failures and computational load by expanding the background.

[0081] Loss function and optimization strategy:

[0082] Loss function: L = Lcls + Lbox + λmask(Lmaskg1 + Lmaskg2 + Lmaskg3) + λtime·T

[0083] Where Lcls is the classification loss (cross-entropy loss), Lbox is the bounding box regression loss (SmoothL1 loss), λmask=1.0, λtime=0.01, and T is the inference time, with a time penalty term to optimize inference speed; dynamic weight adjustment (70% for classification and regression in the first 50 rounds, 60% for segmentation in the last 50 rounds) balances classification accuracy and segmentation speed; transfer learning (ImageNet pre-trained weights) accelerates convergence; mixed precision (FP16+FP32) and dynamic batch size (16-32) reduce memory usage; cosine annealing scheduling (initial learning rate 1×10) −3 (50 rounds per cycle) to ensure training stability.

[0084] Furthermore, the process of training a segmentation network with multi-gradient output capability based on a multi-gradient segmentation dataset includes:

[0085] By introducing a processing time penalty term, Focal Loss, and a dynamic weight adjustment mechanism, and employing transfer learning initialization, mixed precision training, and cosine annealing learning rate scheduling, a multi-gradient segmentation model is obtained.

[0086] Furthermore, the extraction process of the differentiated feature point set includes:

[0087] Morphological optimization, GrabCut post-processing, and edge sub-pixel thinning were performed on the three masks respectively. Then, high-response feature points in the RGB region were extracted according to the retention ratio to form a set of differential feature points corresponding to the g1 complete segmentation mask, the g2 partial background mask, and the g3 large background mask.

[0088] Furthermore, this embodiment performs the following operations for three different masks:

[0089] Morphological optimization (using 3×3 rectangular structuring elements for opening operations to remove noise);

[0090] GrabCut post-processing (5-8 iterations, using mask boundaries as seed points to optimize foreground / background segmentation);

[0091] Edge subpixel refinement (using Zernike moment algorithm, accuracy ≤ 0.5 pixels); extract RGB information of corresponding regions and construct differential feature point sets (g1 retains 90% of feature points, g2 retains 70%, g3 retains 50%), reducing the amount of subsequent computation by reducing the number of low-value feature points;

[0092] Specifically, this embodiment performs post-processing on the three masks output by model M0: First, an opening operation is performed using a 3×3 rectangular structuring element to remove noise; then, the GrabCut algorithm is run iteratively 6 times with the mask boundary as the seed point to optimize the foreground / background segmentation; finally, the Zernike moment algorithm is used to refine the edges to a sub-pixel level, so that the edge positioning accuracy reaches 0.3 pixels.

[0093] After extracting the RGB information of each mask region, the ORB algorithm is used to construct a feature point set: First, the input image is segmented at the pixel level using predefined masks g1, g2, and g3 (corresponding to regions with confidence levels ≥80%, 60%-80%, and <60%, respectively), and the RGB values ​​of the corresponding pixels in each mask region are extracted. The specific steps include: (1) filtering the target region by mask binarization, (2) traversing the pixels of the selected region and recording their RGB three-channel values. (3) Subsequently, the ORB algorithm is applied to detect feature points. Based on the response value, g1 retains 90% of the high-response feature points (for high-precision reconstruction), g2 retains 70% (balancing accuracy and quantity), and g3 retains 50% (reducing redundant calculations). The coordinates and descriptors of all feature points are recorded.

[0094] Furthermore, based on the differentiated feature point sets, the process of performing sparse-dense 3D reconstruction to obtain three sets of point clouds corresponding to the three masks includes:

[0095] COLMAP was used to perform SFM sparse reconstruction and MVS dense reconstruction on the differential feature point set in sequence to obtain g1 point cloud, g2 point cloud and g3 point cloud respectively.

[0096] Furthermore, the 3D reconstruction and residual fusion process in this embodiment includes: using COLMAP to estimate sparse point clouds and camera parameters (intrinsic and extrinsic); performing SFM (Structure of Motion Recovery) and MVS (Multi-View Stereo Matching) on ​​the three sets of gradient images g1, g2, and g3 respectively to generate three sets of point cloud models; and using the ICP algorithm (iterations 20-30 times, convergence threshold 1×10⁻⁶). -4Complete coarse and fine registration; define a residual matrix R (N×3 dimension, where N is the number of point clouds and the elements are the 3D coordinate residuals of the corresponding points), and dynamically optimize the fusion weights based on a reinforcement learning strategy (the state space consists of the point cloud overlap and the mean residual, and the action space consists of the fusion weight coefficients). The weight calculation formula is as follows:

[0097] ;

[0098] Where τ=0.02, Let γ be the 3D coordinate residual of the corresponding point in the i-th point cloud, γ=2, O i The overlap between the i-th group of point clouds and other groups is used to reduce redundant point cloud computing through this fusion strategy, thereby achieving point cloud alignment, fusion and error correction.

[0099] The reinforcement learning strategy employs the PPO algorithm, and the reward function is defined as follows:

[0100] ;

[0101] In the formula, The geometric error of the point cloud is calculated using the average Euclidean distance; The current fusion step takes time; by updating the policy network in each iteration, the fusion weights are optimized towards "low error + short time", reducing the fusion time by 20%-30% compared to the fixed weight policy.

[0102] Specifically, in this embodiment, the preprocessed image (output of the improved Faster-RCNN) is input into COLMAP, feature points are extracted and matched using the SIFT algorithm, and camera intrinsic parameters (focal length, distortion coefficients) and extrinsic parameters (rotation matrix, translation vector) are estimated through bundle adjustment (BA) to generate a sparse point cloud (approximately 5000 points).

[0103] Perform SFM and MVS on the feature point sets g1, g2, and g3 respectively:

[0104] (1) SFM stage: First, based on the previously extracted ORB feature points, feature matching is performed between multi-view images, and mismatched points are removed using the RANSAC algorithm. Then, the 3D spatial coordinates are calculated through triangulation to generate the initial sparse point cloud. Next, the camera intrinsic parameters (focal length, distortion coefficients) and extrinsic parameters (rotation matrix, translation vector) are optimized by combining bundle adjustment (BA) to adjust the position of the point cloud and obtain the initial point clouds of g1, g2, and g3.

[0105] (2) MVS stage: Dense matching is performed using the PatchMatch algorithm. The specific process includes: (1) initializing depth and normal estimation, (2) updating the depth value of pixels and matching cost through iterative optimization, (3) generating a depth map and fusing multi-view depth information, and finally generating a dense point cloud through projection and filtering. g1 retains high-confidence areas (about 100,000 points), g2 balances accuracy and density (about 150,000 points), and g3 covers a larger area but contains more noise (about 200,000 points). The difference in the number of point clouds reflects the influence of the feature point retention ratio (90%, 70%, 50%) and matching complexity.

[0106] The following three algorithms are logically related to each other to achieve high-quality point cloud fusion: point cloud registration, fusion weight calculation and optimization.

[0107] (1) ICP point cloud registration:

[0108] Point cloud alignment is performed using the Iterative Closest Point (ICP) algorithm, which consists of two steps:

[0109] Coarse registration: Initial registration is performed using the Random Sample Consensus (RANSAC) algorithm, with 100 iterations. The initial transformation matrix is ​​estimated based on the correspondence of feature points, and outliers are removed to obtain a coarsely aligned point cloud.

[0110] Fine-grained registration: Based on the coarse-grained registration results, perform ICP iterative optimization, setting 25 iterations and a convergence threshold of 1×10⁻⁶. -4 By minimizing the distance between points, the precise rotation matrix and translation vector are calculated, and finally, aligned point clouds g1, g2, and g3 (approximately 100,000, 150,000, and 200,000 points respectively) are generated, providing consistent spatial coordinates for subsequent fusion.

[0111] Fusion weight calculation:

[0112] Based on the aligned point cloud, the fusion weights g1, g2, and g3 are calculated. The process includes:

[0113] Use a KD-tree to calculate the overlap between each point cloud group and the other two groups. , is defined as the percentage of neighboring points (the number of overlapping points is counted with a neighboring radius of 0.1 meters).

[0114] Combining parameters τ=0.02 and γ=2, the weighted formula is used. ( Assign weights to the point cloud density differences, where w1, w2, and w3 represent the contributions of g1, g2, and g3, respectively.

[0115] Normalize the weights to ensure w1+w2+w3=1. The result is the weight value for each point cloud group, which is used to guide the contribution ratio of subsequent fusion.

[0116] (3) PPO reinforcement learning optimization:

[0117] The Proximal Policy Optimization (PPO) algorithm is used to dynamically adjust the fusion weights, with the optimization goal of "low error + short time consumption".

[0118] State space: defined as (mean overlap, mean residual). The mean overlap is calculated by using a KD tree to determine the average overlap of the three point clouds. The mean residual is calculated by the average Euclidean distance between the point cloud after ICP registration and the reference point cloud.

[0119] Action space: (w1, w2, w3), representing the adjustment range of the fusion weights (0, 1) and the sum of the weights is 1;

[0120] Reward function: r = -( +0.1 ),in To fuse the average Euclidean distance between point clouds and measured values, The fusion process takes time;

[0121] In each iteration, the policy network is updated, and the weights are optimized through gradient ascent to maximize r. The final result is an optimized weight combination (w1*, w2*, w3*), which generates a fused point cloud (approximately 450,000 points), balancing geometric accuracy and computational efficiency.

[0122] The specific logical connection is as follows: ICP registration first aligns the g1, g2, and g3 point clouds spatially to generate a consistent coordinate system; the fusion weight calculation determines the contribution of each point cloud based on the alignment result; the PPO algorithm uses the weights as actions, combined with the registration quality and time consumption feedback, to iteratively optimize the weights and finally generate a high-quality fused point cloud.

[0123] The process of aligning, fusing, and correcting errors in three sets of point clouds based on a residual-driven mechanism to obtain a fused point cloud includes:

[0124] The residual matrices of the three sets of point clouds are calculated, and the point cloud overlap and the mean of the residuals are used as the state space. The fusion weights are dynamically generated through a reinforcement learning strategy. ICP registration and weighted fusion are performed on the three sets of point clouds to obtain the fused point cloud.

[0125] The process of constructing a 3D model of rice based on the fused point cloud includes:

[0126] Based on the fused point cloud and camera parameters, the 3D Gaussian Splatting method is used to assign position, covariance, color and transparency attributes to each three-dimensional point, and a three-dimensional rice model in PLY or STL format is generated through optimization iteration.

[0127] The construction process of the 3D Gaussian Splatting model in this embodiment includes: based on the fused dense point cloud and camera parameters, a 3D model is constructed using the 3DGaussian Splatting method (assigning position, covariance, color, and transparency attributes to each 3D point, iterating 30,000 times using the Adam optimizer with a learning rate of 1×10⁻⁶). -3 (SSIM loss weight 0.2); Generate general format models (such as PLY, STL) through Poisson reconstruction or triangulation techniques, compatible with mainstream agricultural analysis software (such as Agisoft Metashape, PlantEye);

[0128] The fused point cloud (approximately 250,000 points) is a subset obtained by fusing the weighted combination (w1*, w2*, w3*) optimized by the Proximal Policy Optimization (PPO) algorithm with the g1, g2, and g3 point clouds (refer to the aforementioned PPO process), and then processing it through filtering and downsampling. This subset is then input into the 3DGaussianSplatting module. The construction process is as follows:

[0129] Initialization: Initialize each point cloud point as a 3D Gaussian distribution with parameters including position (x, y, z), symmetric covariance matrix Σ (initially the identity matrix), color RGB (inherited from point cloud attributes), and transparency α (initially 1).

[0130] Optimization: The Adam optimizer was used for 30,000 iterations, with a learning rate of 1×10⁻⁶. -3 By combining the SSIM loss (weight 0.2) and L1 projection error, the Gaussian parameters are adjusted to maximize the structural similarity (SSIM) between the rendered image and the input image, while minimizing pixel-level differences.

[0131] Rendering and Verification: A Gaussian distribution is projected onto the 2D image plane using differential rendering technology. The error between the Gaussian distribution and the real image is calculated and fed back to the optimizer. After optimization, the dense point representation of the Gaussian distribution is extracted.

[0132] Subsequently, a triangular mesh model with approximately 500,000 vertices was generated based on the optimized Gaussian distribution point set using the Poisson reconstruction algorithm. After surface reconstruction and smoothing, the model was exported in PLY and STL formats, which can be directly imported into agricultural analysis software such as Agisoft Metashape and PlantEye for further analysis.

[0133] The method in this embodiment also includes system integration and performance optimization: constructing four major modules: "image acquisition - deep learning processing - 3D reconstruction - fusion evaluation", and optimizing processing time through the following methods:

[0134] The modules adopt a zero-copy memory sharing mechanism (to reduce data transfer time);

[0135] A multi-threaded scheduling strategy (dynamic matching of CPU thread count and GPU core count) is designed for the 3D reconstruction and model training steps.

[0136] It employs GPU acceleration (CUDA core parallel computing) and TensorRT inference optimization (model quantization to INT8); supports edge computing (adapted to NVIDIA Jetson series chips) and cloud deployment (distributed model training), and updates the algorithm via OTA upgrades; ultimately realizing the transformation of a single rice plant from image input to 3D model output.

[0137] The system is deployed using Docker containers, and the four main modules share data through zero-copy memory (reducing data transfer time by 90%).

[0138] The multi-threaded scheduling strategy is as follows: image preprocessing is allocated to 2 CPU threads, model inference uses 4 GPU cores, 3D reconstruction is allocated to 8 GPU cores (CUDA parallel computing), and fusion evaluation uses 2 CPU threads. The model is quantized to INT8 using TensorRT, resulting in a 2x speedup for inference.

[0139] The performance optimization is aimed at edge computing scenarios: lightweight model processing (pruning rate of 30%-40%) is adopted, and the dense matching steps in 3D reconstruction are allocated to the cloud GPU cluster. Only image preprocessing and model inference are performed at the edge. The total memory usage of a single rice plant is ≤2GB and the peak power consumption is ≤30W.

[0140] This embodiment specifically presents a three-dimensional imaging method for rice from multiple perspectives that balances processing time and reconstruction accuracy. By integrating deep learning segmentation and three-dimensional reconstruction techniques, it enables the rapid construction and accuracy optimization of three-dimensional models of rice plants, making it suitable for efficient analysis of rice phenotypes in smart agriculture.

[0141] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for three-dimensional imaging of rice from multiple perspectives, balancing processing time and reconstruction accuracy, characterized in that: include: Multi-angle RGB image acquisition and standardized preprocessing were performed on individual rice plants to obtain a standardized image dataset; A multi-gradient segmentation dataset is constructed based on the standardized image dataset, and the multi-gradient segmentation dataset includes at least three segmentation masks with different background expansion ratios; Based on the multi-gradient segmentation dataset, a segmentation network with multi-gradient output capability is trained to obtain a multi-gradient segmentation model that can output three masks in parallel. The multi-gradient segmentation model is used to infer the standardized image dataset to obtain three corresponding masks and extract the set of differential feature points. Based on the differentiated feature point set, sparse-dense 3D reconstruction is performed respectively to obtain three sets of point clouds corresponding to the three masks; Based on the residual-driven mechanism, the three sets of point clouds are aligned, fused, and error-corrected to obtain a fused point cloud. A three-dimensional model of rice is constructed based on the fused point cloud, and the three-dimensional imaging results are output. The process of training a segmentation network with multi-gradient output capability based on the aforementioned multi-gradient segmentation dataset includes: By introducing a processing time penalty term, Focal Loss, and a dynamic weight adjustment mechanism, and employing transfer learning initialization, mixed precision training, and cosine annealing learning rate scheduling, the multi-gradient segmentation model is obtained. The extraction process of the differential feature point set includes: Morphological optimization, GrabCut post-processing, and edge sub-pixel thinning are performed on the three types of masks respectively. Then, high-response feature points in the RGB region are extracted according to the retention ratio to form a set of differential feature points corresponding to the g1 completely segmented mask, the g2 partial background mask, and the g3 large background mask. The process of aligning, fusing, and correcting errors in the three sets of point clouds based on a residual-driven mechanism to obtain a fused point cloud includes: The residual matrices of the three sets of point clouds are calculated, and the point cloud overlap and the mean of the residuals are used as the state space. Fusion weights are dynamically generated through a reinforcement learning strategy, and ICP registration and weighted fusion are performed on the three sets of point clouds to obtain the fused point cloud.

2. The method according to claim 1, characterized in that, The process of acquiring and standardizing RGB images of individual rice plants from multiple angles to obtain a standardized image dataset includes: Centered on a single rice plant, a multi-angle shooting strategy was adopted, taking one image every 10 degrees horizontally and three layers vertically. The original image was obtained by combining an array camera with a turntable for rotating shooting. Based on the original image, data cleaning, lens distortion correction, and color calibration are performed to obtain a standardized image dataset.

3. The method according to claim 1, characterized in that, The process of constructing a multi-gradient segmentation dataset based on the standardized image dataset includes: The standardized image dataset is segmented and labeled at the pixel level to form g1 complete segmentation mask, g2 partial background mask and g3 large background mask, and the mask confidence, occlusion level, growth period label and variety number are recorded.

4. The method according to claim 1, characterized in that, The segmentation network with multi-gradient output capability is an improved Faster-RCNN, and its model structure includes: The shared backbone network, RPN, RoI Head, and three gradient segmentation heads connected in parallel with the RoI Head; The segmentation head dynamically selects the corresponding gradient mask for output based on the confidence level of the proposal output by the RPN.

5. The method according to claim 1, characterized in that, The process of performing sparse-dense 3D reconstruction based on the differentiated feature point set to obtain three sets of point clouds corresponding to the three masks includes: The differential feature point set was sequentially reconstructed using SFM sparse reconstruction and MVS dense reconstruction using COLMAP to obtain point clouds g1, g2 and g3 respectively.

6. The method according to claim 1, characterized in that, The process of constructing a 3D model of rice based on the fused point cloud includes: Based on the fused point cloud and camera parameters, the 3D Gaussian Splatting method is used to assign position, covariance, color and transparency attributes to each three-dimensional point, and a three-dimensional rice model in PLY or STL format is generated through optimization iteration.

Citation Information

Patent Citations

  • Method for improving quality of space-time domain compressed video by using dense network

    CN120434403A

  • Three-dimensional modeling method and apparatus

    WO2025138753A1