Three-dimensional reconstruction method for hilly and mountainous area stereo agricultural scene
By combining multi-segment UAV flight paths with the MultiAttention-3DPreNet network and optimizing the 3DGS reconstruction algorithm using an adaptive hierarchical architecture, the problem of data acquisition and reconstruction in agricultural machinery operations in hilly and mountainous areas was solved. This achieved efficient and low-cost 3D reconstruction and semantic segmentation, improving the safety and intelligence level of agricultural machinery operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING AGRI MECHANIZATION INST MIN OF AGRI
- Filing Date
- 2026-03-23
- Publication Date
- 2026-05-29
AI Technical Summary
Existing 3D reconstruction technology has problems in agricultural machinery operations in hilly and mountainous areas, such as difficulty in balancing data acquisition security and accuracy, complex noise, low reconstruction efficiency, high computing power consumption, and lack of semantic attributes in the model. These problems result in low efficiency and high safety risks in agricultural machinery operations, and cannot meet the needs of intelligent agricultural machinery.
A multi-segment differentiated UAV flight path design and MultiAttention-3DPreNet network are adopted to acquire high-quality multi-view data. The 3DGS reconstruction algorithm is optimized by constructing a 3D model semantic segmentation system through an adaptive hierarchical architecture, dynamic density matching and collaborative pruning strategy, so as to achieve accurate differentiation of key elements such as field ridges, operation roads and crop areas.
It has improved the safety and efficiency of agricultural machinery operations in hilly and mountainous areas, reduced the cost of technology implementation, supported agricultural machinery route planning and intelligent development, and promoted the process of agricultural mechanization and intelligentization.
Smart Images

Figure CN122115764A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart agriculture, specifically a method for three-dimensional reconstruction of three-dimensional agricultural scenes in hilly and mountainous areas. Background Technology
[0002] Hilly and mountainous areas are important production bases for grain, oil and specialty agricultural products in my country, involving 1,429 counties and districts in 19 provinces, autonomous regions and municipalities with nearly 300 million agricultural population. The cultivated land and sown area account for more than 34% of the country's total. However, the region is characterized by low mountains, gentle slopes, valleys and terraces, with a terrain undulation of 50-500m and a slope range of 6°-25°. There are also a large number of terraced fields, slope fissures and gullies, and other fragmented terrain. The appearance of the landscape changes significantly with the seasons. These characteristics pose a severe challenge to agricultural machinery operations. Currently, the comprehensive mechanization rate of farming and harvesting in hilly and mountainous areas is only 50.05%, which is 21.2 percentage points lower than the national average. In some provinces, it is even less than 40%. The core problem is that agricultural machinery operations lack accurate terrain information support. Traditional two-dimensional planar planning methods are completely unsuitable. Existing three-dimensional reconstruction data mostly focus on plains or small-scale orchards and greenhouses, lacking high-quality data covering multiple perspectives in hilly and mountainous areas. The reconstruction model also has difficulty accurately restoring key terrain details such as field ridges and gullies, resulting in unreliable basis for agricultural machinery path planning. The current situation of small fields, varied slopes, and narrow and rugged field roads makes the transfer cost of agricultural machinery high, the operation efficiency low, and prone to safety accidents such as machinery collisions and overturning due to inaccurate terrain judgment.
[0003] When agricultural machinery operates in hilly and mountainous areas, it is necessary to accurately distinguish core scene elements such as field ridges, crop areas, operating roads, and gullies in order to achieve safe obstacle avoidance and precise operation. However, the reconstruction results of existing 3D Gaussian Splatting and variant models only contain spatial geometric information and do not semantically label these key elements closely related to agricultural machinery operations. As a result, agricultural machinery cannot identify obstacles and passable paths in the operating area, making it difficult to achieve intelligent operation planning. It can only rely on manual operation, which reduces the efficiency of operation and amplifies the risks of operation in complex terrain. This cannot meet the needs of intelligent development of agricultural machinery in hilly and mountainous areas.
[0004] Existing data acquisition and reconstruction technologies are poorly adapted to the needs of large-scale agricultural machinery operations in hilly and mountainous areas. At the data acquisition level, low-altitude drone operations are prone to collisions with steep slopes, while high-altitude operations lack precision, making it difficult to balance safety and data integrity. Furthermore, the region experiences frequent rainy seasons and variable lighting, resulting in data collection suffering from ghosting noise caused by drone shake, dynamic noise from human and livestock activities in the fields, and lighting differences at different times of day. High-quality multi-view data is extremely scarce. At the reconstruction technology level, the agricultural scenes in hilly and mountainous areas are vast and rich in geometric elements. 3DGS requires generating tens of millions of Gaussian primitives to accurately represent the terrain, often resulting in training times exceeding 100 hours. Ordinary consumer-grade graphics cards struggle to handle this, easily causing memory overflows and training interruptions, significantly increasing the cost of technology implementation and hindering its large-scale application in agricultural machinery operations. These problems collectively contribute to the severe lag in the development of agricultural mechanization in hilly and mountainous areas.
[0005] To address the aforementioned issues, those skilled in the art have sought solutions using 3DGS variant models. Taking the PlantGaussian model and the 3DGS-Ag model as examples, both are based on 3DGS technology and optimized for agricultural scenarios. They represent the most similar solutions to the technical approach of this invention in the current field of agricultural 3D reconstruction. Their specific details are as follows:
[0006] PlantGaussian model (cross-time and cross-scene crop visualization and reconstruction):
[0007] (1) Data collection: Greenhouse potted plants and small-scale field plots were selected as the collection scenarios. For single crops such as wheat, soybeans and tobacco, multi-view cameras were used to capture image data.
[0008] (2) Data preprocessing: Perform conventional denoising (such as Gaussian filtering) and size normalization on the acquired images, estimate the camera pose and generate sparse point clouds through the SfM algorithm, and use them as input for 3D Gaussian initialization;
[0009] (3) Model training: Based on the original 3DGS framework, the initialization strategy and gradient optimization rate of Gaussian primitives are optimized, and the local detail representation is enhanced for fine structures such as crop leaves and stems.
[0010] (4) Reconstruction output: Generate a three-dimensional visualization model of crops through rendering tools, supporting morphological comparison and scene migration visualization across time series, with the core application being crop growth status monitoring.
[0011] 3DGS-Ag model (high-fidelity reconstruction of a small scene in Taoyuan):
[0012] (1) Data collection: The peach orchard in the plain area was selected as the study area. Low-altitude circling flight of UAV was used to take pictures to obtain multi-view images. The flight altitude was fixed to ensure that the shooting angle evenly covers the canopy and trunk area and avoid the influence of terrain undulation.
[0013] (2) Dataset construction: Three types of small-scale datasets were established, including single trees, multiple trees, and fruiting peach trees. The image resolution was uniform, and blurry and occluded images were removed by manual screening to ensure data quality.
[0014] (3) Model optimization: Based on the 3DGS framework, and considering the characteristics of dense canopy and complex branches of fruit trees, the following optimizations were added: dynamic opacity reset (to reduce Gough element pruning), oversampling (to improve details), and distance weighting (to suppress artifacts);
[0015] (4) Reconstruction output: Generate a high-fidelity 3D model of the peach orchard, which can restore the fruit distribution and branch morphology.
[0016] The applicant believes that the above solution has the following defects:
[0017] Deficiencies in data acquisition and processing technology have resulted in a lack of high-quality reconstruction data for hilly and mountainous areas, as well as a lack of detailed topographic information.
[0018] Existing technologies for 3D reconstruction data acquisition in hilly and mountainous areas often employ single-flight schemes suitable for plains or small-scale scenarios, failing to consider the core features of hilly and mountainous areas. Low-altitude flights are prone to collisions with steep slopes, while high-altitude flights cannot capture detailed topographical information such as field ridges, slope fissures, and gullies. As a result, the acquired data only reflects the general terrain orientation and lacks the topographical details necessary for agricultural machinery operations. Furthermore, existing data denoising methods mostly use single filtering strategies, which cannot specifically address the diverse noise types in data acquired in hilly and mountainous areas, further degrading data quality and making it difficult to support high-precision reconstruction.
[0019] Large-scene reconstruction algorithm defects: low training efficiency, high computing cost, and inability to adapt to the large-scale needs of hilly and mountainous areas.
[0020] Existing 3DGS technology is not optimized for large scenes. Due to the wide range of scenes and rich geometric elements in this area, tens of millions of Gaussian primitives need to be generated to represent the terrain. The lack of a hierarchical reconstruction architecture leads to an indiscriminate increase in primitive size, resulting in long training times. The lack of a dynamic density matching mechanism results in the same primitive density in flat and complex areas, causing a waste of computing power. The lack of a collaborative pruning strategy makes it impossible to efficiently remove redundant primitives, and ordinary consumer-grade graphics cards are prone to memory overflow, which greatly increases the cost of technology implementation and cannot meet the reconstruction needs of large-scale agricultural machinery operations in hilly and mountainous areas.
[0021] The reconstruction model has functional defects: it only contains geometric information and lacks semantic attributes, making it unable to support intelligent navigation for agricultural machinery.
[0022] Existing 3D reconstruction technologies for agricultural scenarios only output geometric information such as spatial coordinates, without constructing a semantic annotation system and lacking a 2D-3D semantic association and fusion mechanism. They cannot accurately distinguish key elements of agricultural machinery operations such as field ridges, operating roads, crop areas, and gullies. As a result, the models can only present terrain undulations and cannot provide semantic guidance such as passable / obstacle avoidance and operating / non-operating areas for agricultural machinery, making it difficult to adapt to the actual needs of intelligent path planning for agricultural machinery. Summary of the Invention
[0023] This invention addresses the problems existing in the background technology by acquiring high-quality multi-view data covering detailed terrain information through differentiated flight path design and a multi-attention mechanism denoising network; it optimizes the large-scene 3DGS reconstruction algorithm by using an adaptive hierarchical architecture, dynamic density matching, and collaborative pruning strategies to shorten training time and reduce computing power consumption while ensuring reconstruction accuracy, thus achieving high efficiency and low cost in large-scene reconstruction in hilly and mountainous areas; and it constructs a 3D model semantic segmentation system to annotate key elements such as field ridges, operating roads, and crop areas, providing dual geometric and semantic support for agricultural machinery, solving the problem of high difficulty in agricultural machinery operation in hilly and mountainous areas, and ultimately promoting the mechanization and intelligent development of agriculture in hilly and mountainous areas.
[0024] Technical solution:
[0025] This invention discloses a method for three-dimensional reconstruction of a three-dimensional agricultural scene in hilly and mountainous areas, the method comprising the following steps:
[0026] S1. Data acquisition: Multi-segment differentiated drone flight paths are used to carry multispectral drones to acquire image data and obtain multi-view RGB image data.
[0027] S2. Perform image data preprocessing on the multi-view RGB image data to obtain preprocessed multi-view RGB image data;
[0028] S3. Input the preprocessed multi-view RGB image data into the improved 3DGS model and export 3D point cloud data; the improved 3DGS model is based on the traditional 3DGS model and embeds architecture layering, dynamic density control, and layered collaborative pruning modules, wherein:
[0029] 1) The 3D Gaussian primitives are first adaptively layered by the architecture layering module, dividing them into coarse-grained and fine-grained layers;
[0030] 2) Dynamic density control calculates the viewpoint overlap rate of adjacent images, reduces Gaussian density in areas of high overlap to avoid redundancy; based on local geometric curvature, it automatically increases the number of Gaussian primitives in areas of high complexity to ensure detail fidelity; and establishes a dynamic density threshold update mechanism to respond in real time to subsequent pruning feedback.
[0031] If the loss value increases after pruning, the dynamic density control module automatically increases the Gaussian density of the corresponding region; if the loss value is stable and the computing power is redundant, the pruning threshold is further optimized to reduce memory usage.
[0032] 3) For coarse-grained layers, the layered collaborative pruning module sets a high transparency threshold and directly deletes redundant Gaussian primitives that contribute little to the image; for fine-grained layers, it introduces a primitive importance metric, retains primitives with importance values higher than the threshold, and removes redundancies.
[0033] Preferably, in S2, image data preprocessing is performed based on the MultiAttention-3DPreNet network; the MultiAttention-3DPreNet network architecture includes a cascaded deblurring attention module RAAM, a moving point interference suppression module MDIM, and an illumination normalization module LNM.
[0034] Specifically:
[0035] Deblurring Attention Module RAAM: First, the spatial attention module locates areas with severe ghosting. Combining the temporal correlation of multiple frames, the temporal model captures the motion patterns between frames. The neural network learns the mapping relationship between the blurred image and the blur kernel, and outputs the blur kernel matrix. Combined with non-blind deconvolution LR, the image is made clear.
[0036] The Motion Interference Suppression Module (MDIM) uses a lightweight semantic segmentation subnetwork for dynamic and static binary classification, outputs a semantic mask, uses channel attention to determine which channels contain key information about the static scene, and calculates element-wise product.
[0037] The Lighting Normalization Module (LNM) uses channel attention weights to perform "targeted correction" on the lighting feature map to maintain global lighting balance and unifies multi-view lighting through normalization.
[0038] Preferably, in S1, the multi-segment differentiated UAV flight path includes full-scene low-altitude circling flight, full-scene high-altitude circling flight, orthographic flight path, low-altitude precision supplementation, and reshooting data.
[0039] Specifically, in S1, the multi-segment structure includes:
[0040] First round of full-area data acquisition: Initial image acquisition was completed according to the flight path;
[0041] Preliminary data processing: Remove severely blurred or occluded invalid images, and standardize, rename, and resize the data;
[0042] Camera pose and initial point cloud generation: The camera pose is calculated using the COLMAP tool to generate an initial sparse point cloud for the camera pose;
[0043] Second round of supplementary data collection: Based on the initial point cloud identification of areas with missing perspectives and insufficient data density, a supplementary shooting route is planned to complete the data supplementation, and finally the original dataset of multi-view images is formed.
[0044] Preferably, the method further includes step S4, semantic segmentation.
[0045] Specifically, S4 includes:
[0046] 1) Render the 3D point cloud data exported from S3 to obtain a 2D multi-view map; the multi-view map is obtained by obtaining a 2D map with semantic labels based on a 2D semantic segmentation model;
[0047] 2) 3D point cloud data is used for 3D semantic segmentation to obtain 3D preprocessed point cloud;
[0048] 3) The semantically labeled 2D image is fused with the 3D preprocessed point cloud to obtain a preliminary semantic point cloud;
[0049] 4) Fill in the gaps in the initial semantic point cloud to obtain the final semantic segmentation result.
[0050] Preferably, the 2D semantic segmentation model is based on a pixel-level semantic annotation mask generated by the LabelMe annotation tool. Each pixel in the mask corresponds to a preset semantic category, including field ridges, work roads, crop areas, and gullies, through a specific value, ensuring that each pixel has a unique semantic label.
[0051] Preferably, the precise association between 2D semantic tags and 3D preprocessed point clouds is achieved by matching the camera parameter projection matrix with the texture.
[0052] Beneficial effects:
[0053] Addressing the challenges of balancing security and accuracy in data collection in hilly and mountainous areas, as well as the complexity of noise, this technology employs a multi-segment differentiated drone flight path design. This approach avoids the risks of operating on steep slopes while achieving comprehensive coverage of detailed terrain features such as field ridges and slope fissures. Furthermore, by combining the MultiAttention-3DPreNet network with targeted suppression of various types of noise, including ghosting, dynamic interference, and lighting differences, it ensures the high quality and integrity of the reconstruction input data. This fundamentally improves the accuracy of terrain reconstruction in 3D models, overcoming the limitations of existing technologies that can only represent the orientation of terrain.
[0054] To address the bottlenecks of low efficiency and high computational consumption in large-scale scene reconstruction, a collaborative optimization strategy for native 3DGS is implemented. Through a closed-loop design involving coarse-grained and fine-grained layered reconstruction, on-demand allocation of Gaussian primitive density, and precise removal of redundant primitives, the training time is significantly shortened and the hardware computational threshold is lowered while ensuring detail fidelity. This makes large-scale implementation of three-dimensional agricultural scene reconstruction in hilly and mountainous areas possible. Meanwhile, the 3D model semantic segmentation function accurately fills the gap in existing technologies that lack semantic attributes, enabling precise differentiation of key elements in agricultural machinery operations such as field ridges, roads, and crop areas. This provides support for agricultural machinery path planning and safe obstacle avoidance, directly echoing the core objective of promoting agricultural mechanization and intelligence in hilly and mountainous areas, and significantly enhancing the practical value and feasibility of the technology. Attached Figure Description
[0055] Figure 1 This is a schematic diagram of multi-route planning in the example.
[0056] Figure 2 This is the overall framework diagram of the preprocessing network (MultiAttention-3DPreNet) of the present invention.
[0057] Figure 3 This is a diagram of the 3DGS collaborative optimization strategy of the present invention.
[0058] Figure 4 This is a framework diagram for the verification and application of the three-dimensional model of the present invention in agricultural machinery navigation.
[0059] Figure 5 This is a schematic diagram of the Ground Truth in the embodiment.
[0060] Figure 6 The image shown is a reconstruction result from an example.
[0061] Figure 7 This is a schematic diagram of the point cloud in the embodiment.
[0062] Figure 8 This is a flowchart of the method of the present invention. Detailed Implementation
[0063] The present invention will be further described below with reference to embodiments, but the scope of protection of the present invention is not limited thereto:
[0064] Combination Figure 8 This invention discloses a method for three-dimensional reconstruction of a three-dimensional agricultural scene in hilly and mountainous areas, the method comprising the following steps:
[0065] S1. Data acquisition: Multi-segment differentiated drone flight routes are used to carry multispectral drones (DJIPHANTOM 4 MULTISPECTRAL multispectral drones) to acquire image data and obtain multi-view RGB image data.
[0066] Combination Figure 1 The multi-route planning diagram shown selects typical hilly and mountainous areas such as Huize County, Qujing City, Yunnan Province (steep slope terraces) and Wangcun Town, Shexian County, Huangshan City, Anhui Province (gentle slope), covering a slope range of 6°-25°. It includes key terrain features for agricultural machinery operations such as field ridges, gullies, and slope fissures, ensuring the representativeness of the dataset. The multi-segment structure includes:
[0067] First round of full-area data acquisition: Initial image acquisition was completed according to the flight path;
[0068] Preliminary data processing: Remove severely blurred or occluded invalid images, and standardize, rename, and resize the data;
[0069] Camera pose and initial point cloud generation: The camera pose is calculated using the COLMAP tool to generate an initial sparse point cloud for the camera pose;
[0070] Second round of supplementary data collection: based on Figure 7 The initial point cloud was used to identify areas with missing viewpoints and insufficient data density. A supplementary shooting route was planned to supplement the data, ultimately forming the original dataset of multi-view images. Figure 5 The image shows example data from the dataset.
[0071] The full-domain data acquisition process includes full-scene low-altitude flight, full-scene high-altitude flight, and orthographic flight path. In this embodiment, the following flight parameters are used:
[0072] (1) Full-scene low-altitude flight: flight speed 2.7M / S, flight altitude 15.2M-16.8M, gimbal pitch angle -20.0°, F: 50% S: 30%, capturing fine landforms such as field ridges and gullies.
[0073] (2) Full-scene surround flight: flight speed 2.4M / S, flight altitude 30.0M-32.0M, gimbal pitch angle -29.8°, F: 60% S: 30%, avoids collision with steep slopes, and covers the entire terrain.
[0074] (3) Orthogonal flight path: flight speed 3.5M / S, flight altitude 45.0M, gimbal pitch angle -90.0°, fixed-point turn, providing global reference view.
[0075] The supplementary data acquisition process includes low-altitude precision supplementation and re-capture data, using the following flight parameters:
[0076] (4) Low-altitude accuracy supplement: Flight speed 3.1M / S, gimbal pitch angle -24.8°, F: 50% S: 30%, quadrant accuracy supplement according to terrain direction.
[0077] (5) Reshoot data: flight speed 3.2M / S, flight altitude 6.0M-8.0M, gimbal pitch angle -20.3°, for precise reshoots in shadowed and obscured areas.
[0078] S2. Perform image data preprocessing on the multi-view RGB image data to obtain preprocessed multi-view RGB image data;
[0079] Combination Figure 2 The preprocessing network (MultiAttention-3DPreNet) consists of three modules: RAAM (Refined Alignment & Anti-blurring with Attention Module), MDIM (Moving Disturber Interference Suppression Module), and LNM (Lighting Normalization Module). Among them:
[0080] The deblurring attention module first locates areas with severe ghosting using a spatial attention module. It then combines the temporal correlation of multiple frames to capture inter-frame motion patterns using a temporal model. A neural network learns the mapping relationship between the blurred image and the blur kernel, outputting a blur kernel matrix. This is then combined with non-blind deconvolution (LR) to sharpen the image.
[0081] Spatial attention:
[0082] Temporal attention:
[0083] in, Spatial attention weights are used to emphasize important spatial locations. conv: convolutional layer; pool: pooling layer; σ: sigmoid activation function; : Flattened temporal feature vector; GRU: Gated Recurrent Unit, used to capture temporal dependencies; Temporal attention weights are used to emphasize important time frames.
[0084] The motion interference suppression module uses a lightweight semantic segmentation subnetwork for dynamic and static binary classification, outputs a semantic mask, and uses channel attention to determine which channels contain key information about the static scene and calculates the element-wise product.
[0085] Channel attention:
[0086] Dynamic interference suppression:
[0087] in, Global average pooling compresses spatial information into channel-level feature descriptions; MLP: Multi-Layer Perceptron, used to learn channel weights; σ: Sigmoid activation function; Channel attention weights are used to emphasize important feature channels. : Mask, used to identify dynamic / interference regions; 0.9: Interference suppression coefficient; : Static characteristics after suppressing dynamic interference.
[0088] The illumination normalization module uses channel attention weights to perform "targeted correction" on the illumination feature map to maintain global illumination balance. It unifies multi-view illumination through normalization, using the following formula:
[0089] in: Features after illumination estimation; Global mean calculation, used to estimate overall light intensity; : Minimal constant, to prevent division by zero.
[0090] S3. Input the preprocessed multi-view RGB image data into the improved 3DGS model and export 3D point cloud data; combined with... Figure 3 The improved 3DGS model is based on the traditional 3DGS model and incorporates architecture layering, dynamic density control, and layered collaborative pruning modules, wherein:
[0091] 1) The 3D Gaussian primitives are first adaptively layered by the architecture layering module, dividing them into coarse-grained and fine-grained layers;
[0092] 2) Dynamic density control calculates the viewpoint overlap rate of adjacent images, reduces Gaussian density in areas of high overlap to avoid redundancy; based on local geometric curvature, it automatically increases the number of Gaussian primitives in areas of high complexity to ensure detail fidelity; and establishes a dynamic density threshold update mechanism to respond in real time to subsequent pruning feedback.
[0093] If the loss value increases after pruning, the dynamic density control module automatically increases the Gaussian density of the corresponding region; if the loss value is stable and the computing power is redundant, the pruning threshold is further optimized to reduce memory usage.
[0094] In a preferred embodiment, the viewpoint overlap rate of adjacent images is calculated using the following formula:
[0095] Overlap(i,j) = |P_i ∩ P_j| / |P_i|
[0096] Where P_i and P_j are the three-dimensional spatial sampling regions corresponding to the i-th and j-th images; overlap rate = (number of pixels in the overlapping region / total number of pixels in a single image) × 100.
[0097] ① Obtain camera parameters: Obtain the intrinsic and extrinsic parameters of each image using COLMAP;
[0098] ② Projecting into three-dimensional space: Projecting the pixels of each image back into three-dimensional space;
[0099] ③ Computational spatial overlap: Take the intersection of the observation areas of adjacent camera positions;
[0100] ④ Quantization overlap rate: Number of overlapping pixels / Total number of pixels.
[0101] In the preferred embodiment, the adaptive determination has a high degree of overlap, and the preliminary judgment is as follows:
[0102] >70%: High overlap, reducing Gaussian density by 20-40%;
[0103] 40-70%: Normal, maintain default density;
[0104] < 40%: Insufficient overlap; density can be increased appropriately.
[0105] Confirmation method:
[0106] # Pseudocode Example
[0107] # Reduce density
[0108] if overlap_rate > THRESHOLD_HIGH:
[0109] gaussian_density *= reduction_factor
[0110] # Increase density
[0111] elif overlap_rate < THRESHOLD_LOW:
[0112] gaussian_density *= increase_factor
[0113] Threshold determination method:
[0114] The initial threshold can be preset according to the scale of the scene (e.g., 60-70%).
[0115] Automatic adjustment based on training feedback: If the loss value increases after pruning → increase the threshold; if the loss is stable and computing power is redundant → decrease the threshold.
[0116] In a preferred embodiment, the distribution of Gaussian primitives is used to automatically increase the number of Gaussian primitives in regions of high complexity based on local geometric curvature.
[0117] Smooth regions (field ridges): have low curvature and reduce primitives.
[0118] Complex regions (edges of fields): high curvature, increasing the number of primitives.
[0119] Specific implementation:
[0120] # pseudocode
[0121] def calculate_curvature(point_cloud, k_neighbors=10):
[0122] # Find k nearest neighbors for each point
[0123] neighbors = find_k_nearest(point, k_neighbors)
[0124] # Calculate the covariance matrix
[0125] cov = covariance(neighbors)
[0126] # Curvature = Minimum eigenvalue / Sum of eigenvalues
[0127] eigenvalues = eig(cov)
[0128] curvature = eigenvalues[0] / sum(eigenvalues)
[0129] return curvature.
[0130] 3) For coarse-grained layers, the layered collaborative pruning module sets a high transparency threshold and directly deletes redundant Gaussian primitives that contribute little to the image; for fine-grained layers, it introduces a primitive importance metric, retains primitives with importance values higher than the threshold, and removes redundancies.
[0131] The core of the improved 3DGS model optimization strategy is the closed-loop linkage of three modules: hierarchical layering provides a hierarchical benchmark for density control / pruning, density control provides redundancy basis for pruning, and pruning inversely optimizes the dynamic density control strategy. This significantly reduces training computational power consumption while preserving detail.
[0132] exist Figure 3In the diagram, the gray area represents the original model's workflow. Embedded optimizations are performed within the original workflow, while the original dynamic density control is improved. The original model used density control based on the error in the projected image, focusing on areas with significant errors. The improved model incorporates a layered differentiation strategy to reduce the waste caused by excessively high Gaussian unit density at coarse-grained levels. The final generated... Figure 6 The 2D display image of the 3D model shown is... Figure 5 Comparison of data from real-world scenarios.
[0133] S4, 2D-3D semantic segmentation, combined with Figure 4 Specifically, it includes:
[0134] 1) Rendering Multi-View Data: The 3D point cloud data exported from S3 (containing X / Y / Z coordinates, RGB color, and surface normals for each point) is rendered into 2D multi-view images using the SIBR_viewer program. Simultaneously, the depth map of each image is exported (recording the 3D height of each pixel). These 2D multi-view images are the rendering results of the 3DGS reconstructed model, not the original noisy images captured by the drone, making them more suitable for semantic analysis from the perspective of agricultural machinery operations.
[0135] 2) 2D Image Semantic Annotation: The LabelMe annotation tool is used to classify and annotate the core areas of the rendered 2D multi-view image, generating a pixel-level semantic annotation mask. Each pixel in the mask corresponds to a preset semantic category (including field ridges, work roads, crop areas, etc.) through a specific value, ensuring that each pixel has a unique semantic label.
[0136] 3) Training the 2D semantic segmentation model: Using the labeled 2D images and semantic masks as training samples, a 2D semantic segmentation model (U-Net++) is trained to learn the correspondence between 2D pixel visual features and semantic categories, enabling it to automatically identify various elements from 2D images. The trained U-Net++ model is then used to automatically perform semantic segmentation on all 2D multi-view images, outputting 2D images with semantic labels.
[0137] 4) 2D-3D Semantic Mapping: Accurate mapping from 2D semantic labels to 3D point clouds is achieved through camera parameter projection matrices. Given the coordinates (u, v) of a pixel in a 2D image, the corresponding 3D spatial coordinates (X, Y, Z) are inversely calculated using the projection matrix, combining the camera's intrinsic parameters (camera properties) and extrinsic parameters (camera spatial pose). The semantic label of this pixel is then assigned to the point with coordinates (X, Y, Z) in the 3D point cloud. This process is repeated to complete the mapping from 2D pixel annotations to 3D point cloud labels. Verification using the camera intrinsic parameter projection matrix and texture matching ensures strict spatial alignment between the 3D point cloud and the 2D semantic annotation mask, providing spatial consistency for subsequent fusion.
[0138] 5) Train the 3D semantic segmentation model: Use 3D point clouds with 2D projected semantic labels as training samples to train a 3D semantic segmentation model (PointNet++), enabling it to perform semantic segmentation directly from 3D point clouds. Use the trained PointNet++ to perform semantic segmentation on the original 3D point clouds and output the 3D semantic segmentation results.
[0139] 6) 2D-3D fusion: The 2D and 3D segmentation results are fused by voting to obtain a preliminary semantic point cloud.
[0140] 7) Fill in the empty points in the preliminary semantic point cloud to obtain the final semantic segmentation result.
[0141] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A method for three-dimensional reconstruction of a three-dimensional agricultural scene in hilly and mountainous areas, characterized in that... The method includes the following steps: S1. Data acquisition: Multi-segment differentiated UAV flight paths are used to carry multispectral UAVs to acquire image data and obtain multi-view RGB image data. S2. Perform image data preprocessing on the multi-view RGB image data to obtain preprocessed multi-view RGB image data; S3. Input the preprocessed multi-view RGB image data into the improved 3DGS model and export 3D point cloud data; the improved 3DGS model is based on the traditional 3DGS model and embeds architecture layering, dynamic density control, and layered collaborative pruning modules, wherein: 1) The 3D Gaussian primitives are first adaptively layered by the architecture layering module, dividing them into coarse-grained and fine-grained layers; 2) Dynamic density control calculates the viewpoint overlap rate of adjacent images, reduces Gaussian density in areas of high overlap to avoid redundancy; based on local geometric curvature, it automatically increases the number of Gaussian primitives in areas of high complexity to ensure detail fidelity; and establishes a dynamic density threshold update mechanism to respond in real time to subsequent pruning feedback. If the loss value increases after pruning, the dynamic density control module automatically increases the Gaussian density of the corresponding region; if the loss value is stable and the computing power is redundant, the pruning threshold is further optimized to reduce memory usage. 3) For coarse-grained layers, the layered collaborative pruning module sets a high transparency threshold and directly deletes redundant Gaussian primitives that contribute little to the image; for fine-grained layers, it introduces a primitive importance metric, retains primitives with importance values higher than the threshold, and removes redundancies.
2. The method according to claim 1, characterized in that... In S2, image data preprocessing is performed based on the MultiAttention-3DPreNet network; the MultiAttention-3DPreNet network architecture includes a cascaded deblurring attention module RAAM, a moving point interference suppression module MDIM, and an illumination normalization module LNM.
3. The method according to claim 2, characterized in that: Deblurring Attention Module RAAM: First, the spatial attention module locates areas with severe ghosting. Combining the temporal correlation of multiple frames, the temporal model captures the motion patterns between frames. The neural network learns the mapping relationship between the blurred image and the blur kernel, and outputs the blur kernel matrix. Combined with non-blind deconvolution LR, the image is made clear. The Motion Interference Suppression Module (MDIM) uses a lightweight semantic segmentation subnetwork for dynamic and static binary classification, outputs a semantic mask, uses channel attention to determine which channels contain key information about the static scene, and calculates element-wise product. The Lighting Normalization Module (LNM) uses channel attention weights to perform "targeted correction" on the lighting feature map to maintain global lighting balance and unifies multi-view lighting through normalization.
4. The method according to claim 1, characterized in that... In S1, the multi-segment differentiated UAV flight routes include full-scene low-altitude circling flight, full-scene high-altitude circling flight, orthographic flight, low-altitude precision supplementation, and reshooting data.
5. The method according to claim 1, characterized in that... In S1, the multi-segment structure includes: First round of full-area data acquisition: Initial image acquisition was completed according to the flight path; Preliminary data processing: Remove severely blurred or occluded invalid images, and standardize, rename, and resize the data; Camera pose and initial point cloud generation: The camera pose is calculated using the COLMAP tool to generate an initial sparse point cloud for the camera pose; Second round of supplementary data collection: Based on the initial point cloud identification of areas with missing perspectives and insufficient data density, a supplementary shooting route is planned to complete the data supplementation, and finally the original dataset of multi-view images is formed.
6. The method according to claim 1, characterized in that... It also includes S4 and semantic segmentation.
7. The method according to claim 6, characterized in that... S4 specifically includes: 1) Render the 3D point cloud data exported from S3 to obtain a 2D multi-view map; the multi-view map is obtained by obtaining a 2D map with semantic labels based on a 2D semantic segmentation model; 2) 3D point cloud data is used for 3D semantic segmentation to obtain 3D preprocessed point cloud; 3) The semantically labeled 2D image is fused with the 3D preprocessed point cloud to obtain a preliminary semantic point cloud; 4) Fill in the gaps in the initial semantic point cloud to obtain the final semantic segmentation result.
8. The method according to claim 7, characterized in that... The 2D semantic segmentation model is based on a pixel-level semantic annotation mask generated by the LabelMe annotation tool. Each pixel in the mask corresponds to a preset semantic category, including field ridges, work roads, crop areas, and gullies, through a specific value, ensuring that each pixel has a unique semantic label.
9. The method according to claim 7, characterized in that... Accurate association between 2D semantic tags and 3D preprocessed point clouds is achieved by matching camera parameter projection matrices with textures.