Three-dimensional gaussian multi-robot collaborative mapping and optimization method and system for complex agroforestry scenarios

By employing a multi-machine collaborative mapping and optimization method, the problems of inconsistent lighting, dynamic disturbances, and sparse viewpoints in complex agricultural and forestry scenes were solved, generating a high-precision 3D Gaussian scene model suitable for digital modeling of agricultural and forestry scenes and monitoring of vegetation growth.

CN122391539APending Publication Date: 2026-07-14DALIAN UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2026-04-20
Publication Date
2026-07-14

Smart Images

  • Figure CN122391539A_ABST
    Figure CN122391539A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of three-dimensional reconstruction and intelligent agriculture, and particularly relates to a three-dimensional Gaussian multi-machine collaborative mapping and optimization method and system for complex agricultural and forestry scenes. In view of the problem of low reconstruction accuracy caused by inconsistent exposure, dynamic disturbance of vegetation and sparse multi-machine view angles in complex agricultural and forestry scenes, the application sequentially carries out multi-machine data preprocessing, spatiotemporal outlier elimination, intermediate frame generation based on FILM and DepthAnything, and 3DGS joint optimization. Dynamic masks are generated through spatiotemporal residual double modal judgment to block dynamic interference at the level of loss function; sparse view angles are completed by using virtual intermediate frames to expand the observation data set. The application can effectively eliminate dynamic artifacts and cavities, realize high-precision three-dimensional reconstruction of complex agricultural and forestry scenes, and be directly applied to digital operation scenes such as vegetation monitoring and automatic path planning of forest land and orchards.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of 3D reconstruction, multi-machine collaborative mapping, 3D Gaussian Splatting (3DGS), outdoor scene perception and smart agriculture technology, specifically involving a 3D Gaussian multi-machine collaborative mapping and optimization method and system for complex agricultural and forestry scenarios. Background Technology

[0002] With the rapid development of smart agriculture and the digitalization of agricultural and forestry scenarios, 3D mapping of complex agricultural and forestry scenarios has become a core support for achieving precise management and improving operational efficiency. 3D mapping of complex agricultural and forestry scenarios is a fundamental step in scenario digitization. As typical complex outdoor scenarios, complex agricultural and forestry scenarios are characterized by high vegetation density, intricate foliage, undulating terrain, and variable lighting conditions, placing extremely high demands on the accuracy, robustness, and efficiency of 3D mapping.

[0003] 3DGS technology, as a 3D reconstruction solution that has emerged in recent years, has gradually replaced traditional point cloud reconstruction and mesh reconstruction methods, becoming the mainstream technology choice for 3D mapping of complex agricultural and forestry scenes due to its outstanding advantages such as fast reconstruction speed, high rendering quality, complete detail preservation, and suitability for large-scale outdoor scenes. Multi-machine collaborative mapping mode, through the synchronous acquisition of data by multiple mobile devices (such as drones and ground inspection robots), can expand the mapping range, improve mapping efficiency, and solve the problems of incomplete coverage and low efficiency of single-device mapping, making it suitable for the rapid digital modeling needs of large-area complex agricultural and forestry scenes. However, when existing 3DGS reconstruction technology and multi-machine collaborative mapping methods are applied in complex agricultural and forestry scenes, they suffer from many intractable technical defects due to the characteristics of the scene itself and the limitations of acquisition conditions, seriously affecting the mapping quality and practical application effects. Specifically, these defects manifest in the following three aspects: (1) Inconsistent illumination and exposure: The natural light intensity of complex agricultural and forestry scenes is affected by factors such as weather, time, and terrain occlusion, and there are obvious dynamic changes. When such images are directly used for 3DGS reconstruction, defects such as noisy Gaussian, image ghosting, and geometric artifacts are easily generated, resulting in texture distortion and blurred details in the reconstructed model, which cannot accurately reflect the branch and leaf structure of agricultural and forestry vegetation and the terrain features of the scene.

[0004] (2) Dynamic disturbance problem: The vegetation, leaves and twigs in complex agricultural and forestry scenarios are light and easily disturbed by a breeze, resulting in high-frequency and irregular swaying, which are typical dynamic disturbance factors.

[0005] (3) Sparse viewpoint problem in multi-camera scenarios: In complex agricultural and forestry scenes, the vegetation density is high and the branches and leaves severely obstruct the view, which strictly limits the flight / inspection paths of multiple acquisition devices, making it impossible to achieve full-view, high-density observation. At the same time, in order to improve mapping efficiency, the viewpoint overlap rate is often reduced when acquiring data from multiple devices, resulting in a sparse distribution of the acquired observation viewpoints. Directly using sparse viewpoint data for 3DGS reconstruction will result in problems such as missing reconstruction of key areas (such as the interior of agricultural and forestry canopy and low-lying areas), distortion of texture and geometric structure, and insufficient model density, which cannot meet the requirements of mapping accuracy for the digitization of agricultural and forestry scenarios.

[0006] Therefore, existing methods are insufficient to achieve stable, complete, and high-precision 3DGS multi-machine collaborative mapping and optimization in complex agricultural and forestry scenarios, which restricts the practical deployment and promotion of 3DGS technology in the field of agricultural and forestry scenario digitization. A targeted optimization solution is urgently needed to solve the above problems. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a 3D Gaussian multi-machine collaborative mapping and optimization method for complex agricultural and forestry scenes. This method specifically addresses issues such as inconsistent exposure, dynamic disturbances caused by wind blowing through vegetation and leaves, and poor reconstruction quality and low robustness due to sparse multi-machine observation perspectives in complex agricultural and forestry scenes. It achieves this through multi-machine data preprocessing, spatiotemporal joint outlier removal, intermediate frame generation fused with FILM and DepthAnything, and joint 3DGS optimization mapping. Figure 4 The core module enables robust and high-precision 3DGS multi-machine collaborative mapping and optimization of complex agricultural and forestry scenes, meeting the actual needs of digital modeling of agricultural and forestry scenes. It can be directly applied to scenarios such as digital modeling of agricultural and forestry scenes, vegetation growth monitoring, scene path planning, precision management and automated operation.

[0008] The technical solution of the present invention is as follows: A three-dimensional Gaussian multi-machine collaborative mapping and optimization method for complex agricultural and forestry scenarios is proposed, with the following specific steps: Step 1: Multi-machine data preprocessing: Acquire complex agricultural and forestry scene image sequences, real depth images and pose data synchronously collected by multiple mobile devices, construct initial observation data, and perform exposure equalization and color alignment processing on the image sequences; Step 2, Spatiotemporal Joint Outlier Removal Mask Generation: Initialize a 3D Gaussian field using initial observation data; use the image sequence after exposure equalization and color alignment as the real color image; perform forward projection rendering using the initialized 3D Gaussian field and the mobile device pose corresponding to the current frame to obtain the predicted color image and predicted depth image; calculate the photometric spatiotemporal residual and geometric depth spatiotemporal residual by subtracting them pixel by pixel from the real color image and the real depth image; when the photometric spatiotemporal residual or geometric depth spatiotemporal residual at a certain pixel position is greater than the set static tolerance threshold, it is determined as a dynamic outlier, and a pixel-level dynamic removal mask matrix is ​​generated accordingly. Step 3: Viewpoint completion based on large motion frame interpolation and depth enhancement: For sparse viewpoint regions of multi-camera observation, a depth estimation model (Depth Anything) is used to perform monocular depth estimation on the real frames of the preprocessed image sequence to generate a depth map. A FILM model is then used in conjunction with the geometric constraints of the depth map to generate virtual intermediate frames between real frame pairs that meet the viewpoint overlap rate and time interval threshold. Valid virtual intermediate frames that have passed depth deviation verification are added to the observation dataset. Step 4, 3DGS Joint Optimization Mapping: The expanded observation dataset is integrated. In the joint optimization stage of pose tracking and mapping, the dynamic removal mask matrix is ​​multiplied by the joint loss function containing color loss and depth loss, so that the pixel loss weight of the pixel identified as a dynamic outlier is forced to zero. When calculating the gradient during backpropagation, the path of the dynamic outlier participating in the iterative update of the 3D Gaussian sphere parameters is cut off from the bottom layer, and finally a 3D Gaussian scene model without dynamic artifacts is output.

[0009] The specific process of performing exposure equalization and color alignment on the image sequence in step 1 is as follows: Calculate the average brightness and color mean of all acquired images; Based on the average brightness and color mean values, brightness stretching and gamma correction are performed on a single frame image; Extract vegetation area features from images and adjust color channel parameters to ensure consistent vegetation colors across images collected by different devices.

[0010] In step 2, when generating the pixel-level dynamic culling mask matrix, the mask value of the pixel position determined as a dynamic outlier is set to 0, and the mask value of the remaining static area is set to 1.

[0011] In step 4, the three-dimensional Gaussian sphere parameters include the center position, covariance matrix, opacity, and spherical harmonic function coefficients; the joint loss function includes a color loss term using L1 loss combined with D-SSIM and a depth loss term using L1 loss.

[0012] A 3D Gaussian multi-machine collaborative mapping and optimization system for complex agricultural and forestry scenes is provided to implement the aforementioned method. It includes a multi-machine data preprocessing module, a spatiotemporal outlier removal module, an intermediate frame generation and completion module, and a 3DGS joint optimization mapping module, executed sequentially. These modules work collaboratively to address exposure, dynamic perturbation, and viewpoint sparsity issues in 3DGS multi-machine mapping of complex agricultural and forestry scenes. The specific details of each module are as follows: The core function of the multi-machine data preprocessing module is to standardize and normalize the raw data of complex agricultural and forestry scenes acquired by multiple acquisition devices, eliminating the negative impacts of differences in equipment and lighting, and providing high-quality input data for subsequent modules. Specifically, multiple acquisition devices (preferably drones and ground inspection robots) are used to simultaneously acquire RGB images and GNSS / IMU pose data in a designated area of ​​complex agricultural and forestry, and relevant acquisition parameters are recorded synchronously to ensure consistency in time and data. Then, the lighting and color features of the image sequence are extracted, and contrast stretching and multi-channel color alignment are performed to obtain an image sequence with balanced exposure and color alignment. Subsequently, the multi-machine pose data is time-stamped, extrinsic parameter calibrated, and anomaly filtered to ensure that the poses of all devices are unified in the same world coordinate system, avoiding reconstruction misalignment and distortion caused by pose deviations. Finally, the images are optimized by denoising and edge enhancement to enhance the details of agricultural and forestry vegetation and terrain, providing reliable support for subsequent processing.

[0013] The aforementioned spatiotemporal outlier removal module addresses mapping artifacts caused by swaying branches and sudden changes in local exposure in complex agricultural and forestry scenes. It employs a "render-observation" residual and bimodal judgment mechanism to physically block dynamic interference, ensuring the purity of static reconstruction. First, forward rendering is performed using an initialized 3D Gaussian field and the current mobile device pose to obtain predicted color and depth images. Second, pixel-level subtraction is performed between the predicted and actual observed images to calculate the photometric spatiotemporal residual and geometric depth residual. Subsequently, a bimodal threshold judgment is introduced to identify pixels with residuals exceeding the limit as dynamic outliers violating geometric consistency (such as wind-blown branches), generating a dynamic removal mask matrix. Finally, in the joint optimization phase, this mask matrix is ​​inverted and directly applied to the loss function, forcibly zeroing the loss weights of dynamic outlier pixels. This approach completely cuts off the path for dynamic noise points to participate in the iterative update of the 3D Gaussian sphere parameters at the gradient level of backpropagation, effectively eliminating ghosting and geometric distortion.

[0014] The intermediate frame generation and completion module addresses the issue of sparse observation perspectives in multi-camera observations by fusing the FILM and DepthAnything models to complete sparse perspectives and increase observation density. First, the DepthAnything model is used to perform high-precision monocular depth estimation and smoothing denoising on preprocessed real frames to obtain reliable geometric constraints. Then, frame pairs with reasonable perspectives and time intervals are selected from the multi-camera real frames as benchmarks. Using the FILM frame interpolation model, combined with depth constraints, virtual intermediate frames that conform to the geometric structure of the real scene are generated. Finally, the generated intermediate frames undergo depth consistency verification and quality assessment. Invalid frames are removed, and valid intermediate frames are added to the observation set to form an expanded observation dataset, solving the problems of reconstruction holes and structural distortion under sparse perspectives.

[0015] The 3DGS joint optimization mapping module jointly optimizes the expanded observation dataset with a static Gaussian field after outlier removal to generate a high-precision, complete 3DGS reconstruction model of complex agricultural and forestry scenes. First, all observation images, poses, and depth data are projected onto the same world coordinate system, establishing a correspondence between pixels and 3D Gaussian points and correcting camera distortion. Then, a joint loss function including color loss, depth loss, and transparency regularization is constructed, and the Gaussian parameters are iteratively optimized using a gradient descent algorithm to ensure a high degree of matching between the reconstructed model and the observation data. Next, the optimized Gaussian field is denoised, smoothed, and fine-tuned to improve model continuity and integrity. Finally, a high-quality reconstruction model without dynamic artifacts or holes is output, meeting the needs of digital applications in complex agricultural and forestry scenes.

[0016] The beneficial effects of this invention are: (1) Strong resistance to dynamic disturbances: Through the spatiotemporal joint outlier removal module, dynamic Gaussian points and exposure noise Gaussian points caused by wind disturbance of vegetation branches and leaves are filtered from both spatial and temporal dimensions. This effectively solves the problems of model blurring, voids and structural disorder caused by dynamic disturbances in traditional 3DGS reconstruction, ensuring the geometric accuracy and integrity of the reconstructed model and adapting to the dynamic interference characteristics of complex agricultural and forestry scenes.

[0017] (2) Effectively solve the problem of sparse viewpoints: The intermediate frame generation technology that integrates FILM and DepthAnything is adopted to generate virtual intermediate frames that conform to geometric constraints without increasing the hardware acquisition cost. This fills the gap in the sparse viewpoints of multiple machines, improves the density and overlap of the observation viewpoints, and solves the problems of missing key areas, geometric distortion, and insufficient model compactness in the reconstruction under sparse viewpoints. It is suitable for the characteristics of dense vegetation and limited viewpoints in complex agricultural and forestry scenes. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the overall process of this invention. Figure 2This is a schematic diagram illustrating the spatiotemporal outlier removal principle of the present invention; Figure 3 This is a schematic diagram illustrating the intermediate frame generation and viewpoint completion of the present invention; Figure 4 The results of 3DGS orchard reconstruction without using the method of this invention are shown, where (a) is a real image captured by a mobile device and (b) is the result reconstructed by the algorithm. Figure 5 The results of 3DGS orchard reconstruction using the method of the present invention are shown, where (a) is a real image captured by a mobile device and (b) is the result reconstructed by the algorithm. Detailed Implementation

[0019] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions. A three-dimensional Gaussian multi-machine collaborative mapping and optimization method for complex agricultural and forestry scenarios, such as... Figure 1 As shown, the specific implementation steps are as follows: Step 1: Multi-machine data acquisition and preprocessing. Three UAVs were used as acquisition devices. In the target complex forest scene, RGB images and GNSS / IMU pose data were acquired synchronously according to the preset grid path. During the acquisition process, the timestamp, device number, and exposure parameters of each frame were recorded. Exposure equalization and color alignment were performed on the acquired images. The average brightness of all images was calculated to be 128, and the average color value (R:110, G:135, B:80) was calculated. Based on this average value, brightness stretching and gamma correction (gamma value 1.2) were performed on each frame of the image, and the color channel parameters were adjusted to unify the color of the forest vegetation. The GNSS / IMU pose data was time-stamped and the camera's intrinsic and extrinsic parameters were calibrated through extrinsic parameter calibration. Abnormal pose data caused by equipment jitter and signal interference were removed, and a unified world coordinate system was established. Gaussian filtering (3×3 filter kernel size) and edge enhancement were performed on the images to obtain preprocessed image data and pose data.

[0020] Step 2: Initial 3DGS Reconstruction and Spatiotemporal Outlier Mask Generation (Core Correction) Based on the preprocessed image and pose data, a 3D Gaussian field is initialized; then, the spatiotemporal consistency check stage begins: the currently initialized 3D Gaussian field is projected and rendered onto the mobile device pose corresponding to the current frame to obtain the predicted color image and depth image; the predicted image is then subtracted pixel by pixel from the observed image acquired by the real sensor. The photometric spatiotemporal residual threshold is set to 0.25, and the geometric depth spatiotemporal residual threshold is set to 0.3m. For swaying leaves or small branches in forest scenes, their residuals often fluctuate drastically. When the photometric or depth residual at a certain pixel location exceeds the above-set threshold, the system identifies it as a dynamic outlier and generates a dynamic removal mask matrix consistent with the image resolution (the effective static region mask value is 1, and the dynamic outlier region mask value is 0), preparing for subsequent gradient blocking. The principle of spatiotemporal outlier removal is as follows: Figure 2 As shown.

[0021] Step 3: Intermediate Frame Generation and Verification. The DepthAnything-V2 model is used to estimate the depth of the preprocessed real frames, generating a 1920×1080 resolution depth map. The depth map is then Gaussian smoothed. Real frame pairs with a viewpoint overlap of 40%-50% and a time interval of 0.5s are selected as the benchmark for intermediate frame generation. The FILM model, combined with depth map constraints, is used to generate virtual intermediate frames, with two intermediate frames generated between each pair of real frames. Depth consistency is verified on the generated intermediate frames. A depth deviation threshold of 0.2m is set, and intermediate frames exceeding the threshold are discarded. Valid intermediate frames with clear textures and consistent colors are selected and added to the observation dataset to form the expanded observation dataset. The principles of intermediate frame generation and viewpoint completion are as follows: Figure 3 As shown.

[0022] Step 4: 3DGS Joint Optimization and Model Output (Core Correction) The expanded observation dataset (real frames + effective intermediate frames) and its pose and depth data are projected onto the world coordinate system to establish the correspondence between pixels and 3D Gaussian points. A joint loss function is constructed, where color loss uses L1 loss combined with D-SSIM, and depth loss uses L1 loss. During backpropagation loss calculation, the dynamic culling mask matrix generated in Step 2 is introduced, and the mask matrix is ​​multiplied by the loss function (i.e., the loss weight for regions with a mask of 0 is forced to zero). The Adam gradient descent algorithm is used to iteratively optimize the Gaussian parameters. Since the gradients of dynamic noise points are cut off by the underlying layer, the Gaussian sphere only relies on the static background for growth and splitting. The number of iterations is set to 1000, and the learning rate is 0.001, until the loss function converges. The optimized Gaussian field is fine-tuned, and the final output is a 3DGS reconstruction model of the forest scene. This model completely eliminates dynamic artifacts and ghosting caused by wind-blown leaves and has no viewpoint blind spots or holes, making it directly applicable to forest vegetation growth monitoring and automated agricultural machinery path planning.

[0023] Figure 4 The results of 3DGS orchard reconstruction without using the method of this invention are shown, where (a) is a real image captured by a mobile device and (b) is the result reconstructed by the algorithm. Figure 5 The images show the 3DGS orchard reconstruction results using the method of this invention, where (a) is a real image captured by a mobile device, and (b) is the result reconstructed by the algorithm. Figure 4 and Figure 5 The comparison shows that the results of traditional 3DGS reconstruction ( Figure 4 Due to the complex agroforestry environment, problems such as blurred and doubled images of branches and leaves, blind spots and voids in the field of view, and uneven color distribution among multiple cameras exist. However, the method of this invention (…) Figure 5 Afterwards, thanks to mechanisms such as spatiotemporal outlier removal, viewpoint completion, and multi-camera color alignment, not only were dynamic artifacts caused by wind effectively eliminated, making the outlines of branches and leaves clear and sharp, but the structural gaps in sparse viewpoints were also filled, achieving natural unity of lighting and color across the entire scene.

Claims

1. A three-dimensional Gaussian multi-machine collaborative mapping and optimization method for complex agricultural and forestry scenarios, characterized in that, The specific steps are as follows: Step 1: Multi-machine data preprocessing: Acquire complex agricultural and forestry scene image sequences, real depth image data and pose data synchronously collected by multiple mobile devices, construct initial observation data, and perform exposure equalization and color alignment processing on the image sequences; Step 2, Spatiotemporal Joint Outlier Removal Mask Generation: Initialize a three-dimensional Gaussian field using the initial observation data; use the image sequence after exposure equalization and color alignment processing as the real color image; use the initialized three-dimensional Gaussian field and the mobile device pose corresponding to the current frame for forward projection rendering to obtain the predicted color image and the predicted depth image; The photometric spatiotemporal residual and geometric depth spatiotemporal residual are calculated by subtracting them pixel by pixel from the real color image and the real depth image. When the photometric spatiotemporal residual or geometric depth spatiotemporal residual at a certain pixel location is greater than the set static tolerance threshold, it is identified as a dynamic outlier and a pixel-level dynamic removal mask matrix is ​​generated accordingly. Step 3: Viewpoint completion based on large motion frame interpolation and depth enhancement: For sparse viewpoint regions of multi-camera observation, the Depth Anything depth estimation model is used to perform monocular depth estimation on the real frames of the preprocessed image sequence to generate a depth map. The FILM image interpolation model is used in combination with the geometric constraints of the depth map to generate virtual intermediate frames between real frame pairs that meet the viewpoint overlap rate and time interval threshold. The valid virtual intermediate frames that have passed the depth deviation verification are added to the observation dataset. Step 4, 3DGS Joint Optimization Mapping: The expanded observation dataset is integrated. In the joint optimization stage of pose tracking and mapping, the dynamic removal mask matrix is ​​multiplied by the joint loss function containing color loss and depth loss, so that the pixel loss weight of the pixel identified as a dynamic outlier is forced to zero. When calculating the gradient during backpropagation, the path of the dynamic outlier participating in the iterative update of the 3D Gaussian sphere parameters is cut off from the bottom layer, and finally a 3D Gaussian scene model without dynamic artifacts is output.

2. The method for three-dimensional Gaussian multi-machine collaborative mapping and optimization for complex agricultural and forestry scenarios according to claim 1, characterized in that, The specific process of performing exposure equalization and color alignment on the image sequence in step 1 is as follows: Calculate the average brightness and color mean of all acquired images; Based on the average brightness and color mean values, brightness stretching and gamma correction are performed on a single frame image; Extract vegetation area features from images and adjust color channel parameters to ensure consistent vegetation colors across images collected by different devices.

3. The method for three-dimensional Gaussian multi-machine collaborative mapping and optimization for complex agricultural and forestry scenarios according to claim 1, characterized in that, In step 2, when generating the pixel-level dynamic culling mask matrix, the mask value of the pixel position determined as a dynamic outlier is set to 0, and the mask value of the remaining static area is set to 1.

4. The method for three-dimensional Gaussian multi-machine collaborative mapping and optimization for complex agricultural and forestry scenarios according to claim 1, characterized in that, In step 4, the three-dimensional Gaussian sphere parameters include the center position, covariance matrix, opacity, and spherical harmonic function coefficients; the joint loss function includes a color loss term using L1 loss combined with D-SSIM and a depth loss term using L1 loss.

5. A three-dimensional Gaussian multi-machine collaborative mapping and optimization system for complex agricultural and forestry scenarios, used to implement the method described in any one of claims 1-4, characterized in that, This includes a multi-machine data preprocessing module, a spatiotemporal outlier removal module, an intermediate frame generation and completion module, and a 3DGS joint optimization mapping module, which are executed sequentially. The multi-machine data preprocessing module standardizes and normalizes the raw data of complex agricultural and forestry scenes acquired by multiple acquisition devices, eliminating the negative impacts of differences in equipment and lighting, and providing high-quality input data for subsequent modules. The aforementioned spatiotemporal outlier removal module addresses mapping artifacts caused by branch and leaf swaying and local exposure abrupt changes in complex agricultural and forestry scenes. Through a "rendering-observation" residual and dual-modal judgment mechanism, it blocks dynamic interference from the physical level, ensuring the purity of static reconstruction. The intermediate frame generation and completion module addresses the issue of sparse observation perspectives in multi-machine observations by fusing the FILM and DepthAnything models to complete sparse perspectives and increase observation density. The 3DGS joint optimization mapping module described above jointly optimizes the expanded observation dataset with the static Gaussian field after removing outliers to generate a high-precision and complete 3DGS reconstruction model of complex agricultural and forestry scenes.

6. A three-dimensional Gaussian multi-machine collaborative mapping and optimization system for complex agricultural and forestry scenarios according to claim 5, characterized in that, The multi-machine data preprocessing module is as follows: First, RGB images and GNSS / IMU pose data are simultaneously acquired in a designated complex agricultural and forestry area using multiple acquisition devices, and relevant acquisition parameters are recorded synchronously to ensure consistency of time and data. Then, the illumination and color features of the image sequence are extracted, and image sequences with balanced exposure and color alignment are obtained through contrast stretching and multi-channel color alignment operations. Subsequently, the pose data of multiple devices are time-stamped, extrinsic parameters are calibrated, and anomalies are filtered to ensure that the poses of all devices are unified in the same world coordinate system. Finally, the images are denoised and edge-enhancing optimized to enhance the details of agricultural and forestry vegetation and terrain.

7. A three-dimensional Gaussian multi-machine collaborative mapping and optimization system for complex agricultural and forestry scenarios according to claim 5, characterized in that, The spatiotemporal outlier removal module is as follows: First, forward rendering is performed using the initialized 3D Gaussian field and the mobile device pose corresponding to the current frame to obtain the predicted color and depth images; second, the predicted images are subtracted from the actual observed images at the pixel level to calculate the photometric spatiotemporal residuals and geometric depth residuals. Subsequently, a dual-modal threshold determination is introduced to identify pixels with residuals exceeding the limit as dynamic outliers that violate geometric consistency, thereby generating a dynamic removal mask matrix. Finally, in the joint optimization stage, the mask matrix is ​​inverted and directly applied to the loss function to force the loss weights of dynamic outliers to zero.

8. A three-dimensional Gaussian multi-machine collaborative mapping and optimization system for complex agricultural and forestry scenarios according to claim 5, characterized in that, The intermediate frame generation and completion module is as follows: First, the Depth Anything V2 model is used to perform high-precision monocular depth estimation and smoothing denoising on the preprocessed real frames to obtain reliable geometric constraints. Then, frame pairs with reasonable viewpoints and time intervals are selected from the multi-camera real frames as benchmarks. Virtual intermediate frames that conform to the geometric structure of the real scene are generated by using the FILM frame interpolation model in combination with depth constraints. Finally, the generated intermediate frames are subjected to depth consistency verification and quality evaluation. After removing invalid frames, the valid intermediate frames are added to the observation set to form an expanded observation dataset.

9. A three-dimensional Gaussian multi-machine collaborative mapping and optimization system for complex agricultural and forestry scenarios according to claim 5, characterized in that, The 3DGS joint optimization mapping module is described in detail below: First, all observed images, poses, and depth data are projected onto the same world coordinate system to establish the correspondence between pixels and 3D Gaussian points and correct camera distortion. Then, a joint loss function including color loss, depth loss, and transparency regularization is constructed, and the Gaussian parameters are iteratively optimized through the gradient descent algorithm to ensure that the reconstruction model matches the observed data. Next, the optimized Gaussian field is denoised, smoothed, and fine-tuned for details. Finally, a high-quality reconstruction model without dynamic artifacts and holes is output to meet the digital application needs of complex agricultural and forestry scenarios.