A digital twin system for complex forest environments based on spatiotemporal multidimensionality and its construction method

By using a digital twin system for complex forest environments based on spatiotemporal multidimensionality, and generating dense disparity maps using binocular image datasets and LANet networks, the problem of low efficiency in traditional forest resource surveys is solved, and high-precision three-dimensional spatial modeling and robust positioning and tracking are achieved.

CN117671175BActive Publication Date: 2026-04-03NORTHEAST FORESTRY UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional forest resource survey methods are inefficient, especially in environments that are inaccessible by humans, have harsh conditions, or are dangerous, making it difficult to achieve high-precision three-dimensional spatial modeling of complex forest environments.

Method used

A digital twin system for complex forest environments based on spatiotemporal multidimensionality is adopted. By combining input module, tracking module, local mapping module, dense mapping module, closed-loop module and global BA module, the system uses a stereo image dataset for localization and tracking, keyframe selection and optimization, and combines it with LANet network to generate a dense disparity map to construct a 3D point cloud dense map.

Benefits of technology

It achieves high-precision 3D spatial modeling in complex forest environments, improves positioning and mapping accuracy, enhances the robustness of the algorithm, and meets the needs of detailed forest resource survey.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117671175B_ABST
    Figure CN117671175B_ABST
Patent Text Reader

Abstract

This invention relates to a digital twin system for complex forest environments based on spatiotemporal multidimensionality and its construction method, belonging to the field of 3D spatial modeling technology for complex forest environments. To improve the accuracy of 3D spatial modeling of complex forest environments, the input module is connected to a tracking module. The tracking module is connected to a local mapping module and a dense mapping module via keyframes. The local mapping module is connected to the dense mapping module and a loop closure module. The loop closure module is connected to a global BA module, and the global BA module is connected to the dense mapping module. This invention combines deep learning with binocular visual SLAM, using the temporal dimension as the main clue to match local sparse map points and estimate the camera's 6DoF pose. A binocular spatial dimension compensation strategy is proposed, using the spatial dimension as a secondary clue to increase the algorithm's tracking robustness. A keyframe selection strategy is designed using inter-frame relative motion and data correlation as constraints. A dense disparity map is generated using the LANet network to construct a dense 3D map, thus realizing a digital twin of the complex forest environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional spatial modeling technology for complex forest environments, specifically relating to a digital twin system for complex forest environments based on spatiotemporal multidimensionality and its construction method. Background Technology

[0002] Traditional forest resource surveys involve manual on-site investigations of forest environments, which are labor-intensive and inefficient. In particular, they can pose significant challenges to surveys in environments that are inaccessible by human intervention, have harsh conditions, or are dangerous.

[0003] In recent years, the rapid development of emerging technologies such as deep learning, machine vision, and artificial intelligence has provided technical support for the creation of digital twins of forest environments. The forest environment digital twin method uses drones to remotely acquire forest images for 3D reconstruction of forest scenes. It can realistically reproduce the complex forest environment from different perspectives, clearly and comprehensively showcasing the density, posture, and spatial location of trees, as well as the internal structure of the 3D forest environment, including the height, thickness, outline, color, and texture of tree trunks. This effectively solves the difficulties in forest resource surveys, such as limited field of vision, overlapping and obstructing trees, inaccessibility by humans, harsh conditions, and dangerous environments. It provides strong evidence for professionals to scientifically regulate tree growth, optimize forest structure, and conduct detailed surveys of forest resources such as stand volume and density. It plays a crucial role in evaluating the economic, ecological, and social value of forests. Summary of the Invention

[0004] The problem this invention aims to solve is to improve the accuracy of three-dimensional spatial modeling of complex forest environments. It proposes a digital twin system for complex forest environments based on spatiotemporal multidimensionality and its construction method.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A digital twin system for complex forest environments based on spatiotemporal multidimensional mapping includes an input module, a tracking module, a local mapping module, a dense mapping module, a closed-loop module, and a global BA module.

[0007] The input module is connected to the tracking module. The tracking module is connected to the local mapping module and the dense mapping module through keyframes. The local mapping module is connected to the dense mapping module and the loop closure module. The loop closure module is connected to the global BA module. The global BA module is connected to the dense mapping module.

[0008] The input module collects stereo image datasets, stereo bag video datasets, or real-time images from stereo cameras as input image data.

[0009] The tracking module locates and tracks each frame of the input image data transmitted by the input module, filters and outputs keyframes;

[0010] The local mapping module receives key frames input from the tracking module, removes redundant map points, generates new map points, optimizes the poses of local map points and key frames, deletes redundant key frames, and sends the optimized key frames to the closed-loop module and the dense mapping module.

[0011] The dense mapping module uses the keyframes output by the tracking module to generate a dense disparity map and establish a preliminary 3D point cloud dense map. It then uses the optimized keyframes input from the local mapping module to update the preliminary 3D point cloud dense map and optimize the keyframe pose and point cloud coordinates.

[0012] The closed-loop module receives optimized keyframes sent by the local mapping module to perform closed-loop detection and correction to eliminate accumulated drift errors. The keyframes detected and corrected in the closed loop are input into the global BA module for further optimization to obtain all 3D map points and keyframe poses.

[0013] The global BA module is used to perform BA optimization on all 3D map points and keyframe poses and update the map to obtain the final 3D point cloud dense map.

[0014] Furthermore, the tracking module finds feature points and matches them with the local map in each frame, solves the pose of the current frame using ICP or PnP methods, and optimizes the pose of the current frame by minimizing the reprojection error, thereby realizing the camera's localization and tracking in each frame. The continuity of the stereo image frames in the temporal dimension serves as the main clue for pose estimation, while the continuity in the spatial dimension serves as an auxiliary clue. The stereo image spatial dimension compensation strategy is used to increase the robustness of the algorithm's tracking, and the relative motion between frames and data correlation are used as constraints to select key frames.

[0015] Furthermore, the dense mapping module uses the LANet network to predict dense disparity maps in areas with weak texture, poor lighting, and occlusion. The dense disparity map is transformed to obtain a depth map, and combined with the key frame pose calculation to obtain a single frame 3D point cloud. A 3D point cloud dense map is generated through point cloud registration, point cloud fusion, and point cloud filtering.

[0016] A method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality, relying on the aforementioned digital twin system for complex forest environments based on spatiotemporal multidimensionality, includes the following steps:

[0017] S1. Use the ZED 2 binocular camera to collect images of planted forests in the Harbin Experimental Forest Farm of Northeast Forestry University and construct the Forest dataset;

[0018] S2. The input module uses the Forest dataset constructed in step S1 as the input image data, and sets the YAML file of the digital twin system for a complex forest environment based on spatiotemporal multidimensional settings according to the parameters calibrated by the binocular camera, so that the parameters of the system processing data are consistent with the parameters of the input image data.

[0019] S3. The tracking module receives the input image data transmitted by the input module in step S2, performs stereo matching on the left and right images of the acquired binocular camera, generates 3D map points and creates an initial map in the first frame, and starts tracking in the next frame. It uses the ICP or PnP method to solve the pose of the current frame image; then it optimizes the pose of the current frame image by minimizing the reprojection error; then it uses the continuity of the binocular image frames in the time dimension as the main clue for pose estimation and the continuity in the spatial dimension as the auxiliary clue to filter and input key frames.

[0020] S4. The local mapping module receives the keyframes input by the tracking module, removes redundant map points, generates new map points, optimizes the poses of local map points and keyframes, deletes redundant keyframes, and sends the optimized keyframes to the closed-loop module and the dense mapping module.

[0021] S5. The closed-loop module receives the optimized keyframes sent by the local mapping module, performs closed-loop detection and correction, eliminates cumulative drift error, and inputs the keyframes detected and corrected in the closed loop into the global BA module for further optimization to obtain all 3D map points and keyframe poses.

[0022] S6. The dense mapping module uses the keyframes output by the tracking module in step S3 to generate a dense disparity map and establish a preliminary 3D point cloud dense map. It uses the optimized keyframes input in the local mapping module in step S4 to update the preliminary 3D point cloud dense map, optimizes the keyframe pose and point cloud coordinates, and finally uses all the 3D map points and keyframe poses obtained in step S5 to update the 3D point cloud dense map.

[0023] Furthermore, the study area in step S1 is located at 126°37′E, 45°43′N, at an altitude of 136–140m in a mid-latitude continental monsoon climate forest. The collected images of the plantations include stereo image datasets, stereo bag video datasets, or real-time images from stereo cameras.

[0024] Furthermore, the specific implementation method of step S3 includes the following steps:

[0025] S3.1. Perform binocular initialization on the input image data transmitted by the input module in step S2. The binocular camera generates 3D map points and creates an initial map in the first frame, and starts tracking in the next frame.

[0026] S3.2. Pose Estimation:

[0027] S3.2.1. Using the continuity of the left eye image frame in the time dimension as the main clue for pose estimation, search for feature matching points in the spatial dimension frame of the right eye image corresponding to the left eye image frame. If a feature matching point of the right eye image frame is found, the pose of the current frame is estimated using the ICP method. If no feature matching point of the right eye image frame is found, proceed to the next step.

[0028] Let the 3D point clouds of two consecutive frames with corresponding data-related feature points in the time dimension be set as follows: the point cloud of the first frame is p = {p1, p2, ... p...} n The point cloud of the second frame is p′={p1′,p′2,...p′}. n}, where the point cloud p,

[0029] In an ideal case, p and p′ satisfy p. i ′=R·p i +t, where R∈SE(3), R represents the rotation matrix, t represents the translation vector, and SE(3) represents the special Euclidean group;

[0030] Due to the presence of noise, p i ′≠R·p i +t, set p and p′ to decentrifuge, and make the decentrifuged point cloud in the first frame. Second frame with centroid decentrifugation point cloud Among them, the average value of the point cloud in the first frame The average value of the point cloud in the second frame The least squares problem (R,t) is solved by constructing a solution, where R is the rotation matrix and t is the translation vector. The calculation expression is as follows:

[0031]

[0032] S3.2.2. When no matching point of the right image frame feature cannot be found in the right image spatial dimension frame, triangulation is performed through multiple frame views, and the pose of the current frame is solved by the PnP method.

[0033] Given the coordinates of n known 3D spatial points and their 2D point observations, select n known map points from the world coordinate system. As a reference point, i can be any one of n, and then four known map points are selected as control points. Control points and reference points are associated using a weighted summation method, calculated as follows:

[0034]

[0035] Where, α ijLet be the correlation coefficient between the i-th known point and the j-th control point.

[0036] set up and If the reference points and control points are in the camera coordinate system, then the reference points and control points in the camera coordinate system are associated by a weighted sum, and the calculation expression is:

[0037]

[0038] Reference point in camera coordinate system Its corresponding pixel The projection equation describes this as follows:

[0039]

[0040] Among them, w i K is the scale factor, and K is the camera intrinsic parameter;

[0041] Then we have:

[0042]

[0043] Among them, f u f is the scaled focus coordinate on the u-axis. v The scaled coordinates of the focus on the v-axis;

[0044] Coefficient matrix A 2×12 The vector h is constructed from the weighting coefficients of the reference point, pixel coordinates, and camera intrinsic parameters. 12×1 The 12 coordinate values ​​of 4 control points in the camera coordinate system and structure;

[0045] Therefore, a reference point Projected onto camera pixels u i Construct a linear equation A 2×12 ·h 12×1 =0, then there are n reference points Projected onto pixel u i Construct a system of linear equations A 2n×12 ·h 12×1 =0;

[0046] Solve the equation to find vector h. The solution process uses the least squares and SVD methods. Solve the equation to obtain the coordinate values ​​of the four control points in the camera coordinate system. Combine the known coordinate values ​​of the four control points in the world coordinate system, and then use the ICP algorithm to find the transformation relationship (R,t) between the four control points in the two coordinate systems.

[0047] S3.3. Optimize the pose of the current frame image obtained in step S3.2. Minimize the reprojection error to optimize the pose of the current frame. Project the 3D points onto the 2D points and construct the reprojection error equation with the observed 2D points. Use the Gauss-Newton optimization method to iteratively find the optimal solution.

[0048] Map points in the world coordinate system Through the transition matrix Convert to camera coordinates: The pixel coordinates are obtained by projecting the camera coordinates onto the camera model. The calculation expression is as follows:

[0049]

[0050] in, These are the estimated pixel coordinates;

[0051] Establish the relationship between world coordinates and pixel coordinates: Right now 2D point observation coordinates The error term is obtained by subtracting the values ​​of all error terms. A least-squares problem is constructed using the sum of all error terms. Then, the reprojection error is minimized, and the camera pose T is solved using the Gauss-Newton optimization algorithm. The calculation expression is as follows:

[0052]

[0053] S3.4. Keyframe Filtering:

[0054] S3.4.1. Select the inter-frame relative motion amount including rotation and displacement changes and the number of matched feature points as constraints for screening keyframes.

[0055] S3.4.2. Set the relative motion Tran(R,t) between the current frame and the previous keyframe to be a function of (R,t), then the calculation expression is:

[0056] Tran(R,t)=(1-α)||t||+αmin(2π-||R||,||R||) (8)

[0057] Where α is the inter-frame motion transformation factor:

[0058]

[0059] Where w is the rotation angle, w = min(2π - ||R||, ||R);

[0060] S3.4.3. Take the Euclidean spatial distance as the inter-frame translation and rotation changes respectively, and let α∈[0,1] as the inter-frame motion transformation factor. The value increases exponentially with the camera rotation angle. When α is large, the inter-frame relative motion depends on α. ​​The value of α and the specific expression of Tran are determined by the range of w=min(2π-||R||,||R||). The calculation expression is:

[0061]

[0062] S3.4.4. When the camera turns at... During this period, a binocular spatial dimension compensation strategy is adopted, using the corresponding right-eye image frame in the spatial dimension as a clue for connection, inserting the right-eye matching image frame to compensate for the lost field of view in the temporal dimension, and continuing tracking.

[0063] S3.4.5. The system presets a maximum threshold of η and a minimum threshold of ξ. It compares the calculated Tran with the preset thresholds: when Tran < ξ, Frame... ckey ≠Frame cur ;

[0064] When ξ≤Tran≤η, Frame ckey =Frame cur , where Frame ckey For candidate keyframes, Frame cur The current frame for the left eye;

[0065] When Tran > η, if At that time, Frame key =Frame cur ;if At that time, Frame key =Frame rcur -1;

[0066] Among them, Frame cur For the current frame of the left eye, Frame rcur -1 represents the right eye in the previous frame. ckey For candidate keyframes, Frame key For keyframes;

[0067] When the relative motion between frames Tran > η, a keyframe is inserted; when ξ ≤ Tran ≤ η, redundant keyframes are avoided, and the relative motion between frames and data association are used as constraints.

[0068] S3.4.6. Set constraints for two filtering keyframes: The constraint for the first filtering keyframe is Frame. ckeyThe number of co-visible feature points tracked between the previous and previous keyframes meets the following condition:

[0069] Track(Frame key -1,Frame ckey )>τ f (10);

[0070] The constraint for the second keyframe selection is Frame. ckey The number of tracked near points is less than the threshold τ t And create more than τ c A new proximity.

[0071] Furthermore, in step S4, the local mapping module optimizes multiple keyframes with co-view relationships and their corresponding map points, and re-matches the multiple keyframes with co-view relationships to obtain new map points.

[0072] Furthermore, the specific implementation method of step S5 includes the following steps:

[0073] S5.1. Loop Closure Detection: Use Bag-of-Words (BoW) to accelerate matching, query the dataset to see if there are loops, and then calculate the SE3 pose between the current keyframe and the loop closure candidate keyframes;

[0074] S5.2. Closed-loop correction: Corrects cumulative drift through closed-loop fusion and the Essential Graph.

[0075] Furthermore, the specific implementation method of step S6 includes the following steps:

[0076] S6.1. The dense mapping module uses step S3 to trace the keyframes output by the module to generate a dense disparity map:

[0077] The dense mapping module embeds a binocular stereo matching network LANet into D-SLAM to generate a dense disparity map. The binocular stereo matching network LANet consists of five parts: ResNet, an AM attention module, a matching cost construction module, a 3D CNN aggregation module, and a disparity prediction module. Feature extraction uses ResNet as the backbone network. The AM attention module includes SAM and CAM components, and the 3D CNN aggregation module includes a basic structure.

[0078] The basic structure of the Res_CAM_SAM_k512_Base combined module consists of 12 convolutional layers with a kernel size of 3×3×3. It performs BN and ReLU. The disparity prediction module uses the SoftArgmin function to obtain disparity estimates in a regression manner and generates a dense disparity map.

[0079] S6.2. The dense mapping module generates a dense disparity map based on the key frames sent by the tracking module according to the method in step S6.1. It calculates the pose of the key frames to obtain a single-frame 3D point cloud. Through point cloud registration, point cloud fusion and point cloud filtering, a preliminary 3D point cloud dense forest environment map is generated.

[0080] S6.3 The preliminary 3D point cloud dense forest environment map obtained in step S6.2 is optimized by using the optimized keyframes in the local mapping module to obtain a locally optimized 3D point cloud dense map.

[0081] S6.4. The locally optimized 3D point cloud dense map obtained in step S6.3 is updated by performing global BA optimization using all keyframe poses and 3D map points optimized in the global BA module, and the final 3D point cloud dense map is obtained.

[0082] The beneficial effects of this invention are:

[0083] The present invention discloses a digital twin system for complex forest environments based on spatiotemporal multidimensionality. It combines deep learning with binocular visual SLAM, designs a key frame selection strategy, estimates the camera 6DoF pose based on minimizing feature point reprojection error, and uses the LANet network to generate a dense disparity map to construct a 3D point cloud dense map to realize a digital twin of complex forest environments.

[0084] The present invention describes a digital twin system for complex forest environments based on spatiotemporal multidimensionality. It uses the temporal dimension of binocular images as the main clue to match local sparse map points to estimate the 6DoF (Degrees of Freedom) rigid body pose of the camera, and optimizes the camera pose by minimizing the visual feature reprojection error to achieve zero-drift positioning.

[0085] The present invention provides a digital twin system for complex forest environments based on spatiotemporal multidimensionality. Using the spatial dimension of binocular images as a secondary clue, a binocular spatial dimension compensation strategy is proposed. The right eye matching image frame is inserted to compensate for the lost field of view in the temporal dimension and continue tracking. This increases the robustness of the algorithm tracking when the camera turns too much, and to a certain extent solves the problem of camera tracking failure caused by excessive turning angle.

[0086] The present invention provides a digital twin system for complex forest environments based on spatiotemporal multidimensionality. Adapting to scenario requirements, it designs key frame selection strategies and corresponding algorithms with relative motion between frames and data association as constraints. This ensures the quality of key frame tracking while achieving the goal of having both constraints and less information redundancy between key frames and other key frames in the local map. This improves the system's positioning and mapping accuracy while also taking into account the running speed.

[0087] The present invention discloses a digital twin system for complex forest environments based on spatiotemporal multidimensionality. It utilizes the LANet network to predict dense disparity maps of key frames, combines key frame poses to generate single-frame dense point clouds, and uses point cloud registration, point cloud fusion, and point cloud filtering techniques to construct a dense map of complex forest spatial environments to achieve digital twin of complex forest environments.

[0088] The proposed D-SLAM system for complex forest environments based on spatiotemporal multidimensionality has been evaluated on three datasets: EuRoc, KITTI, and Forest. The results show that the proposed D-SLAM system can run at a speed of 30 normal frames and 3 key frames per second, with a positioning accuracy of several centimeters. It outperforms some mainstream models in terms of tracking accuracy and robustness, and is the most accurate SLAM solution in most cases. Its running speed and 3D dense point cloud meet the requirements for real-time positioning and dense mapping in complex forest ecosystems.

[0089] This invention discloses a digital twin system for complex forest environments based on spatiotemporal multidimensionality. It combines deep learning with binocular visual SLAM, using the temporal dimension as the main clue to match local sparse map points to estimate the camera's 6DoF pose, and the spatial dimension as a secondary clue. A binocular spatial dimension compensation strategy is proposed to increase the robustness of the algorithm's tracking. A key frame selection strategy is designed with inter-frame relative motion and data correlation as constraints. A dense disparity map is generated using the LANet network to construct a 3D point cloud dense map to realize a digital twin of complex forest environments. Attached Figure Description

[0090] Figure 1 This is a schematic diagram of the structure of a digital twin system for complex forest environments based on spatiotemporal multidimensionality, as described in this invention.

[0091] Figure 2 This is a flowchart illustrating a method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality, as described in this invention.

[0092] Figure 3 This is a schematic diagram of the right eye spatial dimension compensation of the present invention;

[0093] Figure 4 This is a schematic diagram of the network structure of LANet according to the present invention;

[0094] Figure 5 The estimated trajectory diagram of the construction method of digital twin system of forest complex environment based on spatiotemporal multidimensional forest described in this invention on EuRoC dataset, where (a) corresponds to the first sequence, (b) corresponds to the second sequence, (c) corresponds to the third sequence, (d) corresponds to the fourth sequence, (e) corresponds to the fifth sequence, (f) corresponds to the sixth sequence, (g) corresponds to the seventh sequence, (h) corresponds to the eighth sequence, and (i) corresponds to the ninth sequence;

[0095] Figure 6 The pose error of the trajectory generated on the KITTI 05 sequence according to the construction method of the digital twin system of forest complex environment based on spatiotemporal multidimensional as described in this invention is calculated by the EVO evaluation tool. Among them, (a) is the absolute pose error trajectory diagram, (b) is the numerical waveform diagram of absolute pose error (APE), (c) is the relative pose error trajectory diagram, and (d) is the numerical waveform diagram of relative pose error.

[0096] Figure 7 The image shows the results of dense 01 sequence mapping on the KITTI dataset for the construction method of a digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in this invention. (a) is the left-eye RGB image, (b) is the visualization disparity map generated by the LANet network, (c) is the feature points tracked by the system, (d) is the trajectory map estimated by the system, (e) is the sparse 3D point cloud map generated by the system, and (f) is the final dense 3D point cloud map generated by the system.

[0097] Figure 8 This invention describes a method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality, using a local 3D point cloud dense map from four perspectives in the KITTI01 sequence.

[0098] Figure 9 This is a schematic diagram of the dense mapping process in Forest, illustrating the construction method of a digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in this invention. (a) is the left-eye RGB image, (b) is the visual disparity map generated by the LANet network, (c) is the feature points tracked by the system, (d) is the trajectory map estimated by the system, (e) is the sparse 3D point cloud map generated by the system, and (f) is the final dense 3D point cloud map generated by the system.

[0099] Figure 10 This invention describes a method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality, using dense 3D point cloud maps of local areas on Forest from four viewpoints. Detailed Implementation

[0100] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention; that is, the described specific embodiments are merely a part of the embodiments of the invention, and not all of them. The components of the specific embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations, and the invention may also have other embodiments.

[0101] Therefore, the following detailed description of specific embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected specific embodiments of the invention. All other specific embodiments obtained by those skilled in the art based on these specific embodiments without inventive effort are within the scope of protection of this invention.

[0102] To further understand the invention's content, features, and effects, the following specific embodiments are provided, along with accompanying drawings. Figure 1 -Appendix Figure 10 Detailed explanation is as follows: Specific implementation method one:

[0104] A digital twin system for complex forest environments based on spatiotemporal multidimensional mapping includes an input module, a tracking module, a local mapping module, a dense mapping module, a closed-loop module, and a global BA module.

[0105] The input module is connected to the tracking module. The tracking module is connected to the local mapping module and the dense mapping module through keyframes. The local mapping module is connected to the dense mapping module and the loop closure module. The loop closure module is connected to the global BA module. The global BA module is connected to the dense mapping module.

[0106] The input module collects stereo image datasets, stereo bag video datasets, or real-time images from stereo cameras as input image data.

[0107] The tracking module locates and tracks each frame of the input image data transmitted by the input module, filters and outputs keyframes;

[0108] The local mapping module receives key frames input from the tracking module, removes redundant map points, generates new map points, optimizes the poses of local map points and key frames, deletes redundant key frames, and sends the optimized key frames to the closed-loop module and the dense mapping module.

[0109] The dense mapping module uses the keyframes output by the tracking module to generate a dense disparity map and establish a preliminary 3D point cloud dense map. It then uses the optimized keyframes input from the local mapping module to update the preliminary 3D point cloud dense map and optimize the keyframe pose and point cloud coordinates.

[0110] The closed-loop module receives optimized keyframes sent by the local mapping module to perform closed-loop detection and correction to eliminate accumulated drift errors. The keyframes detected and corrected in the closed loop are input into the global BA module for further optimization to obtain all 3D map points and keyframe poses.

[0111] The global BA module is used to perform BA optimization on all 3D map points and keyframe poses and update the map to obtain the final 3D point cloud dense map.

[0112] Furthermore, the tracking module finds feature points and matches them with the local map in each frame, solves the pose of the current frame using ICP or PnP methods, and optimizes the pose of the current frame by minimizing the reprojection error, thereby realizing the camera's localization and tracking in each frame. The continuity of the stereo image frames in the temporal dimension serves as the main clue for pose estimation, while the continuity in the spatial dimension serves as an auxiliary clue. The stereo image spatial dimension compensation strategy is used to increase the robustness of the algorithm's tracking, and the relative motion between frames and data correlation are used as constraints to select key frames.

[0113] Furthermore, the dense mapping module uses the LANet network to predict dense disparity maps in areas with weak texture, poor lighting, and occlusion. The dense disparity map is transformed to obtain a depth map, and combined with the key frame pose calculation to obtain a single frame 3D point cloud. A 3D point cloud dense map is generated through point cloud registration, point cloud fusion, and point cloud filtering.

[0114] Furthermore, the loop closure detection uses Bag-of-Words (BoW) to accelerate matching, queries the dataset to detect whether a loop is closed, and calculates the SE3 pose between the current keyframe and the loop closure candidate keyframe; the loop closure correction uses loop closure fusion and EssentialGraph optimization to correct the cumulative drift, and starts a Full BA thread to perform BA optimization on all map points and keyframes. Specific Implementation Method Two:

[0116] A method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality, relying on the digital twin system for complex forest environments based on spatiotemporal multidimensionality described in Specific Implementation Method 1, includes the following steps:

[0117] S1. Use the ZED 2 binocular camera to collect images of planted forests in the Harbin Experimental Forest Farm of Northeast Forestry University and construct the Forest dataset;

[0118] Furthermore, the study area in step S1 is located at 126°37′E, 45°43′N, at an altitude of 136–140m in a mid-latitude continental monsoon climate forest. The collected images of the plantations include stereo image datasets, stereo bag video datasets, or real-time images from stereo cameras.

[0119] S2. The input module uses the Forest dataset constructed in step S1 as the input image data, and sets the YAML file of the digital twin system for a complex forest environment based on spatiotemporal multidimensional settings according to the parameters calibrated by the binocular camera, so that the parameters of the system processing data are consistent with the parameters of the input image data.

[0120] S3. The tracking module receives the input image data transmitted by the input module in step S2, performs stereo matching on the left and right images of the acquired binocular camera, generates 3D map points and creates an initial map in the first frame, and starts tracking in the next frame. It uses the ICP or PnP method to solve the pose of the current frame image; then it optimizes the pose of the current frame image by minimizing the reprojection error; then it uses the continuity of the binocular image frames in the time dimension as the main clue for pose estimation and the continuity in the spatial dimension as the auxiliary clue to filter and input key frames.

[0121] Furthermore, the specific implementation method of step S3 includes the following steps:

[0122] S3.1. Perform binocular initialization on the input image data transmitted by the input module in step S2. The binocular camera generates 3D map points and creates an initial map in the first frame, and starts tracking in the next frame.

[0123] S3.2. Pose Estimation:

[0124] S3.2.1. Using the continuity of the left eye image frame in the time dimension as the main clue for pose estimation, search for feature matching points in the spatial dimension frame of the right eye image corresponding to the left eye image frame. If a feature matching point of the right eye image frame is found, the pose of the current frame is estimated using the ICP method. If no feature matching point of the right eye image frame is found, proceed to the next step.

[0125] Let the 3D point clouds of two consecutive frames with corresponding data-related feature points in the time dimension be set as follows: the point cloud of the first frame is p = {p1, p2, ... p...} n The point cloud of the second frame is p′={p′1,p′2,...p′}. n}, where the point cloud p,

[0126] In an ideal case, p and p′ satisfy p′. i =R·p i +t, where R∈SE(3) R represents the rotation matrix, t represents the translation vector, and SE(3) represents the special Euclidean group;

[0127] Due to the presence of noise, p′ i ≠R·p i +t, set p and p′ to decentrifuge, and make the decentrifuged point cloud in the first frame. Second frame with centroid decentrifugation point cloud Among them, the average value of the point cloud in the first frame The average value of the point cloud in the second frame The least squares problem (R,t) is solved by constructing a solution, where R is the rotation matrix and t is the translation vector. The calculation expression is as follows:

[0128]

[0129] S3.2.2. When no matching point of the right image frame feature cannot be found in the right image spatial dimension frame, triangulation is performed through multiple frame views, and the pose of the current frame is solved by the PnP method.

[0130] Given the coordinates of n known 3D spatial points and their 2D point observations, select n known map points from the world coordinate system. As a reference point, i can be any one of n, and then four known map points are selected as control points. Control points and reference points are associated using a weighted sum, calculated as follows:

[0131]

[0132] Where, α ij Let be the correlation coefficient between the i-th known point and the j-th control point.

[0133] set up and If the reference points and control points are in the camera coordinate system, then the reference points and control points in the camera coordinate system are associated by a weighted sum, and the calculation expression is:

[0134]

[0135] Reference point in camera coordinate system Its corresponding pixel The projection equation describes this as follows:

[0136]

[0137] Among them, w i K is the scale factor, and K is the camera intrinsic parameter;

[0138] Then we have:

[0139]

[0140] Among them, f u f is the scaled focus coordinate on the u-axis. v The scaled coordinates of the focus on the v-axis;

[0141] Coefficient matrix A 2×12 The vector h is constructed from the weighting coefficients of the reference point, pixel coordinates, and camera intrinsic parameters. 12×1 The 12 coordinate values ​​of 4 control points in the camera coordinate system and structure;

[0142] Therefore, a reference point Projected onto camera pixels u i Construct a linear equation A 2×12 ·h 12×1 =0, then there are n reference points Projected onto pixel u i Construct a system of linear equations A 2n×12 ·h 12×1 =0;

[0143] Solve the equation to find vector h. The solution process uses the least squares and SVD methods. Solve the equation to obtain the coordinate values ​​of the four control points in the camera coordinate system. Combine the known coordinate values ​​of the four control points in the world coordinate system, and then use the ICP algorithm to find the transformation relationship (R,t) between the four control points in the two coordinate systems.

[0144] S3.3. Optimize the pose of the current frame image obtained in step S3.2. Minimize the reprojection error to optimize the pose of the current frame. Project the 3D points onto the 2D points and construct the reprojection error equation with the observed 2D points. Use the Gauss-Newton optimization method to iteratively find the optimal solution.

[0145] Map points in the world coordinate system Through the transition matrix Convert to camera coordinates: The pixel coordinates are obtained by projecting the camera coordinates onto the camera model. The calculation expression is as follows:

[0146]

[0147] in, These are the estimated pixel coordinates;

[0148] Establish the relationship between world coordinates and pixel coordinates: Right now 2D point observation coordinates The error term is obtained by subtracting the values ​​of all error terms. A least-squares problem is constructed using the sum of all error terms. Then, the reprojection error is minimized, and the camera pose T is solved using the Gauss-Newton optimization algorithm. The calculation expression is as follows:

[0149]

[0150] S3.4. Keyframe Filtering:

[0151] A keyframe selection strategy is designed based on the complex spatial environment of the forest to avoid excessive information redundancy due to excessive image overlap. Simultaneously, the image overlap cannot be too small, ensuring a certain number of shared feature points to prevent tracking loss. Under the constraints of shared viewpoints, the goal is to guarantee keyframe tracking quality while achieving a balance between constraints and minimal information redundancy between the keyframe and other keyframes in the local map.

[0152] S3.4.1. Select the inter-frame relative motion amount including rotation and displacement changes and the number of matched feature points as constraints for screening keyframes.

[0153] S3.4.2. Set the relative motion Tran(R,t) between the current frame and the previous keyframe to a function of pose (R,t), then the calculation expression is:

[0154] Tran(R,t)=(1-α)||t||+αmin(2π-||R||,||R||) (8)

[0155] Where α is the inter-frame motion transformation factor;

[0156]

[0157] Where w is the rotation angle, w = min(2π - ||R||, ||R||);

[0158] S3.4.3. Take the Euclidean spatial distance as the inter-frame translation and rotation changes respectively, and let α∈[0,1] as the inter-frame motion transformation factor. The value increases exponentially with the camera rotation angle. When α is large, the inter-frame relative motion depends on α. ​​The value of α and the specific expression of Tran are determined by the range of w=min(2π-||R||,||R||). The calculation expression is:

[0159]

[0160] S3.4.4. When the camera turns at... During this period, the camera almost loses its associated viewpoint, and feature points cannot be matched in the temporal dimension, causing camera tracking to fail. To increase the robustness of the system's tracking, a binocular spatial dimension compensation strategy is adopted. The corresponding right-eye image frame in the spatial dimension is used as a cue for connection, and the right-eye matching image frame is inserted to compensate for the lost viewpoint in the temporal dimension, thus continuing tracking.

[0161] S3.4.5. The system presets a maximum threshold of η and a minimum threshold of ξ. It compares the calculated Tran with the preset thresholds: when Tran < ξ, Frame... ckey ≠Frame cur ;

[0162] When ξ≤Tran≤η, Frame ckey =Frame cur , where Frame ckey For candidate keyframes, Frame cur The current frame for the left eye;

[0163] When Tran > η, if At that time, Frame key =Frame cur ;if At that time, Frame key =Frame rcur -1;

[0164] Among them, Frame cur For the current frame of the left eye, Frame rcur -1 represents the right eye in the previous frame. ckey For candidate keyframes, Frame key For keyframes;

[0165] When the relative motion between frames Tran > η, a keyframe is inserted; when ξ ≤ Tran ≤ η, redundant keyframes are avoided, and the relative motion between frames and data association are used as constraints.

[0166] When the relative motion between frames, Tran > η, it indicates a significant change in the camera's field of view, and keyframes should be inserted promptly; otherwise, tracking will fail. When ξ ≤ Tran ≤ η, it represents the camera's normal motion range. In this case, redundant keyframes should be avoided, using the relative motion between frames and data correlation as constraints, combined with the candidate keyframes calculated above.

[0167] S3.4.6. Set constraints for two filtering keyframes: The constraint for the first filtering keyframe is Frame. ckey The number of co-visible feature points tracked between the previous and previous keyframes meets the following condition:

[0168] Track(Frame key -1,Frame ckey )>τ f (10);

[0169] The constraint for the second keyframe selection is Frame. ckey The number of tracked near points is less than the threshold τ t And create more than τ c A new point of proximity;

[0170] S4. The local mapping module receives the keyframes input by the tracking module, removes redundant map points, generates new map points, optimizes the poses of local map points and keyframes, deletes redundant keyframes, and sends the optimized keyframes to the closed-loop module and the dense mapping module.

[0171] Furthermore, in step S4, the local mapping module optimizes multiple keyframes with co-view relationships and their corresponding map points, and re-matches the multiple keyframes with co-view relationships to obtain new map points.

[0172] S5. The closed-loop module receives the optimized keyframes sent by the local mapping module, performs closed-loop detection and correction, eliminates cumulative drift error, and inputs the keyframes detected and corrected in the closed loop into the global BA module for further optimization to obtain all 3D map points and keyframe poses.

[0173] Furthermore, the specific implementation method of step S5 includes the following steps:

[0174] S5.1. Loop Closure Detection: Use Bag-of-Words (BoW) to accelerate matching, query the dataset to see if there are loops, and then calculate the SE3 pose between the current keyframe and the loop closure candidate keyframes;

[0175] S5.2. Closed-loop correction: Corrects cumulative drift through closed-loop fusion and the Essential Graph;

[0176] S6. The dense mapping module uses the keyframes output by the tracking module in step S3 to generate a dense disparity map and establish a preliminary 3D point cloud dense map. It uses the optimized keyframes input in the local mapping module in step S4 to update the preliminary 3D point cloud dense map, optimizes the keyframe pose and point cloud coordinates, and finally uses all the 3D map points and keyframe poses obtained in step S5 to update the 3D point cloud dense map.

[0177] Furthermore, the specific implementation method of step S6 includes the following steps:

[0178] S6.1. The dense mapping module uses step S3 to trace the keyframes output by the module to generate a dense disparity map:

[0179] The dense disparity mapping module embeds a stereo matching network LANet into D-SLAM to generate a dense disparity map. The stereo matching network LANet consists of five parts: ResNet, AM attention module, matching cost construction module, 3D CNN aggregation module, and disparity prediction module. Feature extraction uses ResNet as the backbone network. The AM attention module includes SAM and CAM. The 3D CNN aggregation module includes a basic structure Res_CAM_SAM_k512_Base combined module with 12 convolutional layers and a kernel size of 3×3×3, performing BN and ReLU. The disparity prediction module uses the SoftArgmin function to obtain disparity estimates through regression and generates a dense disparity map.

[0180] S6.2. The dense mapping module generates a dense disparity map based on the key frames sent by the tracking module according to the method in step S6.1. It calculates the pose of the key frames to obtain a single-frame 3D point cloud. Through point cloud registration, point cloud fusion and point cloud filtering, a preliminary 3D point cloud dense forest environment map is generated.

[0181] S6.3 The preliminary 3D point cloud dense forest environment map obtained in step S6.2 is optimized by using the optimized keyframes in the local mapping module to obtain a locally optimized 3D point cloud dense map.

[0182] S6.4. The locally optimized 3D point cloud dense map obtained in step S6.3 is updated by performing global BA optimization using all keyframe poses and 3D map points optimized in the global BA module, and the final 3D point cloud dense map is obtained.

[0183] Experiments and evaluations were conducted using the method described in this embodiment:

[0184] 1. Visual Odometry Localization Accuracy Experiments: The performance of D-SLAM will be evaluated on multiple sequences of two popular datasets. To demonstrate the robustness of the proposed system, the estimated camera-generated trajectory and map are compared with the real trajectory. Furthermore, the experimental results of this system are compared with some state-of-the-art SLAM systems using standard evaluation metrics. All D-SLAM experiments were run on a Dell Notebook G33590 with an Intel Core i7-9750H CPU (2.6GHz) and 16GB of memory, using only the CPU.

[0185] The estimated trajectory on the EUROC dataset is obtained as follows: Figure 5As shown, the dark blue estimated trajectories and red ground truth trajectories are displayed for nine sequences on the EuRoC dataset. By comparing with the ground truth, the method of this implementation exhibits good trajectory accuracy on these sequences.

[0186] The RMSATE error comparison uses the root mean square error RMSATE

[37] to measure accuracy. The estimated trajectory was aligned with the ground truth using the SE(3) transform, and the value is the average of 5 executions. Other results were reported by the authors of each system and compared with the GT for all frames in the trajectory. The results are shown in Table 1, where dashes indicate experimental failures:

[0187] Table 1 Comparison of RMSATE errors on the EuRoC dataset.

[0188]

[0189]

[0190] As shown in Table 1, both ORB-SLAM2 and VINS-Fusion failed to track certain parts of the V2_03_difficult sequence due to severe motion blur. Even BASALT, a binocular vision-inertial navigation system, could not complete the tracking task on this sequence because one of the cameras lost some frames. However, the system in this embodiment, by employing a binocular image spatial dimension compensation strategy, can use the right eye image to compensate for the lost field of view of the camera to some extent when the above situation occurs, and successfully tracked the difficult V2_03_difficult sequence with an error of 0.468.

[0191] 2. Error plot of the test sequences in the KITTI dataset, as shown below. Figure 6As shown, Absolute Pose Error (APE) calculates the difference between the estimated pose of the SLAM system and the true pose of the camera, and is suitable for evaluating the accuracy of the algorithm and the global consistency of the camera trajectory. Relative Pose Error (RPE)... RPE (Real-Performance Error) calculates the difference between the estimated pose change and the actual pose change at two identical time stamps, making it suitable for evaluating system drift and the local accuracy of camera trajectory. Since the KITTI dataset depicts large outdoor scenes in urban and highway environments, its overall APE error is significantly higher than that of the indoor dataset EuRoC. Larger errors occur near curves or at trajectory edges where loops do not occur. The mean error of APE in the translation direction for this sequence is 1.310438 m, the median error is 1.184211 m, the root mean square error (RMSE) is 1.469650 m, and the standard deviation (std) is 0.665299 m. Compared to APE, RPE is smaller, with a mean error of 0.015108 m, a median error of 0.013262 m, a root mean square error (RMSE) of 0.018214 m, and a standard deviation (std) of 0.010172 m.

[0192] KITTI's average relative error comparison uses average relative translation error and rotation error as performance evaluation metrics. Translation error t rel Rotational error R, expressed as a percentage. rel The translation amount is also expressed as deg / 100m, and the results are shown in Table 2:

[0193] Table 2 shows the average relative error on the KITTI dataset.

[0194]

[0195]

[0196] Table 2 shows two sequences with significant errors: the 01 sequence and the 08 sequence. Neither of these sequences exhibits loop closure. The 01 sequence is the only highway sequence in the KITTI dataset. Due to its high speed and low frame rate, few near points can be tracked in this sequence, making it difficult to estimate translation. Various methods have shown significant errors in their t-interference. rel Both are relatively large. However, because there are many distant points that can be tracked for a long time, the orientation rotation can be accurately estimated. ORB-SLAM2 is able to achieve R... rel A good error of 0.21 deg / 100m is achieved, while R... relThe error is smaller, with a value of 0.19 deg / 100m. No loop closures were observed in the 08 sequence. PL-SLAM, lacking Local Basis Allocation (BA), could not correct trajectory drift in a timely manner. While ORB-SLAM2 and Stereo LSD-SLAM performed BA in each frame, the lack of loop closures in this sequence prevented the execution of full BA, resulting in uncorrected global errors and a larger cumulative error. The system in this implementation, due to its accurate pose estimation in each preceding frame, exhibits less severe drift even without loop closure correction. D-SLAM achieved local t... rel Average 0.64m, R rel With an average accuracy of 0.20, this implementation achieves more accurate results compared to some mainstream binocular stereo systems, and has a significant advantage in most cases.

[0197] 3. Dense Mapping Section

[0198] Dense mapping on the KITTI dataset, such as Figure 7 As shown, Figure 7 This is the dense mapping effect of 0-1 sequences from the KITTI dataset. The sequence represents real-time images of a highway. Due to the high speed and low frame rate, this type of scene involves significant camera translation, minimal rotation, and no loops, presenting a considerable challenge. Figure 7 It can be observed that the projected trajectory (c) of this sequence has good accuracy in the straight sections of the highway, but there is some slight drift near the final curve. This is because the lack of closed-loop Full BA leads to an increase in the cumulative error of the trajectory, resulting in greater drift. In the sparse point cloud map (d), red dots represent the common observation points of the shared view keyframes, i.e., the reference map points, while black dots represent all map points generated by the keyframes. (f) is the overall dense point cloud map generated after the previous steps.

[0199] Figure 8 These are local dense point cloud maps from four viewing angles on the KITTI01 sequence, clearly showing details of the highway, including dashed lines, zebra crossings, tree shadows, and roadside grass. Because the KITTI dataset is a high-resolution image dataset, the stereo sensor has a long baseline, and the stereo images are corrected, resulting in relatively clear dense point cloud maps.

[0200] Dense Mapping on the Forest Dataset: Forest is a large forest scene dataset with low-texture images. Tree trunk features are not very obvious and the number of feature points is not large enough. In order to have enough nearby feature points to ensure tracking accuracy and dense mapping effect, as many keyframes as possible need to be inserted, while avoiding redundancy. A keyframe selection strategy has been designed for the characteristics of forest scenes. Figure 9It is a dense mapping process on the Forest dataset. Figure 10 It shows Figure 9 The details are observed from four different perspectives after it is rotated. Figure 10 The map can clearly reproduce the structure of the forest interior, including the density, posture, and spatial location of trees; the height, thickness, outline, color, and texture of tree trunks; and the color and density of leaves, canopy, and ground surface. These dense maps realistically reflect the appearance and structure of the forest interior, providing important data for forestry exploration.

[0201] This implementation explores accurate 6DoF (Degrees of Freedom) pose estimation of forest ecological spatial environments in a D-SLAM system using a low-cost binocular camera. The lightweight localization mode uses only the Tracking thread to track unmapped areas, achieving zero drift. A dense mapping thread is added, enabling not only sparse visual feature tracking but also the reconstruction of the internal structure of the dense 3D point cloud forest ecological spatial environment. Leveraging the advantage of rapid depth information acquisition by binocular cameras, initialization and absolute scale information can be obtained in the first frame. Keyframes are selected using inter-frame relative motion and data correlation as constraints, and a binocular image spatial dimension compensation strategy is employed to improve tracking robustness under adverse conditions such as large rotations, rapid motion, and insufficient texture. A shared view is used to control tracking and mapping within a local shared view area, unaffected by the global map size, enabling real-time operation in large scenes. An essential graph is used to optimize pose for loop closure detection, resulting in low time consumption and high accuracy. Compared to direct methods, this implementation can be used for wide baseline feature matching and is more suitable for 3D reconstruction scenes with high depth accuracy requirements. Our system runs at 30 normal frames and 3 keyframes per second, achieving a localization accuracy of several centimeters on the EUROC dataset and a local t-value on the KITTI dataset. rel Average 0.64m, R rel With an average accuracy of 0.20, the system's positioning accuracy and robustness are superior to several current mainstream methods, and it has significant advantages in most cases.

[0202] The system utilizes a consumer-grade computing platform, enabling real-time operation on a CPU. Its densely constructed maps clearly reproduce the structure of forest ecosystems, meeting the accuracy and speed requirements for UAV positioning and mapping. Furthermore, it is more reliable even under signal congestion, making it a powerful supplement and alternative to currently expensive commercial GNSS / INS navigation systems. Since pure visual SLAM is susceptible to changes in lighting and interference from movement speed, combining multi-sensor information and multi-level complementarity can compensate for these shortcomings, significantly improving the overall system performance.

[0203] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0204] Although this application has been described above with reference to specific embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of this application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in this application can be combined with each other in any way. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, this application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A digital twin system for complex forest environments based on spatiotemporal multi-dimensionality, characterized in that: It includes an input module, a tracking module, a local mapping module, a dense mapping module, a closed-loop module, and a global BA module; The input module is connected to the tracking module. The tracking module is connected to the local mapping module and the dense mapping module through keyframes. The local mapping module is connected to the dense mapping module and the loop closure module. The loop closure module is connected to the global BA module. The global BA module is connected to the dense mapping module. The input module collects stereo image datasets, stereo bag video datasets, or real-time images from stereo cameras as input image data. The tracking module locates and tracks each frame of the input image data transmitted by the input module, filters and outputs keyframes; The local mapping module receives key frames input from the tracking module, removes redundant map points, generates new map points, optimizes the poses of local map points and key frames, deletes redundant key frames, and sends the optimized key frames to the closed-loop module and the dense mapping module. The dense mapping module uses the keyframes output by the tracking module to generate a dense disparity map and establish a preliminary 3D point cloud dense map. It then uses the optimized keyframes input from the local mapping module to update the preliminary 3D point cloud dense map and optimize the keyframe pose and point cloud coordinates. The closed-loop module receives optimized keyframes sent by the local mapping module to perform closed-loop detection and correction to eliminate accumulated drift errors. The keyframes detected and corrected in the closed loop are input into the global BA module for further optimization to obtain all 3D map points and keyframe poses. The global BA module is used to perform BA optimization on all 3D map points and keyframe poses and update the map to obtain the final 3D point cloud dense map. The tracking module finds feature points and matches them with the local map in each frame, solves the pose of the current frame using ICP or PnP methods, and optimizes the pose of the current frame by minimizing the reprojection error, thus realizing the camera's localization and tracking in each frame. The continuity of the stereo image frames in the temporal dimension serves as the main clue for pose estimation, while the continuity in the spatial dimension serves as an auxiliary clue. The stereo image spatial dimension compensation strategy is used to increase the robustness of the algorithm's tracking, and the relative motion between frames and data correlation are used as constraints to select key frames.

2. The digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in claim 1, characterized in that: The dense mapping module uses the LANet network to predict dense disparity maps in areas with weak texture, poor lighting, and occlusion. The dense disparity map is transformed to obtain a depth map, and combined with the pose calculation of key frames to obtain a single frame of 3D point cloud. A 3D point cloud dense map is generated through point cloud registration, point cloud fusion, and point cloud filtering.

3. A method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality, implemented using the digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in any one of claims 1-2, characterized in that, Includes the following steps: S1. Use the ZED 2 binocular camera to collect images of planted forests in the Harbin Experimental Forest Farm of Northeast Forestry University and construct the Forest dataset; S2. The input module uses the Forest dataset constructed in step S1 as input image data, and sets the YAML file of the digital twin system for complex forest environments based on spatiotemporal multidimensional settings according to the parameters calibrated by the binocular camera, so that the parameters of the system processing data are consistent with the parameters of the input image data. S3. The tracking module receives the input image data transmitted by the input module in step S2, performs stereo matching on the left and right images of the acquired binocular camera, generates 3D map points and creates an initial map in the first frame, and starts tracking in the next frame. It uses the ICP or PnP method to solve the pose of the current frame image; then it optimizes the pose of the current frame image by minimizing the reprojection error; then it uses the continuity of the binocular image frames in the time dimension as the main clue for pose estimation and the continuity in the spatial dimension as the auxiliary clue to filter and input key frames. S4. The local mapping module receives keyframes input from the tracking module, removes redundant map points, generates new map points, optimizes the poses of local map points and keyframes, deletes redundant keyframes, and sends the optimized keyframes to the closed-loop module and the dense mapping module. S5. The closed-loop module receives the optimized keyframes sent by the local mapping module, performs closed-loop detection and correction, eliminates cumulative drift error, and inputs the keyframes detected and corrected in the closed loop into the global BA module for further optimization to obtain all 3D map points and keyframe poses. S6. The dense mapping module uses the keyframes output by the tracking module in step S3 to generate a dense disparity map and establish a preliminary 3D point cloud dense map. It uses the optimized keyframes input in the local mapping module in step S4 to update the preliminary 3D point cloud dense map, optimizes the keyframe pose and point cloud coordinates, and finally uses all the 3D map points and keyframe poses obtained in step S5 to update the 3D point cloud dense map.

4. The method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in claim 3, characterized in that, The study area in step S1 is located at 126°37′E, 45°43′N, at an altitude of 136–140m in a mid-latitude continental monsoon climate forest. The collected images of the plantations include stereo image datasets, stereo bag video datasets, or real-time images from stereo cameras.

5. The method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in claim 4, characterized in that, The specific implementation method of step S3 includes the following steps: S3.

1. Perform binocular initialization on the input image data transmitted by the input module in step S2. The binocular camera generates 3D map points and creates an initial map in the first frame, and starts tracking in the next frame. S3.

2. Pose Estimation: S3.2.

1. Using the continuity of the left eye image frame in the time dimension as the main clue for pose estimation, search for feature matching points in the spatial dimension frame of the right eye image corresponding to the left eye image frame. If a feature matching point of the right eye image frame is found, the pose of the current frame is estimated using the ICP method. If no feature matching point of the right eye image frame is found, proceed to the next step. Let the 3D point clouds of two consecutive frames with corresponding data-related feature points in the time dimension be set as follows: the first frame point cloud is... The second frame point cloud is Point clouds ; In ideal circumstances and Between ,in, R represents the rotation matrix, and t represents the translation vector. Indicates a special European group; Due to the presence of noise ,set up and Decentrifuge the first frame of the point cloud. The second frame decentroid point cloud Among them, the average value of the point cloud in the first frame The average value of the point cloud in the second frame The least squares problem (R,t) is solved by constructing a least squares problem, where R is the rotation matrix and t is the translation vector. The calculation expression is as follows: (1); S3.2.

2. When no matching point of the right image frame feature cannot be found in the right image spatial dimension frame, triangulation is performed through multiple frame views, and the pose of the current frame is solved using the PnP method. Given the coordinates of n known 3D spatial points and their 2D point observations, select n known map points from the world coordinate system. As a reference point, i can be any one of n, and then four known map points are selected as control points. For j=1,2,3,4, the control points and reference points are associated using a weighted summation method, and the calculation expression is: (2) in, Let be the correlation coefficient between the i-th known point and the j-th control point. ; set up and If the reference points and control points are in the camera coordinate system, then the reference points and control points in the camera coordinate system are associated by a weighted sum, and the calculation expression is: (3) Reference point in camera coordinate system Its corresponding pixel The projection equation describes this as follows: (4) Among them, w i K is the scale factor, and K is the camera intrinsic parameter; Then we have: (5) Among them, f u f is the scaled focus coordinate on the u-axis. v The scaled coordinates of the focus on the v-axis; coefficient matrix Constructed from the weighting coefficients of the reference point, pixel coordinates, and camera intrinsic parameters, the vector... The 12 coordinate values ​​of 4 control points in the camera coordinate system , , and structure; Therefore, a reference point Projected onto camera pixels Construct a linear equation So, n reference points Projected onto pixels Construct a system of linear equations ; Solve equations to find vectors The solution process uses the least squares and SVD methods to solve the equations and obtain the coordinate values ​​of the four control points in the camera coordinate system. Combined with the known coordinate values ​​of the four control points in the world coordinate system, the ICP algorithm is used to find the transformation relationship (R,t) between the four control points in the two coordinate systems. S3.

3. Optimize the pose of the current frame image obtained in step S3.

2. Minimize the reprojection error to optimize the pose of the current frame. Project the 3D points onto the 2D points and construct the reprojection error equation with the observed 2D points. Use the Gauss-Newton optimization method to iteratively find the optimal solution. Map points in the world coordinate system Through the transition matrix Convert to camera coordinates: The camera coordinates are projected onto the camera model to obtain the pixel coordinates, and the calculation expression is as follows: (6) in, These are the estimated pixel coordinates; Establish the relationship between world coordinates and pixel coordinates: ,Right now , and the coordinates of the 2D point observation The error terms are obtained by subtracting them. A least-squares problem is constructed using the sum of all error terms. Then, the reprojection error is minimized, and the camera pose is solved using the Gauss-Newton optimization algorithm. The calculation expression is: (7) S3.

4. Keyframe Filtering: S3.4.

1. Select the inter-frame relative motion amount including rotation and displacement changes and the number of matched feature points as constraints for keyframe selection. S3.4.

2. Set the relative motion between the current frame and the previous keyframe. It is about The function is evaluated by the following expression: (8) in, Inter-frame motion transformation factor: ; in, The rotation angle is [value]. ; S3.4.

3. Take the Euclidean spatial distance as the inter-frame translation and rotation changes respectively, and let... As an inter-frame motion transformation factor, its value increases exponentially with the camera rotation angle. The range of values ​​is determined The value and The specific form of the expression, the calculation expression is: (9) S3.4.

4. When the camera turns at... During this period, a binocular spatial dimension compensation strategy is adopted, using the corresponding right-eye image frame in the spatial dimension as a clue for connection, inserting the right-eye matching image frame to compensate for the lost field of view in the temporal dimension, and continuing tracking. S3.4.

5. The system's preset maximum threshold is... Minimum threshold is For the obtained Compare with the set threshold: when hour, ; when hour, ,in These are candidate keyframes. The current frame for the left eye; when At that time, if hour, ;if hour, ; in, For the current frame of the left eye, The right eye in the previous frame. These are candidate keyframes. For keyframes; When the relative motion between frames Insert a keyframe when; At the same time, avoid redundant keyframes and use the relative motion between frames and data association as constraints. S3.4.

6. Set constraints for two filtering keyframes: The constraint for the first filtering keyframe is... The number of co-visible feature points tracked between the previous keyframe and the previous keyframe meets the following condition: (10); The constraint for the second keyframe selection is: The number of tracked near points is less than the threshold. And create more A new proximity.

6. The method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in claim 5, characterized in that, In step S4, the local mapping module optimizes multiple keyframes with co-view relationships and their corresponding map points, and re-matches the multiple keyframes with co-view relationships to obtain new map points.

7. The method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in claim 6, characterized in that, The specific implementation method of step S5 includes the following steps: S5.

1. Loop Closure Detection: Use Bag-of-Words (BoW) to accelerate matching, query the dataset to see if there are loops, and then calculate the SE3 pose between the current keyframe and the loop closure candidate keyframes; S5.

2. Closed-loop correction: Corrects cumulative drift through closed-loop fusion and the Essential Graph.

8. The method for constructing a digital twin system for complex forest environments based on spatiotemporal multidimensionality as described in claim 7, characterized in that, The specific implementation method of step S6 includes the following steps: S6.

1. The dense mapping module uses the keyframes output by the tracking module in step S3 to generate a dense disparity map: The dense disparity mapping module embeds a stereo matching network LANet into D-SLAM to generate a dense disparity map. The stereo matching network LANet consists of five parts: ResNet, AM attention module, matching cost construction module, 3DCNN aggregation module, and disparity prediction module. Feature extraction uses ResNet as the backbone network. The AM attention module includes SAM and CAM. The 3D CNN aggregation module includes a basic structure Res_CAM_SAM_k512_Base combined module with 12 convolutional layers and a kernel size of 3×3×3, performing BN and ReLU. The disparity prediction module uses the Soft Argmin function to obtain disparity estimates through regression and generates a dense disparity map. S6.

2. The dense mapping module generates a dense disparity map based on the key frames sent by the tracking module according to the method in step S6.

1. It calculates the pose of the key frames to obtain a single-frame 3D point cloud. Through point cloud registration, point cloud fusion and point cloud filtering, a preliminary 3D point cloud dense forest environment map is generated. S6.3 The preliminary 3D point cloud dense forest environment map obtained in step S6.2 is optimized by using the optimized keyframes in the local mapping module to obtain a locally optimized 3D point cloud dense map. S6.

4. The locally optimized 3D point cloud dense map obtained in step S6.3 is updated by performing global BA optimization using all keyframe poses and 3D map points optimized in the global BA module, and the final 3D point cloud dense map is obtained.

Citation Information

Patent Citations

  • Visual SLAM method based on deep learning in dynamic environment

    CN116563340A