Unmanned aerial vehicle simultaneous localization and mapping method, device and equipment and storage medium
Patent Information
- Application Number
- CN202411690933.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2044-11-25
AI Technical Summary
[0005]基于此,有必要针对现有技术在无人机低空遥感领域中同步定位与建图效果较差的技术问题,提出了一种无人机同步定位与建图方法、装置、设备及存储介质
[0019] The proposed UAV synchronous localization and mapping method acquires image data collected by the UAV in a low-altitude environment, then performs four-dimensional Gaussian initialization based on the image data, followed by rendering the four-dimensional Gaussian map using four-dimensional Gaussian rendering technology. If new image data is acquired, camera pose estimation is obtained based on odometry, image gradient sampling methods, and the new image data. The optimal parameters of the camera pose are updated based on reprojection error. Finally, the four-dimensional Gaussian map is densified and filtered, and the map is updated based on the densified and filtered four-dimensional Gaussian map and the updated camera pose estimation. This invention ensures high-precision and robust autonomous pose estimation for UAVs in dynamic, large-scale outdoor environments, and also captures dynamic spatiotemporal features of the environment for mapping, thus meeting the real-time environmental detection and analysis needs of emergency scenarios in the field.
Smart Images

Figure CN119762664B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of UAV synchronous positioning and mapping technology, and in particular to a method, apparatus, device and storage medium for UAV synchronous positioning and mapping. Background Technology
[0002] Global extreme climate change has brought enormous challenges to forest fire prevention and control. Especially with accelerated urbanization, the interaction between human activities and forests is becoming increasingly frequent, leading to a gradual increase in the frequency and intensity of forest fires, posing a significant threat to natural ecosystems, socio-economic conditions, and public health. Firefighters are the main force in forest fire rescue. However, complex geographical environments and variable weather conditions make fire scenes highly unpredictable, making it difficult for firefighters to conduct safe and efficient firefighting operations, leading to larger-scale fires and even casualties. Therefore, it is necessary to use sensing technologies to detect forest wildfires, thereby providing comprehensive information for forest fire prevention and control. Traditional forest scene sensing technologies are mainly based on remote sensing satellites and lookout tower platforms. These technologies have limitations in their performance in detecting wildfires in the field. For example, lookout towers have limited field of view and are expensive to build. Furthermore, they are easily damaged by fires, resulting in additional maintenance costs. Remote sensing satellites have the advantage of a wide monitoring field of view. However, they also suffer from high operating costs, long replay cycles, and limited image resolution. This makes real-time monitoring of forest fires in specific high-risk areas very difficult.
[0003] Mapping technology based on real-time UAV imagery is expected to overcome the aforementioned limitations and is considered one of the most promising methods for future field monitoring. This technology acquires color and depth images of forest scenes using low-altitude UAVs, and then performs 3D reconstruction of the target scene offline using point cloud extraction and surface reconstruction techniques. These solutions can provide emergency commanders with high-precision, high-fidelity real-time forest scene perception, supporting core services such as forest fire spread prediction and risk warning, thereby helping to "fight forest fires early, extinguish them small, and extinguish them quickly," improving firefighting efficiency and reducing disaster losses. Furthermore, because it does not use expensive lidar to provide geometric information, but instead employs more economical and portable depth cameras, this technology also has a greater prospect for rapid deployment in the field of emergency rescue and disaster relief.
[0004] However, existing simultaneous localization and mapping (SMR) schemes based on depth cameras still have technical shortcomings in the field of UAV low-altitude remote sensing. Firstly, these schemes are mainly designed for indoor scenes and are difficult to perform accurate tracking and high-precision mapping in large-scale outdoor low-altitude environments. Secondly, most existing vision-based tracking schemes only utilize paired images for camera pose estimation, making it difficult to provide long-term point tracking to cope with the dynamic characteristics of outdoor environments. Thirdly, existing simultaneous mapping schemes mainly use methods such as artifact filtering to perform 3D modeling of static scenes, making it difficult to fully capture the dynamic features in the scene, which are crucial for emergency rescue analysis and simulation. Summary of the Invention
[0005] Based on this, it is necessary to address the technical problem of poor synchronous positioning and mapping performance of existing technologies in the field of UAV low-altitude remote sensing, and to propose a method, apparatus, device, and storage medium for UAV synchronous positioning and mapping. Firstly, a method for UAV synchronous positioning and mapping is provided, the method comprising:
[0006] Acquire image data collected by drones in low-altitude environments;
[0007] Perform four-dimensional Gaussian initialization based on the image data;
[0008] The four-dimensional Gaussian rendering technique is used to render the four-dimensional Gaussian map to obtain a map represented by the four-dimensional Gaussian.
[0009] If new image data is acquired, the camera pose estimate is obtained based on odometry, image gradient sampling method and the new image data, and the optimal parameters of the camera pose are updated based on the reprojection error.
[0010] The four-dimensional Gaussian is compacted and filtered, and the map represented by the four-dimensional Gaussian is updated based on the compacted and filtered four-dimensional Gaussian and the updated camera pose estimate.
[0011] Secondly, a device for simultaneous localization and mapping of unmanned aerial vehicles (UAVs) is provided, the device comprising:
[0012] The acquisition module is used to acquire image data collected by drones in low-altitude environments;
[0013] An initialization module is used to perform four-dimensional Gaussian initialization based on the image data;
[0014] The rendering module is used to render a four-dimensional Gaussian map based on four-dimensional Gaussian rendering technology to obtain a map represented by a four-dimensional Gaussian map.
[0015] The first update module is used to obtain a camera pose estimate based on odometry, image gradient sampling method and new image data if new image data is acquired, and update the optimal parameters of the camera pose based on the reprojection error.
[0016] The second update module is used to densify and filter the four-dimensional Gaussian, and update the map represented by the four-dimensional Gaussian based on the densified and filtered four-dimensional Gaussian and the updated camera pose estimate.
[0017] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described UAV synchronous positioning and mapping method.
[0018] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described UAV synchronous positioning and mapping method.
[0019] The proposed UAV synchronous localization and mapping method acquires image data collected by the UAV in a low-altitude environment, then performs four-dimensional Gaussian initialization based on the image data, followed by rendering the four-dimensional Gaussian map using four-dimensional Gaussian rendering technology. If new image data is acquired, camera pose estimation is obtained based on odometry, image gradient sampling methods, and the new image data. The optimal parameters of the camera pose are updated based on reprojection error. Finally, the four-dimensional Gaussian map is densified and filtered, and the map is updated based on the densified and filtered four-dimensional Gaussian map and the updated camera pose estimation. This invention ensures high-precision and robust autonomous pose estimation for UAVs in dynamic, large-scale outdoor environments, and also captures dynamic spatiotemporal features of the environment for mapping, thus meeting the real-time environmental detection and analysis needs of emergency scenarios in the field. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] in:
[0022] Figure 1 This is an application environment diagram of the UAV synchronous localization and mapping method in one embodiment;
[0023] Figure 2 This is a flowchart of a UAV synchronous localization and mapping method in one embodiment;
[0024] Figure 3 This is a structural block diagram of a UAV synchronous positioning and mapping device in one embodiment;
[0025] Figure 4 This is a structural block diagram of a computer device in one embodiment;
[0026] Figure 5 This is a structural block diagram of a computer device in another embodiment. Detailed Implementation
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0028] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] The UAV synchronous positioning and mapping method provided in this embodiment of the invention can be applied to, for example... Figure 1In this application environment, client 110 communicates with server 120 via a network. Server 120 can acquire image data collected by UAVs in low-altitude environments through client 110, then perform four-dimensional Gaussian initialization based on the image data, followed by rendering the four-dimensional Gaussian map using four-dimensional Gaussian rendering technology to obtain a four-dimensional Gaussian map. If new image data is acquired, camera pose estimation is obtained based on odometry, image gradient sampling methods, and the new image data. The optimal parameters of the camera pose are updated based on reprojection error. Finally, the four-dimensional Gaussian map is densified and filtered, and the map represented by the four-dimensional Gaussian map is updated based on the densified and filtered four-dimensional Gaussian map and the updated camera pose estimation. This invention can ensure high-precision and robust autonomous pose estimation of UAVs in dynamic large-scale outdoor environments, and can also capture the dynamic spatiotemporal features of the environment for mapping, in order to meet the real-time environmental detection and analysis needs of emergency scenarios in the wild. Client 110 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server 120 can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.
[0031] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a method for simultaneous localization and mapping of unmanned aerial vehicles (UAVs) according to an embodiment of the present invention includes the following steps:
[0032] Step S101: Acquire image data collected by drones in a low-altitude environment;
[0033] In this embodiment, image data can be collected in a low-altitude forest environment. Simulated environment data is used to simulate various properties that may occur in real-world scenes, efficiently testing the performance of the algorithm. Real-time data is used to verify the performance of the mapping algorithm and the realism of the simulation dataset. The main collection objects include color images and depth images from the drone's depth camera. The simulation dataset is collected using the drone simulation software Airsim on a forest environment (Electric Dreams Environment) modeled by Unreal Engine 5, at different heights ranging from 5 meters to 50 meters, with different collection routes including straight lines, circular paths, and random paths. The real-world data collection equipment consists of a DJI M350RTK drone equipped with a visible light camera and a depth camera, among other equipment. The DJI M350RTK drone is a professional-grade multi-rotor aircraft with high stability and precise positioning capabilities. The DJI M350RTK's RTK (Real-Time Kinematic) technology enables it to provide high-precision location information in low-altitude environments, which is crucial for data collection and map creation. The visible light camera and depth camera provide RGB-D images, thereby generating accurate 3D point cloud data. These data contain three-dimensional coordinate information, suitable for accurate mapping of urban features such as buildings, roads, and vegetation. Furthermore, the system is ready to operate immediately upon startup, making it applicable to various fields. Data acquisition process: The drone takes off and performs a carpet-like scan in the low-altitude forest environment, simultaneously acquiring pose information and RGB-D images. The image data is then transmitted to the synchronous positioning and mapping system in a time-stamped sequence.
[0034] Step S102: Perform four-dimensional Gaussian initialization based on the image data;
[0035] In this embodiment, for the first frame, the tracking step is skipped, and the camera pose is set to the identity matrix, that is, the current position is taken as the origin of the world coordinate system. In the mapping part, all pixels in the first frame are used to initialize a new four-dimensional Gaussian. The mean and covariance matrix ∑ of each Gaussian are parameterized as four scalars μ = (μ... x μ y μ z μ t ) and a four-dimensional covariance matrix ∑=RSS T R T Specifically, marginalizing the four-dimensional Gaussian distribution along X, Y, Z, and T dimensions yields one-dimensional Gaussian distributions in X, Y, Z, and T dimensions, respectively. The mean of these one-dimensional Gaussian distributions is μ. x μ y μ z μ tThe expression represents the mean of the numerical distribution of the four-dimensional Gaussian along the X, Y, Z, and T axes, i.e., the mean position of the Gaussian in three-dimensional space and time. R in the covariance matrix is the four-dimensional rotation matrix, and S is the scale factor. Through this parameterization, the four-dimensional Gaussian can be described by the shape and position of an ellipsoid, which can be arbitrarily rotated in both space and time. This representation method can flexibly capture complex motion changes in dynamic scenes.
[0036] Step S103: Render the four-dimensional Gaussian map using the four-dimensional Gaussian rendering technique to obtain a map represented by the four-dimensional Gaussian map;
[0037] In this embodiment, based on the four-dimensional Gaussian representation, a high-fidelity reconstructed RGB-D image is obtained using four-dimensional Gaussian rendering technology as the map of the four-dimensional Gaussian representation. Specifically, the four-dimensional Gaussian rendering technology uses a given time t (initialized to time 0) and viewpoint. Assuming there are N Gaussians in the scene, for the i-th four-dimensional Gaussian p i (x, y, z, t) is first decomposed into a conditional three-dimensional Gaussian p i (x, y, z|t) (representing the probability distribution of the Gaussian in three-dimensional space at time t) and a marginal one-dimensional Gaussian p i (t) (representing the probability distribution of the Gaussian over time). Then, the conditional 3D Gaussian is projected onto a 2D pixel plane to obtain the 2D Gaussian distribution p of the Gaussian at pixel (u, v) at time t. i (u, v|t). Finally, these N Gaussian terms are rendered to obtain the visual perspective. The RGB-D image below needs to be fused with the color c that varies with time t and viewpoint d (3D vector). i (d, t), two-dimensional Gaussian distribution p i (u, v|t), a one-dimensional time-dependent Gaussian distribution p i (t) and the opacity α of the Gaussian point i .
[0038] Step S104: If new image data is acquired, the camera pose estimate is obtained based on odometry, image gradient sampling method and the new image data, and the optimal parameters of the camera pose are updated based on the reprojection error.
[0039] In this embodiment, camera pose tracking aims to estimate the camera pose of the new incoming image data, thereby minimizing the image and depth reconstruction error of the new image data relative to the camera pose parameters within time t+1.
[0040] In one embodiment, step S104 includes: the odometry uses an image gradient sampling method to select key points of new image data, performs bidirectional key point tracking in the most recent multi-frames to obtain the camera pose, and minimizes the reprojection error through gradient optimization to update the optimal parameters for obtaining the camera pose.
[0041] In this embodiment, firstly, the odometry uses an image gradient sampling method to select key points, ensuring the uniqueness and good distribution of features. Then, the tracking model performs bidirectional key point tracking in the most recent multiple frames, using trajectory quality assessment for efficient filtering, retaining only high-quality points for optimization. The filtering is based on the uncertainty, visibility, and dynamics of the point trajectory. Next, gradient optimization is used to minimize reprojection error, updating the optimal parameters for the camera pose. A sliding window bundle adjustment can be used to optimize the camera pose and 3D point positions, ensuring the reliability of the trajectory.
[0042] Specifically, assuming the camera intrinsics are K, the current image sequence is represented as S. BA The extrinsic parameter of the i-th image is T. i The extrinsic parameter of the j-th image is T. j The position of the nth feature point in the i-th image is x. i,n The corresponding depth d i,n First, the two-dimensional pixel position (Model) of the nth feature point projected onto other images (i.e., the jth image) is obtained through long-term arbitrary point tracking visual odometry. i→j (x i,n The corresponding two-dimensional pixel position is calculated by reprojection. The distance error between the two locations is calculated. For each feature point in i, this distance error is calculated and summed with weights to obtain the reprojection error from the i-th image to the j-th image. Furthermore, S is calculated. BA The reprojection error between all image pairs in the image sequence is summed, and the formula for this reprojection error is summarized as follows:
[0043]
[0044] Among them, ||·|| ρ It is a distance metric, using Manhattan distance, with weights w. i→j,n The calculation results are derived from the long-term arbitrary point tracking visual odometry, indicating the importance of the reprojection error from the i-th image to the j-th image for pose and 3D point optimization.
[0045] Step S105: Compact and filter the four-dimensional Gaussian, and update the map represented by the four-dimensional Gaussian based on the compacted and filtered four-dimensional Gaussian and the updated camera pose estimate.
[0046] In this embodiment, after the camera pose of the newly input image frame is determined, four-dimensional Gaussian compaction and filtering are performed. The purpose of compaction is to add a new Gaussian volume to fit previously unseen scene elements or enhance the fitting effect. The purpose of filtering is to remove abnormal Gaussians and Gaussians with poor fitting effects. Regarding Gaussian filtering, due to the use of inverse depth, the calculated points will have uneven distribution and contain some outliers, which will introduce errors into the optimization. To address this problem, a selection method is proposed to sample points and remove outliers, specifically as follows: 1) Divide the image into g×g blocks. 2) Obtain the average depth d of each block. mean and standard deviation d std 3) Only retain one Gaussian with the maximum weight and a mean depth of d, where d satisfies |dd|. min |<3d std Then, other Gaussians in this block are removed. For Gaussian densification, which pixels should be densified is determined using masking and surface information guidance. For locations where the map is not dense enough, or where new geometry should precede the currently estimated geometry (i.e., the ground truth depth precedes the predicted depth, and the depth error is greater than 50 times the median depth error (MDE), a mask is added. Next, the rendering gradient value and the number of Gaussians for the current pixel are calculated. If the current pixel satisfies the conditions of having a mask added, and the rendering gradient value is higher than the median gradient value of the current keyframe and the Gaussians are dense, a new Gaussian is added following the same process as the initialization of the first frame.
[0047] After determining the camera pose and the number and location of Gaussians, real-time mapping is performed. This step aims to update the parameters of the four-dimensional Gaussian map based on the currently estimated online camera pose set, minimizing reprojection error and maximizing the model's fit to the real scene. This is accomplished through differentiable rendering and gradient-based optimization. In this step, the camera pose is fixed, and the Gaussian parameters are updated. Optimization is hot-started only from the most recently constructed map, without optimizing all previous keyframes; instead, frames that might affect the newly added Gaussians are selected. Each latest frame is selected as a keyframe, and k frames are selected for optimization, including the current frame, the most recent keyframe, and the k-2 previous keyframes with the highest overlap with the current frame. The degree of overlap is determined by obtaining the point cloud of the current frame's depth map and the number of points within the view frustum of each keyframe. The objective loss function is optimized using gradient descent in this stage. The loss function is mainly used to optimize the reconstruction and rendering quality of dynamic scenes. It combines multiple loss terms, including L1 loss and structural similarity loss (SSIM), to ensure that the difference between the generated image and the real image is minimized while maintaining image smoothness and consistency.
[0048] This invention has the following advantages: 1. Low cost: It can achieve real-time detection of low-altitude environments within 50 meters without using radar, which is cost-effective. 2. Robustness: The proposed camera pose estimation algorithm has good robustness in outdoor dynamic scenes, accurately distinguishing between static and dynamic elements in the scene, achieving high-precision and robust tracking and camera pose estimation, thus providing accurate geographic information reference for 3D models. 3. High efficiency: The proposed outdoor dynamic scene mapping algorithm can perform real-time 3D reconstruction of outdoor dynamic scenes such as forests and fields at a theoretical rendering speed of up to 400 frames per second, greatly improving the reconstruction speed and meeting the needs of rapid modeling in emergency scenarios. 4. It can provide more accurate geographic environmental information for forest fire prevention and decision-making, enhancing the intelligence and accuracy of forest fire prevention. In addition, this solution can also be used in multiple application fields such as urban planning, environmental monitoring, cultural heritage protection, and emergency response.
[0049] Please see Figure 3 As shown, in one embodiment, a UAV synchronous positioning and mapping device is provided, the device comprising:
[0050] Acquisition module 10 is used to acquire image data collected by drones in low-altitude environments;
[0051] Initialization module 20 is used to perform four-dimensional Gaussian initialization based on the image data;
[0052] The rendering module 30 is used to render a four-dimensional Gaussian map based on the four-dimensional Gaussian rendering technology to obtain a map represented by a four-dimensional Gaussian map.
[0053] The first update module 40 is used to obtain a camera pose estimate based on odometry, image gradient sampling method and new image data if new image data is acquired, and update the optimal parameters of the camera pose based on the reprojection error.
[0054] The second update module 50 is used to densify and filter the four-dimensional Gaussian, and update the map represented by the four-dimensional Gaussian based on the densified and filtered four-dimensional Gaussian and the updated camera pose estimate.
[0055] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a UAV synchronous localization and mapping method on the server side.
[0056] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a client-side method for simultaneous localization and mapping of unmanned aerial vehicles (UAVs).
[0057] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the following steps:
[0058] Acquire image data collected by drones in low-altitude environments;
[0059] Perform four-dimensional Gaussian initialization based on the image data;
[0060] The four-dimensional Gaussian rendering technique is used to render the four-dimensional Gaussian map to obtain a map represented by the four-dimensional Gaussian.
[0061] If new image data is acquired, the camera pose estimate is obtained based on odometry, image gradient sampling method and the new image data, and the optimal parameters of the camera pose are updated based on the reprojection error.
[0062] The four-dimensional Gaussian is compacted and filtered, and the map represented by the four-dimensional Gaussian is updated based on the compacted and filtered four-dimensional Gaussian and the updated camera pose estimate.
[0063] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps:
[0064] Acquire image data collected by drones in low-altitude environments;
[0065] Perform four-dimensional Gaussian initialization based on the image data;
[0066] The four-dimensional Gaussian rendering technique is used to render the four-dimensional Gaussian map to obtain a map represented by the four-dimensional Gaussian.
[0067] If new image data is acquired, the camera pose estimate is obtained based on odometry, image gradient sampling method and the new image data, and the optimal parameters of the camera pose are updated based on the reprojection error.
[0068] The four-dimensional Gaussian is compacted and filtered, and the map represented by the four-dimensional Gaussian is updated based on the compacted and filtered four-dimensional Gaussian and the updated camera pose estimate.
[0069] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0070] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0071] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0072] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for simultaneous localization and mapping of unmanned aerial vehicles (UAVs), characterized in that, The UAV synchronous positioning and mapping method includes: Acquire image data collected by drones in low-altitude environments; Perform four-dimensional Gaussian initialization based on the image data; A map represented by a four-dimensional Gaussian representation is obtained by rendering a four-dimensional Gaussian map using four-dimensional Gaussian rendering technology. If new image data is acquired, the camera pose estimate is obtained based on odometry, image gradient sampling method and the new image data, and the optimal parameters of the camera pose are updated based on the reprojection error. The four-dimensional Gaussian is compacted and filtered, and the map represented by the four-dimensional Gaussian is updated based on the compacted and filtered four-dimensional Gaussian and the updated camera pose estimate. The steps for performing four-dimensional Gaussian initialization based on the image data include: In the first frame of the image data, all pixels are used to initialize a new four-dimensional Gaussian, where the mean of each four-dimensional Gaussian is... Covariance Matrix Parameterized into four scalars and a four-dimensional covariance matrix ; in, These refer to the mean and covariance matrix of a one-dimensional Gaussian distribution obtained by marginalizing the four-dimensional Gaussian distribution along the X, Y, Z, and T axes, respectively. It is a four-dimensional rotation matrix. It is the scale scaling factor; The densification process involves determining which pixels should be densified using masks and surface information. Specifically, for locations where the map is not dense enough, or where a new geometry should exist in front of the currently estimated geometry (i.e., where the actual ground depth is ahead of the predicted depth and the depth error is greater than 50 times the median depth error), a mask is added. Next, the rendering gradient value and Gaussian density of the current pixel are calculated. If the current pixel satisfies the requirement of adding a mask, and the rendering gradient value is higher than the median gradient value of the current keyframe and the Gaussians are dense, then a new Gaussian density is added following the same process as the initialization of the first frame. The filtering method determines sampling points and removes outliers based on the following selection method: 1) Divide the image into g×g blocks; 2) Obtain the average depth of each block. and standard deviation ; 3) Only retain one Gaussian with the maximum weight and a depth-average value of d, where d satisfies And remove other Gaussians in this block.
2. The UAV synchronous positioning and mapping method according to claim 1, characterized in that, If new image data is acquired, the steps of obtaining a camera pose estimate based on odometry, image gradient sampling methods, and the new image data, and updating the optimal parameters of the camera pose based on the reprojection error, include: The odometry uses an image gradient sampling method to select key points in new image data, performs bidirectional key point tracking in the most recent multiple frames to obtain the camera pose, and minimizes the reprojection error through gradient optimization to update the optimal parameters for obtaining the camera pose.
3. A device for simultaneous localization and mapping of unmanned aerial vehicles (UAVs), characterized in that, The method for simultaneous localization and mapping of unmanned aerial vehicles (UAVs) as described in any one of claims 1-2 is employed; the UAV simultaneous localization and mapping device comprises: The acquisition module is used to acquire image data collected by the UAV in a low-altitude environment; An initialization module is used to perform four-dimensional Gaussian initialization based on the image data; The rendering module is used to render a four-dimensional Gaussian map based on four-dimensional Gaussian rendering technology to obtain a map represented by a four-dimensional Gaussian map. The first update module is used to obtain a camera pose estimate based on odometry, image gradient sampling method and new image data if new image data is acquired, and update the optimal parameters of the camera pose based on the reprojection error. The second update module is used to densify and filter the four-dimensional Gaussian, and update the map represented by the four-dimensional Gaussian based on the densified and filtered four-dimensional Gaussian and the updated camera pose estimate.
4. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the UAV synchronous positioning and mapping method as described in any one of claims 1 to 2.
5. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the UAV synchronous positioning and mapping method as described in any one of claims 1 to 2.
Citation Information
Patent Citations
Unmanned aerial vehicle navigation map construction system and method based on image three-dimensional reconstruction technology
CN111599001A
Dynamic scene rendering method and system based on three-dimensional decomposition Hash coding
CN118840471A
Environment reconstruction method and device
WO2021179745A1