Unmanned aerial vehicle landing positioning reconstruction method based on three-dimensional Gaussian snowball throwing neural rendering
Through the three-dimensional Gaussian snowball throwing neural rendering technology, the problem of insufficient visual autonomous landing accuracy in traditional drone navigation systems is solved, and high-precision positioning and runway reconstruction of drones in complex environments is realized.
Patent Information
- Application Number
- CN202411984177.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
In traditional drone navigation systems, GPS application scenarios are limited and easily disturbed, and inertial navigation errors have cumulative effects, making it difficult to achieve high-precision visual autonomous landing.
The drone landing positioning reconstruction method is adopted with three-dimensional Gaussian snowball throwing nerve rendering. The landing runway is represented through the ellipsoid form, the drone camera position is estimated, the loss function is established, the keyframe image is filtered, the three-dimensional Gaussian mapping optimization is carried out, the smooth loss function is constructed, and the optimal rendered image is finally obtained.
It realizes high-precision positioning and runway reconstruction of drone near-land landing in the case of GPS or inertial navigation failure, and has the advantages of passive measurement, low cost, high precision, comprehensive measurement and anti-electromagnetic interference.
Smart Images

Figure CN120047610A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for reconstructing the landing position of an unmanned aerial vehicle (UAV) by three-dimensional Gaussian snowball neural rendering, belonging to the technical fields of computer vision and image processing. Background Art
[0002] Traditional UAV navigation systems are based on GPS and inertial navigation systems, both of which have disadvantages. The application scenarios of GPS are limited and it is vulnerable to interference, and the errors of inertial navigation have a cumulative effect and the accuracy cannot be maintained. Vision-based autonomous landing is difficult to achieve in the context of the prior art by estimating the camera pose and reconstructing the runway according to the real-time landing scene images captured by the on-board camera. Summary of the Invention
[0003] The technical problem solved by the present invention is: aiming at the problem that traditional image reconstruction means in the current prior art are difficult to maintain accuracy and are limited by vision, a method for reconstructing the landing position of an unmanned aerial vehicle by three-dimensional Gaussian snowball neural rendering is proposed.
[0004] The present invention solves the above technical problems through the following technical solutions:
[0005] A method for reconstructing the landing position of an unmanned aerial vehicle by three-dimensional Gaussian snowball neural rendering, comprising:
[0006] Representing the landing runway in the form of an ellipsoid, determining the three-dimensional coordinates of the landing runway in the world coordinate system, rasterizing the landing runway at different pose perspectives, and rendering to obtain the rendered image of the landing runway;
[0007] Estimating the pose of the UAV camera during the landing process according to the rendered image;
[0008] Tracking the landing process of the UAV, and establishing a loss function by using the pose of the UAV camera during the landing process;
[0009] Screening key frame images in the rendered images of the landing process;
[0010] Traversing the rendered images of the UAV landing process, and optimizing the three-dimensional Gaussian mapping according to each key frame image;
[0011] Constructing a smooth loss function for the UAV landing runway;
[0012] Constructing a final loss function according to the smooth loss function and the loss function;
[0013] Performing three-dimensional Gaussian mapping on the landing runway according to the final loss function to obtain the optimal rendered image.
[0014] The landing runway is represented in the form of an ellipsoid as G, and its attributes include color c, semantics sem, and opacity a. The rendering process is as follows:
[0015] The landing runway is abstractly represented as an ellipsoid G with optical properties. The center mean μ of the ellipsoid G is used as the three-dimensional coordinates of the runway in the world coordinate system, and the variance Σ is used as the axis lengths of the ellipsoid centered on the three-dimensional coordinates of the runway. The landing runway is rasterized at different pose perspectives, and the 3D ellipsoid is projected onto the 2D image to be rendered through the pose, and the rendered image of the landing runway is obtained.
[0016] The method for obtaining the 2D image to be rendered is as follows:
[0017]
[0018] μ I = π(T CW ·μ W ), Σ I = JWΣ W W T J T
[0019] In the formula, i is the ellipsoid that can be projected onto the current pixel, c i is the color of the i-th ellipsoid, a i is the opacity of the i-th ellipsoid, and j is the j-th ellipsoid before the i-th ellipsoid is projected onto the pixel. μ I and μ W are the projected pixel coordinates and world coordinates of the center of the ellipsoid respectively. π represents the projection operator, T CW represents the pose of the camera, that is, the transformation matrix between the camera coordinate system and the world coordinate system. J represents the Jacobian matrix of the projection matrix, and W represents the rotation part in T CW ;
[0020] Among them, the pixel value of the 2D image to be rendered is determined by weighting the ellipse color and opacity.
[0021] The method for estimating the pose T CW of the UAV camera during the landing process based on the rendered image is as follows:
[0022]
[0023] In the formula, D represents the derivative, μ C represents the coordinates of the center of the ellipsoid in the camera coordinate system, T CW represents the pose to be solved, T ∈ SE(3), × represents the skew-symmetric matrix of the three-dimensional vector, represents the i-th column of the matrix.
[0024] The loss function is:
[0025] E pho = ||I(G, T CW ) - I||
[0026] E sem = CE||I sem (G, T CW ) - I sem ||
[0027] Wherein, I(G, T CW ) represents the RGB image obtained by rendering the ellipsoid G through T CW , I represents the RGB image observed in the current frame; I sem (G, T CW ) represents the semantic image obtained by rendering the ellipsoid G through T CW ; I sem represents the semantic image predicted in the current frame, CE is the cross-entropy loss function, E pho represents the error between the rendered image and the actually acquired image, E sem represents the semantic image error between the rendered image and the actually acquired image.
[0028] The method for screening key-frame images is as follows:
[0029] Preset the number of frames at the edge of the sliding window, select the i-th frame image and the j-th frame image in the sliding window, and determine whether it is a key frame through the Gaussian co-visibility between frames, where:
[0030] The calculation methods for the Gaussian intersection-over-union ratio and the overlap factor of the i-th frame image and the j-th frame image are as follows:
[0031]
[0032] Wherein, in the j-th frame image, when it is determined that the Gaussian intersection-over-union ratio IOU cov () < 0.9 and the overlap factor OC cov () < 0.3, the j-th frame image is a key-frame image; the traversal of frame numbers is judged by adjusting the values of i and j.
[0033] The method for optimizing the three-dimensional Gaussian mapping is as follows:
[0034] Whenever a key-frame image is generated, reconstruct the original three-dimensional Gaussian mapping, perform random sampling through the depth map predicted in the current frame, and the random sampling error follows a normal distribution. The sampling points p in the obtained region follow N(D p , 0.2σ D ), and adjust the three-dimensional Gaussian mapping according to the obtained sampling points.
[0035] The smoothing loss function is as follows:
[0036]
[0037] Wherein, for the i-th 3DGS optimization, ensure that the z values of the nearest N 3DGSs are consistent.
[0038] The final loss function is as follows:
[0039] E total = λ color E pho + λ sem E sem + λ smooth E smooth
[0040] In the formula, λ color is the color coefficient, λ sem is the loss coefficient, λ smooth is the smooth loss coefficient. Set λ color = 1, λ sem = 0.06, λ smooth = 1.
[0041] The advantages of the present invention compared with the prior art are as follows:
[0042] The method for drone landing positioning and reconstruction based on three-dimensional Gaussian snowball throwing neural rendering provided by the present invention deduces the pose backpropagation gradient chain derivative process in the three-dimensional Gaussian snowball throwing neural rendering process, and designs a real-time tracking and mapping algorithm, which can solve the positioning and runway reconstruction detection problems during the near-section landing of a drone when GPS or inertial navigation fails. Vision-based autonomous landing is based on the real-time landing scene images captured by an on-board camera, and estimates the camera pose and reconstructs the runway, with advantages such as passive measurement, low cost, high precision, comprehensive measurement, and anti-electromagnetic interference. Brief Description of the Drawings
[0043] Figure 1 is a schematic diagram of the drone positioning and reconstruction process based on three-dimensional snowball throwing neural rendering provided by the present invention. Detailed Embodiments
[0044] A method for drone landing positioning and reconstruction based on three-dimensional Gaussian snowball throwing neural rendering deduces the pose backpropagation gradient chain derivative process in the three-dimensional Gaussian snowball throwing neural rendering process, and designs a real-time tracking and mapping algorithm, which can solve the positioning and runway reconstruction detection problems during the near-section landing of a drone when GPS or inertial navigation fails.
[0045] The steps of the method for drone landing positioning and reconstruction based on three-dimensional Gaussian snowball throwing neural rendering are as follows:
[0046] Represent the landing runway in the form of an ellipsoid, determine the three-dimensional coordinates of the landing runway in the world coordinate system, rasterize the landing runway from different pose perspectives, and render to obtain the rendered image of the landing runway;
[0047] Estimate the pose of the drone camera during the landing process according to the rendered image;
[0048] Track the landing process of the UAV, and establish a loss function using the pose of the UAV camera during the landing process;
[0049] Filter key-frame images from the rendered images during the landing process;
[0050] Traverse the rendered images of the UAV landing process, and optimize the three-dimensional Gaussian mapping according to each key-frame image;
[0051] Construct a smooth loss function for the UAV landing runway;
[0052] Construct a final loss function according to the smooth loss function and the loss function;
[0053] Perform three-dimensional Gaussian mapping on the landing runway according to the final loss function to obtain the optimal rendered image.
[0054] The landing runway is represented as an ellipsoid G in the form of an ellipsoid, and its attributes include color c, semantics sem, and opacity a. The rendering process is as follows:
[0055] Abstractly represent the landing runway as an ellipsoid G with optical attributes. Take the center mean μ of the ellipsoid G as the three-dimensional coordinates of the runway in the world coordinate system, and take the variance Σ as the axis length of the ellipsoid centered on the three-dimensional coordinates of the runway. Perform rasterization of the landing runway from different pose perspectives, project the 3D ellipsoid onto the 2D image to be rendered through the pose, and obtain the rendered image of the landing runway.
[0056] The method for obtaining the 2D image to be rendered is:
[0057]
[0058] μ I = π(T CW · μ W ), Σ I = JWΣ W W T J T
[0059] In the formula, i is the ellipsoid that can be projected onto the current pixel, c i is the color of the i-th ellipsoid, a i is the opacity of the i-th ellipsoid, and j is the j-th ellipsoid before the i-th ellipsoid is projected onto this pixel. μ I and μ W are the projected pixel coordinates and world coordinates of the center of the ellipsoid respectively, π represents the projection operator, T CW represents the pose of the camera, that is, the transformation matrix between the camera coordinate system and the world coordinate system, J represents the Jacobian matrix of the projection matrix, and W represents the rotation part in T CW ;
[0060] Among them, the pixel values of the 2D image to be rendered are determined by weighting the ellipse color and opacity.
[0061] The method for estimating the UAV camera pose T during the landing process based on the rendered image is as follows: CW as follows:
[0062]
[0063]
[0064] In the formula, D represents the derivative, μ C represents the coordinates of the center of the ellipsoid in the camera coordinate system, T CW represents the pose to be solved, T ∈ SE(3), × represents the skew-symmetric matrix of a three-dimensional vector, represents the i-th column of the matrix.
[0065] The loss function is:
[0066] E pho = ||I(G, T CW ) - I||
[0067] E sem = CE||I sem (G, T CW ) - I sem ||
[0068] In the formula, I(G, T CW ) represents the RGB image obtained by rendering the ellipsoid G through T CW , I represents the RGB image observed in the current frame; I sem (G, T CW ) represents the semantic image obtained by rendering the ellipsoid G through T CW ; I sem represents the predicted semantic image in the current frame, CE is the cross-entropy loss function, E pho represents the error between the rendered image and the actually acquired image, and E sem represents the semantic image error between the rendered image and the actually acquired image.
[0069] The method for screening key-frame images is as follows:
[0070] Preset the number of frames at the edge of the sliding window, select the i-th frame image and the j-th frame image in the sliding window, and determine whether it is a key frame through the Gaussian co-visibility between frames, where:
[0071] The calculation methods for the Gaussian intersection-over-union ratio and the overlap factor of the i-th frame image and the j-th frame image are as follows:
[0072]
[0073] In the formula, in the j-th frame image, it is judged that when the current intersection over union (IOU) cov () < 0.9 and the overlap factor OC cov () < 0.3, the j-th frame image is a key frame image; by adjusting the values of i and j, the traversal of frame numbers is judged.
[0074] The method for optimizing the three-dimensional Gaussian mapping is as follows:
[0075] Whenever a key frame image is generated, the original three-dimensional Gaussian mapping is reconstructed. Random sampling is performed through the depth map predicted by the current frame, and the random sampling error follows a normal distribution. The sampling points p in the region follow N(D p , 0.2σ D ), and the three-dimensional Gaussian mapping is adjusted according to the obtained sampling points.
[0076] The smoothing loss function is:
[0077]
[0078] In the formula, for the i-th 3DGS optimization, ensure that the z values of the nearest N 3DGSs are consistent.
[0079] The final loss function is:
[0080] E total = λ color E pho + λ sem E sem + λ smooth E smooth
[0081] In the formula, λ color is the color coefficient, λ sem is the loss coefficient, λ smooth is the smoothing loss coefficient. Set λ color = 1, λ sem = 0.06, λ smooth = 1.
[0082] The following is further described in conjunction with the accompanying drawings of the specification and preferred embodiments:
[0083] In the current embodiment, based on the unmanned aerial vehicle landing positioning and reconstruction technology of visual three-dimensional Gaussian splatting (3DGS) neural rendering, the flowchart of 3DGS neural rendering is as Figure 1 shown:
[0084] (1) 3DGS neural rendering forward process
[0085] The scene representation of 3DGS neural rendering is expressed by a series of anisotropic 3D guassians (denoted as G). The properties of G include: color (c), semantics (sem), and opacity (a). The rendering process is as follows: First, the runway is abstractly represented as an ellipsoidal representation with optical properties: the mean μ (center of the ellipsoid) of the ellipsoidal G is the three-dimensional coordinates of the runway in the world coordinate system, and the variance Σ is the axis length of the ellipsoid centered on the three-dimensional coordinates of the runway. Then, rasterization is performed in the body rendering manner for different pose perspectives to obtain the rendered image, denoted as C p 。Volume rendering is to project the 3D ellipsoid onto the 2D image to be rendered according to the pose of the frame to be rendered. The pixel value of the 2D image to be rendered is obtained by weighting the ellipse color and opacity:
[0086]
[0087] μ I =π(T CW ·μ W ), Σ I =JWΣ W W T J T (2)
[0088] where i is the ellipsoid that can be projected onto the current pixel, c i is the color of the i-th ellipsoid, a i is the opacity of the i-th ellipsoid, and j is the j-th ellipsoid before the i-th ellipsoid is projected onto this pixel. μ I and μ W are the projected pixel coordinates and world coordinates of the center of the ellipsoid respectively. π represents the projection operator, T CWCW represents the pose of the camera (the transformation matrix between the camera coordinate system and the world coordinate system), J represents the Jacobian matrix of the projection matrix, and W represents the rotation part in T CW .
[0089] Since the reconstructed runway is planar, 3DGS can be simplified to 2DGS, that is, the axis length of the ellipsoid in the vertical direction is set to 0, and the ellipsoid is simplified to an ellipse, and the simplified parameters improve the operation efficiency.
[0090] (2) Camera pose estimation
[0091] The neural rendering method continuously trains the 3DGS parameters and the poses of the corresponding frame cameras through gradient descent, making the rendered images consistent with the runway images collected by the drone during landing. Therefore, how to represent the camera pose and its derivative derivation is crucial. 3DGS realizes the rasterization and gradient backpropagation processes through CUDA. This method derives the gradient propagation formula of the pose and implements the pose gradient propagation in the rasterization direction propagation in the CUDA code of 3DGS rasterization, achieving the final near-ground landing positioning and reconstruction of the drone.
[0092] The pose expression based on the Lie group SE(3) in the rasterization process is easy to calculate and project. As the tangent space expression of the Lie group, the Lie algebra can well satisfy the 6 degrees of freedom of the pose (including 3 translational degrees of freedom x, y, z and 3 rotational degrees of freedom pitch angle, yaw angle and roll angle), and converts the constrained optimization problem into an unconstrained problem. Therefore, it is used as the pose derivative for the backpropagation of 3DGS rasterization. Based on the forward process of 3DGS rendering, the pose mainly acts in the forward process of the 3DGS transformation from the world coordinate system to the camera coordinate system. Therefore, the chain derivative of the pose gradient backpropagation is also based on this process, and the derivation is as follows:
[0093]
[0094] Among them, D represents the derivative, and μ C represents the coordinates of the center of the ellipsoid in the camera coordinate system, and T CW represents the pose to be solved. Let T CW ∈SE(3), then the derivative expression of this pose matrix, which is the Lie group on the manifold, is:
[0095]
[0096] Based on the relationship between the Lie group and the Lie algebra, the derivative expressions of the pose matrix T CW in formulas (3) and (4) are:
[0097]
[0098] Among them, × represents the skew-symmetric matrix of the three-dimensional vector, represents the i-th column of the matrix.
[0099] (3) Tracking
[0100] is to perform real-time tracking on the pose of the camera. During the tracking process, only the pose of the current frame camera is optimized without updating 3DGS. This method uses the error between the rendered frame of the current frame and the actually collected frame as the loss function:
[0101] E pho = ||I(G, T CW)-I|| (7)
[0102] E sem =CE||I sem (G,T CW )-I sem || (8)
[0103] where I(G,T CW ) represents the RGB image obtained by rendering the ellipsoid G through T CW , I represents the RGB image observed in the current frame; I sem (G,T CW ) represents the semantic image obtained by rendering the ellipsoid G through T CW ; I sem represents the semantic image predicted for the current frame (predicted through the Mask2former algorithm). CE is the cross-entropy loss function.
[0104] (4) Key frame acquisition
[0105] To improve the operation efficiency, instead of using all video frames for online 3DGS and camera pose optimization, key frames are selected through a sliding window. Specifically, whether a frame is a key frame is determined by the Gaussian co-visibility between frames. Key frames can provide a wider baseline in addition to co-visibility to ensure the error of multi-view constraints. The co-visibility is defined by the Gaussian intersection over union and the overlap factor between two frames i and j:
[0106]
[0107] At the latest frame j, if IOU cov () < 0.9 and OC cov () < 0.3, then frame j is a key frame and participates in the joint optimization of pose and 3DGS.
[0108] (5) Addition and deletion of 3DGS
[0109] To ensure that there are sufficient 3DGS (point clouds) to provide constraints on co-visibility during tracking, addition and deletion of 3DGS need to be performed after a key frame is added. Specifically, whenever a new key frame comes in, the method adds 3DGS to the area that can only be seen by the current new key frame for the 3DGS reconstruction of the original 3DGS and optimizes and deletes the original 3DGS. For the addition of 3DGS, when a new key frame is generated, through the depth map predicted by the current frame, random sampling is performed on the depth map, and the random sampling error is assumed to follow a normal distribution, that is, the sampling points p in the visible area follow N(D p , 0.2σ D) For the newly added points in the non-overlapping field of view, calculate their depths using a normal distribution with a higher variance. The points sampled in the upsampled depth map are the added 3DGSs. For the deletion of 3DGSs, if the 3DGS of a newly added point cannot be observed by three frames other than the current frame, it will be deleted.
[0110] (6) Mapping
[0111] After continuously completing the pose tracking and key-frame optimization between frames in the front end, the runway reconstruction process (mapping) in the back end is carried out in parallel. The loss function for mapping is the same as the loss function during tracking (see formulas (7)(8)). During mapping, both the 3DGSs and the poses of all key frames are optimized based on color and semantics simultaneously. For the runway where the drone lands, prior information on ground smoothing is added, and a smoothing loss function is added:
[0112]
[0113] For the i-th 3DGS, the ground heights of the N nearest 3DGSs around it should be consistent (the runway is flat). Based on the above loss function, the weighted final loss function for mapping is:
[0114] E total =λ color E pho +λ sem E sem +λ smooth E smooth (12)
[0115] According to experience, set λ color =1, λ sem =0.06, λ smooth =1. Use the Adam optimizer to optimize the camera pose and 3DGS parameters.
[0116] Through the above steps, the camera pose tracking and runway reconstruction of the drone during landing based on 3DGS are completed.
[0117] Vision-based autonomous landing is based on the real-time landing scene images captured by the on-board camera. By estimating the camera pose and reconstructing the runway, it has the advantages of passive measurement, low cost, high precision, comprehensive measurement, and anti-electromagnetic interference.
[0118] Although the present invention has been disclosed above in preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications to the technical solution of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and decorations made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the protection scope of the technical solution of the present invention.
[0119] The content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.
Claims
1. A three-dimensional Gaussian snowball neural rendering method for UAV landing positioning reconstruction, characterized in that include: The landing runway is represented in the form of an ellipsoid, the three-dimensional coordinates of the landing runway in the world coordinate system are determined, the landing runway is rasterized at different postures and viewing angles, and a rendering image of the landing runway is obtained by rendering; Estimate the drone camera pose during landing based on the rendered image; Track the landing process of the drone and use the drone camera position during the landing process to establish the loss function; Filter keyframe images in the rendered images of the landing process; Traverse the rendered images of the drone landing process and perform three-dimensional Gaussian mapping optimization based on each key frame image; Construct a smooth loss function for the drone landing runway; Construct the final loss function based on the smoothed loss function and the loss function; The landing runway is 3D-Gaussian mapped according to the final loss function to obtain the optimal rendering image.
2. The method for reconstructing the landing position of a UAV based on three-dimensional Gaussian snowball neural rendering according to claim 1, characterized in that: The landing runway is represented by an ellipsoid G, and its attributes include color c, semantic sem, and opacity a. The rendering process is: The landing runway is abstractly represented as an ellipsoid G with optical properties. The central mean μ of the ellipsoid G is used as the three-dimensional coordinates of the runway in the world coordinate system, and the variance Σ is used as the axis length of the ellipsoid centered on the three-dimensional coordinates of the runway. The landing runway is rasterized at different postures and perspectives, and the 3D ellipsoid is projected onto the 2D image to be rendered through the posture to obtain the rendered image of the landing runway.
3. The method for reconstructing the landing position of a UAV based on three-dimensional Gaussian snowball neural rendering according to claim 1, characterized in that: The method for obtaining the 2D image to be rendered is: m I =π(T CW ·m W ),S I =JWΣ W W T J T Where i is the ellipsoid that can be projected to the current pixel, c i is the color of the i-th ellipsoid, a i is the opacity of the i-th ellipsoid, and j is the j-th ellipsoid projected before the i-th ellipsoid is projected to the pixel. I and μ W are the projected pixel coordinates and world coordinates of the ellipsoid center, π represents the projection operator, T CW represents the camera's position, that is, the transformation matrix from the camera coordinate system to the world coordinate system, J represents the Jacobian matrix of the projection matrix, and W represents T CW The rotating part in The pixel value of the 2D image to be rendered is determined by weighting the ellipse color and opacity.
4. The method for reconstructing the landing position of a UAV based on three-dimensional Gaussian snowball neural rendering according to claim 3 is characterized in that: Estimate the drone camera pose T during landing based on the rendered image CW The method is: In the formula, D represents the derivative, μ C represents the coordinates of the ellipsoid center in the camera coordinate system, T CW represents the pose to be solved, T∈SE(3), × represents the antisymmetric matrix of the three-dimensional vector, Represents the i-th column of the matrix.
5. The method for reconstructing the landing position of an unmanned aerial vehicle using three-dimensional Gaussian snowball neural rendering according to claim 3 is characterized in that: The loss function is: E pho =||I(G,T CW )-I|| It is sem =CE||I sem (G,T CW )-I sem || In the formula, I(G,T CW ) indicates that the CW The RGB image obtained by rendering the ellipsoid G, I represents the RGB image observed in the current frame; I sem (G,T CW ) indicates that the CW The semantic image obtained by rendering the ellipsoid G; I sem represents the semantic image predicted by the current frame, CE is the cross entropy loss function, E pho Represents the error between the rendered image and the actual captured image, E sem Represents the semantic image error between the rendered image and the actual captured image.
6. The method for reconstructing the landing position of a UAV based on three-dimensional Gaussian snowball neural rendering according to claim 3 is characterized by: The method for filtering key frame images is: The number of frames at the edge of the sliding window is preset, and the i-th frame image and the j-th frame image in the sliding window are selected. The Gaussian common view between the frames is used to determine whether they are key frames, where: The calculation method of Gaussian intersection and union ratio and overlap factor of the i-th frame image and the j-th frame image is: In the formula, in the j-th frame image, the intersection and union (IOU) is determined. cov ()<0.9, overlap factor OC cov When ()<0.3, the j-th frame image is the key frame image; the traversal frame number judgment is realized by adjusting the values of i and j.
7. The method for reconstructing the landing position of a UAV based on three-dimensional Gaussian snowball neural rendering according to claim 3 is characterized by: The method for optimizing the three-dimensional Gaussian mapping is: Whenever a key frame image is generated, the original 3D Gaussian map is reconstructed, and the depth map predicted by the current frame is randomly sampled, and the random sampling error obeys the normal distribution. The sampling point p of the acquisition area obeys N(D p , 0.2σ D ), and adjust the three-dimensional Gaussian map according to the obtained sampling points.
8. The method for reconstructing the landing position of an unmanned aerial vehicle using three-dimensional Gaussian snowball neural rendering according to claim 7 is characterized in that: The smoothing loss function is: Where, for the i-th 3DGS optimization, ensure that the z values of the most recent N 3DGSs remain consistent.
9. The method for reconstructing the landing position of a UAV based on three-dimensional Gaussian snowball neural rendering according to claim 8, characterized in that: The final loss function is: E total =λ color E pho +λ sem E sem +λ smooth E smooth In the formula, λ color is the color coefficient, λ sem is the loss coefficient, λ smooth is the smoothing loss coefficient, set λ color =1,λ sem =0.06,λ smooth =1.