Unmanned aerial vehicle indoor positioning method and system based on three-dimensional Gaussian splash rendering scene

The 3D Gaussian Sphere rendering method addresses the challenges of indoor drone positioning by using cloud-based feature matching and pose estimation with RGB-D images, achieving high-precision positioning with reduced computational requirements.

CN120318481AActive Publication Date: 2025-07-15CHINA ORDNANCE SCI INST

Patent Information

Application Number
CN202510810577.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In indoor environments, traditional drone positioning methods such as visual SLAM, LiDAR-SLAM and wireless signal-assisted positioning have problems such as complex calculations, high cost, limited accuracy or susceptible to environmental interference, especially in the absence of initial position estimation, which is difficult to achieve high-precision positioning.

Method used

Using a three-dimensional Gaussian splash rendering method, a 2D voxelized 3DGS model is deployed in the cloud, and a panoramic view is constructed using RGB-D images, combining Superpoint feature matching and PNP optimization to achieve high-precision positioning without initial pose estimation.

Benefits of technology

It realizes efficient and accurate positioning of drones under low computing resource requirements, reduces dependence on high-performance equipment, improves real-time and accuracy of positioning, and supports multi-machine collaborative tasks and downstream applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318481A_ABST
    Figure CN120318481A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle indoor positioning method and system based on a three-dimensional Gaussian splash rendering scene. The method comprises the steps that a cloud deploys a 2D voxelized 3DGS model, reconstructs a scene and configures unmanned aerial vehicle communication; the unmanned aerial vehicle rotates at a fixed point to shoot 12 RGB-D images, and an image library of a panoramic view is constructed; uploading the internal reference of the unmanned aerial vehicle camera and the acquired image library to a cloud; judging an area where the unmanned aerial vehicle is located according to the uploaded image library; selecting an optimal query image from the image database based on the feature density, and obtaining initial pose estimation through feature matching of the adjacent 3DGS model and the image database; the image under the current pose estimation is rendered through 3DGS, and updated pose estimation is obtained through feature matching and PNP optimization iteration; and transmitting the updated pose estimation calculated by the cloud to the unmanned aerial vehicle terminal to obtain unmanned aerial vehicle positioning. According to the invention, high-precision and low-resource-demand unmanned aerial vehicle positioning without initial pose estimation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robot positioning, and particularly to a method and system for indoor positioning of an unmanned aerial vehicle (UAV) based on a three-dimensional Gaussian splash rendering scene. Background Art

[0002] In the past few years, with the rapid development of the UAV industry and artificial intelligence, great progress has been made in the research and practice of UAVs. However, in an indoor environment, due to the weakening or even complete loss of GPS signals, traditional GNSS-based positioning methods are not applicable, and a large number of solutions such as SLAM require an initial pose estimate. How to achieve UAV positioning without a rough initial pose estimate in an indoor environment has become the research focus.

[0003] In the prior art, the mainstream indoor positioning methods include visual SLAM, LiDAR-SLAM, visual positioning based on feature point matching, and wireless signal assisted positioning. Among them, the visual SLAM method relies on a map constructed in real time, with complex calculations and high requirements for computing resources, especially in large-scale scenarios, there are cumulative errors; the LiDAR-SLAM method relies on high-precision lidar equipment, with high costs, and is prone to false matching in textureless environments; the wireless signal assisted positioning method is greatly affected by the environment, and the positioning accuracy is limited. Summary of the Invention

[0004] In view of this, the purpose of the embodiments of the present invention is to provide a method and system for indoor positioning of a UAV based on a three-dimensional Gaussian splash rendering scene, which can achieve high-precision and low-resource-demand UAV positioning without an initial pose estimate, not only eliminates the need for the UAV to high-computing-power devices, but also can share the reconstructed 3DGS scene to support other downstream tasks. The present invention aims to complete the initialization positioning task in the scene where 3DGS reconstruction has been completed through LiDAR and a camera. This framework uses 3DGS to represent the three-dimensional space scene instead of the traditional point cloud model, can render the scene picture more precisely, and optimize the positioning accuracy. At the same time, an architecture of cloud storage, calculation, terminal collection, and query is designed to avoid excessive requirements for the memory and computing power of the UAV terminal, and also enables the positioning task to be executed on different UAV terminals. Compared with traditional positioning methods such as SLAM, this method does not require the position information of the UAV at the previous moment, and only uses RGB-D images to align the position of the UAV with the reconstructed scene. The implementation of the present invention not only provides an efficient and accurate UAV positioning solution, but also provides strong support for the technological progress and industrial development of related fields.

[0005] In the first aspect, an embodiment of the present invention provides a method for indoor positioning of a UAV based on a three-dimensional Gaussian splash rendering scene, which includes:

[0006] S100, Deploy the 2D voxelized 3DGS model in the cloud for scene reconstruction and configure the drone communication.

[0007] S200, The drone takes 12 RGB-D images by rotating at a fixed point to build an image library of the panoramic view.

[0008] S300, Upload the internal parameters of the drone camera and the obtained image library to the cloud.

[0009] S400, The cloud determines the area where the drone is located based on the uploaded image library.

[0010] S500, The cloud selects the optimal query image from the image library based on the feature density, and obtains the initial pose estimate through the feature matching between the adjacent 3DGS model and the image database.

[0011] S600, The cloud renders the image under the current pose estimate through 3DGS, and obtains the updated pose estimate through feature matching and PNP optimization iteration.

[0012] S700, Transmit the updated pose estimate calculated by the cloud to the drone terminal to obtain the drone positioning.

[0013] The present invention has no additional requirements for the memory and computing power of the drone, and realizes the efficient positioning task without initial pose estimate through cloud computing.

[0014] Combined with the first aspect, the embodiment of the present invention provides the first possible implementation manner of the first aspect. Among them, the step of deploying the 2D voxelized 3DGS model in the cloud for scene reconstruction and configuring the drone communication in S100 includes:

[0015] S110, Import the parameters of the 3DGS model into the cloud and save them in the form of 2D voxels. Among them, the training of the parameters of the 3DGS model is obtained by collecting, transmitting and calculating the scene images in advance.

[0016] S120, Import the semantic table of the objects in the scene corresponding to each voxel grid into the cloud to obtain the reconstructed scene.

[0017] S130, Configure the drone wireless communication module in the cloud.

[0018] Combined with the first aspect, the embodiment of the present invention provides the second possible implementation manner of the first aspect. Among them, the step of the drone taking 12 RGB-D images by rotating at a fixed point to build an image library of the panoramic view in S200 includes:

[0019] S210, The drone keeps the spatial coordinates unchanged and rotates one circle along the horizontal direction.

[0020] S220, save the images captured by the drone camera and the depth information obtained by the sensor every 30°, forming 12 RGB-D images with panoramic views, and obtaining an image library of panoramic views.

[0021] Combined with the first aspect, the embodiment of the present invention provides a third possible implementation manner of the first aspect. Among them, uploading the internal parameters of the drone camera and the obtained image library to the cloud in S300 includes:

[0022] S310, upload the internal parameters of the drone camera to the cloud.

[0023] S320, upload the RGB-D image and the corresponding rotation angle to the cloud.

[0024] Combined with the first aspect, the embodiment of the present invention provides a fourth possible implementation manner of the first aspect. Among them, the cloud determines the area where the drone is located according to the uploaded image library in S400 includes:

[0025] S410, the cloud identifies the objects in the RGB-D image and saves the semantic information of the objects.

[0026] S420, compare and identify the semantic categories of the objects obtained in the RGB-D image with the object semantic tables of each voxel region, and select the candidate 3DGS regions.

[0027] Combined with the first aspect, the embodiment of the present invention provides a fifth possible implementation manner of the first aspect. Among them, the cloud selects the optimal query image from the image library based on feature density, and obtains the initial pose estimation through the feature matching between the neighboring 3DGS model and the image database includes:

[0028] S510, the cloud extracts the Superpoint features of the RGB-D images at 12 viewpoints from the image library, and selects the viewpoint with the densest features as the query image.

[0029] S520, extract the 3DGS parameters and the corresponding image database in the candidate 3DGS regions.

[0030] S530, brute-force search the corresponding image database and perform feature matching with the query image, and select the pose corresponding to the closest database image as the initial pose estimation.

[0031] Combined with the first aspect, the embodiment of the present invention provides a sixth possible implementation manner of the first aspect. Among them, the cloud renders the image under the current pose estimation through 3DGS, and obtains the updated pose estimation through feature matching and PNP optimization iteration includes:

[0032] S610, the cloud renders the image under the current pose estimation through 3DGS.

[0033] S620, extract the feature points of the rendered image, match the feature points with the query image, and adopt the PNP algorithm to refine and iterate to obtain an updated pose estimation.

[0034] Combined with the first aspect, the embodiment of the present invention provides a seventh possible implementation manner of the first aspect, wherein, the obtaining of the UAV positioning by transmitting the updated pose estimation calculated by the cloud to the UAV terminal in S700 includes:

[0035] S710, transmit the updated pose estimation calculated by the cloud and the angle of the query image to the UAV terminal.

[0036] S720, the UAV calculates the current position and camera pose of the UAV according to the camera position parameters and the query image angle.

[0037] Combined with the first aspect, the embodiment of the present invention provides an eighth possible implementation manner of the first aspect, wherein,

[0038] The UAV communication supports data sharing between different UAV terminals.

[0039] The reconstructed scene is constructed once and can also be used for other downstream tasks other than positioning.

[0040] In the second aspect, the embodiment of the present invention further provides a UAV indoor positioning system based on a three-dimensional Gaussian splash rendering scene, which includes:

[0041] A scene reconstruction module, used to deploy a 2D voxelized 3DGS model in the cloud for reconstructing the scene and configuring UAV communication.

[0042] An image library construction module, used for the UAV to rotate and shoot 12 RGB-D images at a fixed point to construct an image library of a panoramic view.

[0043] An image parameter acquisition module, used to upload the internal parameters of the UAV camera and the obtained image library to the cloud.

[0044] A region judgment module, used for the cloud to judge the region where the UAV is located according to the uploaded image library.

[0045] An initial pose estimation module, used for the cloud to select an optimal query image from the image library based on the feature density, and obtain an initial pose estimation through feature matching of the adjacent 3DGS model and the image database.

[0046] The updated pose estimation module is used for the cloud to render the image under the current pose estimation through 3DGS, and obtain the updated pose estimation through feature matching and PNP optimization iteration.

[0047] The positioning implementation module is used to transmit the updated pose estimation calculated by the cloud to the UAV terminal to obtain UAV positioning.

[0048] The beneficial effects of the embodiments of the present invention are:

[0049] Traditional visual SLAM or image matching-based positioning methods usually require initial pose estimation. The present invention can quickly obtain the accurate position and attitude of the UAV in the case of unknown initial pose through panoramic RGB-D image acquisition, 3DGS model matching, and PNP optimization, improving the adaptability and stability of autonomous positioning.

[0050] Compared with traditional point cloud and voxel grid reconstruction methods, the 3DGS of the present invention represents the scene with Gaussian ellipsoids, can quickly achieve high-quality rendering, effectively improving the real-time performance and accuracy of indoor positioning. The reconstruction of the 3DGS scene is once-only, supporting multiple UAV terminals to share the same scene model, not only improving the efficiency of multi-UAV collaborative tasks, but also providing high-quality scene information for other downstream applications.

[0051] The present invention adopts a cloud computing architecture, handing over computationally intensive tasks (such as feature extraction, matching, and pose optimization) to the server for processing, enabling the UAV terminal to achieve high-precision positioning without a high-performance processor, and being applicable to small UAVs with limited computing resources. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a flowchart of the UAV indoor positioning method based on three-dimensional Gaussian splash rendering of the scene of the present invention. Detailed Embodiments

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0055] Please refer to Figure 1 The first embodiment of the present invention provides a method for indoor positioning of a UAV based on a three-dimensional Gaussian splash rendering scene, which includes: S100, deploying a 2D voxelized 3DGS model on the cloud to reconstruct the scene and configure UAV communication; S200, the UAV rotates at a fixed point to shoot 12 RGB-D images to build an image library of a panoramic view; S300, uploading the internal parameters of the UAV camera and the acquired image library to the cloud; S400, the cloud determines the area where the UAV is located based on the uploaded image library; S500, the cloud selects the optimal query image from the image library based on feature density, and obtains an initial pose estimate by matching the features of the neighboring 3DGS model and the image database; S600, the cloud renders the image under the current pose estimate through 3DGS, and obtains an updated pose estimate through feature matching and PNP optimization iteration; S700, transmitting the updated pose estimate calculated by the cloud to the UAV terminal to obtain UAV positioning.

[0056] The present invention has no additional requirements on the memory and computing power of the drone, and realizes efficient non-initial pose estimation and positioning tasks through cloud computing.

[0057] Among them, the cloud-based deployment of a 2D voxelized 3DGS model in S100 is used to reconstruct the scene, and the configuration of drone communication includes: S110, importing the parameters of the 3DGS model to the cloud, and saving them in the form of 2D voxels, so that it can efficiently render large-scale scenes with less memory requirements, wherein the training of the parameters of the 3DGS model is obtained by the prior collection, transmission and calculation of the scene image; S120, importing the semantic table of objects in the scene corresponding to each voxel grid to the cloud to obtain the reconstructed scene; S130, configuring the drone wireless communication module to the cloud, connecting the drone to the local area network of the current scene server, and ensuring that images and other data can be transmitted normally.

[0058] Among them, the 3DGS training process for map representation and visual relocalization is: first, create a color point cloud map from LiDAR scans, images and postures, which is used as the initialization of 3DGS, and then perform incremental training on the initialization map. In a large-scale spatial environment, the 3DGS map is stored as a 2D voxel map, using a KD tree architecture to support fast spatial queries. For each sub-map, that is, a voxel containing 3DGS, the category semantic features of the instance are extracted by extracting the object category features from the RGB image. And save the confidence of each semantic and store it in the voxel as an additional attribute outside the 3DGS parameters. Used for subsequent rough position determination.

[0059] Among them, for the 12 RGB-D images captured by the UAV during fixed-point rotation in S200, the construction of the panoramic view image library includes: S210, the UAV keeps the spatial coordinates unchanged and rotates one circle horizontally; S220, save the images captured by the UAV camera and the depth information obtained by the sensor every 30°, forming 12 RGB-D images with a panoramic view, and obtaining the panoramic view image library. After rotating 30° horizontally, use lidar and camera data to record the perspective information in the current direction and calculate the RGB-D image.

[0060] Among them, for uploading the internal parameters of the UAV camera and the obtained image library to the cloud in S300, it includes: S310, upload the internal parameters of the UAV camera to the cloud, modify the camera internal parameters in the configuration file according to the camera model and upload them. If the camera internal parameters are unknown, the camera internal parameters can be obtained through camera calibration; S320, upload the RGB-D images and the corresponding rotation angles to the cloud.

[0061] Among them, for the cloud to judge the area where the UAV is located according to the uploaded image library in S400, it includes: S410, the cloud identifies the objects in the RGB-D images and saves the semantic information of the objects; S420, compare and identify the semantic categories of the objects obtained in the RGB-D images with the object semantic tables of each voxel area, and select the candidate 3DGS areas.

[0062] First, use the original pose data to accurately locate the query position on the global map. This data may come from various sources. Using the original pose as a reference, we retrieve the global 3DGS map voxels that are most likely to contain the exact position of the query image. For the case where the current 3DGS voxel is not clear, a method based on object category matching is used to select candidate 3DGS voxels. Based on the YOLOv8 object detection method, semantic segmentation and object segmentation tasks are performed on the given RGB image. Compare with the object categories and confidences saved in the voxels with 3DGS parameters, and select the voxel with the highest coincidence degree as the candidate area.

[0063] Among them, for the cloud to select the optimal query image from the image library based on feature density and obtain the initial pose estimation through feature matching between the adjacent 3DGS model and the image database in S500, it includes: S510, the cloud extracts the Superpoint features of the RGB-D images at 12 perspectives from the image library, and selects the perspective with the densest features as the query image; S520, extract the 3DGS parameters and the corresponding image database in the candidate 3DGS area; S530, perform brute-force search on the corresponding image database and perform feature matching with the query image, and select the pose corresponding to the closest database image as the initial pose estimation.

[0064] For further tasks, only one RGB-D image is required to complete the positioning task. Therefore, the best view for positioning needs to be selected from the panoramic view. Denser Superpoint features mean the richness of the geometric structure and color of the image. Therefore, the view with the densest Superpoint feature points is selected as the query view for positioning.

[0065] For the 3DGS voxels of the candidate region, in addition to containing 3DGS parameters for scene reconstruction, they also contain a large amount of image data (which can also be obtained through 3DGS rendering). Brute-force search the image data in the database, match the query image with it, use the normalized cross-correlation value as the measurement criterion, select the image closest to the query image, and read its corresponding pose. Some code examples are as follows:

[0066] # Superpoint feature extraction

[0067] def extract_superpoint_features(image):

[0068] sift = cv2.SIFT_create()

[0069] keypoints, descriptors = sift.detectAndCompute(image, None)

[0070] return keypoints, descriptors

[0071] # Select the view with the densest features as the query image

[0072] def select_query_image(images):

[0073] max_features = 0

[0074] query_image = None

[0075] query_descriptors = None

[0076] for img in images:

[0077] _, descriptors = extract_superpoint_features(img)

[0078] if descriptors is not None and len(descriptors) > max_features:

[0079] max_features = len(descriptors)

[0080] query_image = img

[0081] query_descriptors = descriptors

[0082] return query_image, query_descriptors

[0083] # Read the 3DGS parameters of the candidate regions and the image database

[0084] def load_image_database(database_path):

[0085] image_paths = glob.glob(database_path + ' / *.jpg')

[0086] database = {}

[0087] for path in image_paths:

[0088] img = cv2.imread(path, cv2.IMREAD_GRAYSCALE)

[0089] _, descriptors = extract_superpoint_features(img)

[0090] if descriptors is not None:

[0091] database[path] = descriptors

[0092] return database

[0093] # Calculate the most similar database image

[0094] def find_best_match(query_descriptors, database):

[0095] best_match = None

[0096] best_score = -1

[0097] for path, descriptors in database.items():

[0098] if descriptors is not None:

[0099] sim_score = cosine_similarity(query_descriptors,descriptors).max()

[0100] if sim_score > best_score:

[0101] best_score = sim_score

[0102] best_match = path

[0103] return best_match, best_score

[0104] Among them, the cloud described in S600 renders the image under the current pose estimation through 3DGS, and the updated pose estimation obtained through feature matching and PNP optimization iteration includes: S610, the cloud renders the image under the current pose estimation through 3DGS; S620, extracts the feature points of the rendered image, matches the feature points with the query image, and adopts the PNP algorithm to refine and iterate to obtain the updated pose estimation.

[0105] 3DGS has the ability to quickly and accurately render images from a specified perspective. Feature points are extracted from the images obtained by 3DGS and matched with the feature points extracted from the query image. Since the image depth information can be obtained from 3DGS, for the matched feature points, the pose of the query image can be refined in the PNP manner as the estimated pose for the next iteration. Some code examples are as follows:

[0106] def render_3dgs_image(pose, scene_model):

[0107] """

[0108] Render the image under the current pose estimation through 3DGS.

[0109] :param pose: Current pose estimation (rotation + translation matrix)

[0110] :param scene_model: 3DGS scene model

[0111] :return: Rendered RGB image

[0112] """

[0113] vis = o3d.visualization.Visualizer()

[0114] vis.create_window(visible=False)

[0115] vis.add_geometry(scene_model)

[0116] ctr = vis.get_view_control()

[0117] ctr.convert_from_pinhole_camera_parameters(pose)

[0118] vis.poll_events()

[0119] vis.update_renderer()

[0120] image = vis.capture_screen_float_buffer(do_render=True)

[0121] vis.destroy_window()

[0122] return np.asarray(image * 255, dtype=np.uint8)

[0123] def extract_features(image):

[0124] """

[0125] Extract feature points of the image.

[0126] :param image: Input RGB image

[0127] :return: Key points and descriptors

[0128] """

[0129] sift = cv2.SIFT_create()

[0130] gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

[0131] keypoints, descriptors = sift.detectAndCompute(gray, None)

[0132] return keypoints, descriptors

[0133] def match_features(desc1, desc2):

[0134] """

[0135] Match the feature points of two images.

[0136] :param desc1: Descriptor of the query image

[0137] :param desc2: Descriptor of the rendered image

[0138] :return: Indices of the matched feature points

[0139] """

[0140] bf = cv2.BFMatcher(cv2.NORM_L2, crossCheck=True)

[0141] matches = bf.match(desc1, desc2)

[0142] matches = sorted(matches, key=lambda x: x.distance)

[0143] return matches

[0144] def refine_pose(matches, kp1, kp2, camera_matrix):

[0145] """

[0146] Optimize the pose estimation through the PNP algorithm.

[0147] :param matches: Matched feature points

[0148] :param kp1: Keypoints of the query image

[0149] :param kp2: Rendered image key points

[0150] :param camera_matrix: Camera intrinsic matrix

[0151] :return: Optimized rotation matrix and translation vector

[0152] """

[0153] obj_pts = np.array([kp1[m.queryIdx].pt for m in matches], dtype=np.float32)

[0154] img_pts = np.array([kp2[m.trainIdx].pt for m in matches], dtype=np.float32)

[0155] _, rvec, tvec, _ = cv2.solvePnPRansac(obj_pts, img_pts, camera_matrix, None)

[0156] R_mat = cv2.Rodrigues(rvec)[0]

[0157] return R_mat, tvec

[0158] Combined with the first aspect, the embodiments of the present invention provide a seventh possible implementation manner of the first aspect. Among them, the obtaining of the UAV positioning by transmitting the updated pose estimation calculated by the cloud to the UAV terminal in S700 includes: S710, transmitting the updated pose estimation calculated by the cloud and the angle of the query image to the UAV terminal; S720, the UAV calculates the current position and camera pose of the UAV according to the camera position parameters and the query image angle.

[0159] Among them, the UAV communication supports data sharing between different UAV terminals; the reconstruction scene is constructed once and can also be used for other downstream tasks other than positioning.

[0160] The second embodiment of the present invention provides a UAV indoor positioning system based on three-dimensional Gaussian splash rendering of a scene, which includes: a scene reconstruction module for deploying a 2D voxelized 3DGS model in the cloud for reconstructing the scene and configuring UAV communication; an image library construction module for the UAV to rotate at a fixed point to capture 12 RGB-D images to construct an image library of a panoramic view; an image parameter acquisition module for uploading the internal parameters of the UAV camera and the obtained image library to the cloud; a region judgment module for the cloud to judge the region where the UAV is located according to the uploaded image library; an initial pose estimation module for the cloud to select an optimal query image from the image library based on feature density, and obtain an initial pose estimation through feature matching between the adjacent 3DGS model and the image database; an updated pose estimation module for the cloud to render the image under the current pose estimation through 3DGS, and obtain an updated pose estimation through feature matching and PNP optimization iteration; a positioning implementation module for transmitting the updated pose estimation calculated by the cloud to the UAV terminal to obtain UAV positioning.

[0161] The embodiments of the present invention aim to protect a UAV indoor positioning method and system based on three-dimensional Gaussian splash rendering of a scene, having the following effects:

[0162] 1. The present invention proposes a UAV indoor positioning method based on three-dimensional Gaussian splash rendering of a scene, which can achieve efficient and accurate indoor positioning under low computational resource requirements. By importing 3DGS parameters into the cloud and storing them in the form of a 2D voxel grid, large-scale scenes can be efficiently rendered with less memory occupancy. At the same time, the KD-tree architecture is used to accelerate spatial queries, thereby improving the computational efficiency.

[0163] 2. During the UAV positioning process, the present invention adopts a method without initial pose estimation, uses panoramic acquisition technology to obtain 12 RGB-D images, and selects the optimal viewing angle as the query image based on Superpoint feature point analysis. This method not only reduces the requirements for the local computing ability of the UAV, but also realizes efficient and real-time positioning through cloud computing. In addition, this method combines the YOLOv8 object detection algorithm to extract object category features from the RGB image and match them with the 3DGS voxel semantic table, so as to quickly determine the approximate position of the UAV in the candidate region and realize a semantic-enhanced positioning mechanism.

[0164] 3. In terms of pose estimation, the present invention adopts an initial pose estimation method based on database retrieval, searches and matches images in the 3DGS database by brute-force search, and selects the pose corresponding to the closest database image as the initial estimated value. The virtual view under the current pose is rendered using 3DGS, feature-matched with the query image, and the pose estimation is iteratively optimized in combination with the PNP algorithm, thereby further improving the positioning accuracy. The application of the cloud computing mode enables the drone to quickly upload data and receive calculation results without relying on high-performance local computing units, effectively reducing the storage and computing power requirements for the drone terminal.

[0165] The computer program product of the drone indoor positioning method and device based on three-dimensional Gaussian splash rendering scene provided by the embodiments of the present invention includes a computer-readable storage medium storing program codes, and the instructions included in the program codes can be used to execute the methods in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated herein.

[0166] Specifically, the storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the storage medium is run, it can execute the above-mentioned drone indoor positioning method based on three-dimensional Gaussian splash rendering scene, so as to achieve drone positioning with high precision and low resource requirements without initial pose estimation, eliminating the need for the drone to high-computing-power devices.

[0167] If the said function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0168] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or make equivalent replacements for some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims described.

Claims

1. A method for indoor positioning of an unmanned aerial vehicle based on three-dimensional Gaussian splash rendering of a scene, characterized in that, Including: S100, a 3DGS model with 2D voxelization deployed in the cloud, used for reconstructing a scene and configured for drone communication; S200, the drone takes 12 RGB-D images by rotating at a fixed point to build an image library of a panoramic view; S300, uploading the internal parameters of the drone camera and the obtained image library to the cloud; S400, the cloud determines the area where the drone is located according to the uploaded image library; S500, the cloud selects the optimal query image from the image library based on feature density, and obtains an initial pose estimate through feature matching between the adjacent 3DGS model and the image database; S600, the cloud renders the image under the current pose estimate through 3DGS, and obtains an updated pose estimate through feature matching and PNP optimization iteration; S700, transmitting the updated pose estimate calculated by the cloud to the drone terminal to obtain drone positioning.

2. The UAV indoor positioning method based on three-dimensional Gaussian splash rendering of a scene according to claim 1, wherein, The 3DGS model with 2D voxelization deployed in the cloud in S100, used for reconstructing a scene and configured for drone communication includes: S110, importing the parameters of the 3DGS model into the cloud and saving them in the form of 2D voxels, where the training of the parameters of the 3DGS model is obtained by pre-collecting, transmitting and calculating scene images; S120, importing the semantic table of objects in the scene corresponding to each voxel grid into the cloud to obtain a reconstructed scene; S130, configuring a drone wireless communication module for the cloud.

3. The UAV indoor positioning method based on three-dimensional Gaussian splash rendering of a scene according to claim 1, characterized in that, The drone in S200 takes 12 RGB-D images by rotating at a fixed point to build an image library of a panoramic view includes: S210, the drone keeps the spatial coordinates unchanged and rotates horizontally for one circle; S220, saving the images captured by the drone camera and the depth information obtained by the sensor every 30° to form 12 RGB-D images with a panoramic view, and obtaining an image library of a panoramic view.

4. The method for indoor positioning of an unmanned aerial vehicle based on three-dimensional Gaussian splash rendering of a scene according to claim 1, characterized in that, The uploading of the internal parameters of the drone camera and the obtained image library to the cloud in S300 includes: S310, uploading the internal parameters of the drone camera to the cloud; S320, uploading the RGB-D images and the corresponding rotation angles to the cloud.

5. The method for indoor positioning of an unmanned aerial vehicle based on three-dimensional Gaussian splash rendering of a scene according to claim 2, wherein The cloud in S400 determines the area where the drone is located according to the uploaded image library includes: S410, the cloud identifies the objects in the RGB-D images and saves the semantic information of the objects; S420, comparing and identifying the semantic categories of the objects obtained in the RGB-D images with the semantic tables of the objects in each voxel area, and selecting candidate 3DGS areas.

6. The UAV indoor positioning method based on three-dimensional Gaussian splash rendering scene according to claim 5, characterized in that, The cloud in S500 selects the optimal query image from the image library based on feature density, and obtains an initial pose estimate through feature matching between the adjacent 3DGS model and the image database includes: S510, the cloud extracts the Superpoint features of the RGB-D images at 12 viewpoints from the image library, and selects the viewpoint with the densest features as the query image; S520, extracting the 3DGS parameters and the corresponding image database in the candidate 3DGS areas; S530, brute-force search the corresponding image database and perform feature matching with the query image, and select the pose corresponding to the closest database image as the initial pose estimate.

7. The indoor positioning method of an unmanned aerial vehicle based on three-dimensional Gaussian splash rendering of a scene according to claim 6, wherein In S600, the cloud renders the image under the current pose estimate through 3DGS, and the updated pose estimate obtained through feature matching and PNP optimization iteration includes: S610, the cloud renders the image under the current pose estimate through 3DGS; S620, extract the feature points of the rendered image, perform feature point matching with the query image, and adopt the PNP algorithm to refine and iterate to obtain the updated pose estimate.

8. The method for indoor positioning of an unmanned aerial vehicle based on three-dimensional Gaussian splash rendering of a scene according to claim 7, wherein In S700, the transmission of the updated pose estimate calculated by the cloud to the UAV terminal to obtain UAV positioning includes: S710, transmit the updated pose estimate calculated by the cloud and the angle of the query image to the UAV terminal; S720, the UAV calculates the current position and camera pose of the UAV according to the camera position parameters and the query image angle.

9. The UAV indoor positioning method based on three-dimensional Gaussian splash rendering scene according to claim 2, characterized in that the UAV communication supports data sharing between different UAV terminals; the reconstructed scene is constructed once.

10. An indoor positioning system for drones based on three-dimensional Gaussian splash rendering of a scene, characterized in that, Including: A scene reconstruction module, which is used to deploy a 2D voxelized 3DGS model in the cloud for reconstructing the scene and configuring UAV communication; An image library construction module, which is used for the UAV to take 12 RGB-D images by fixed-point rotation to construct an image library of a panoramic view; An image parameter acquisition module, which is used to upload the internal parameters of the UAV camera and the obtained image library to the cloud; A region judgment module, which is used for the cloud to judge the region where the UAV is located according to the uploaded image library; An initial pose estimation module, which is used for the cloud to select the optimal query image from the image library based on feature density, and obtain the initial pose estimation through feature matching between the adjacent 3DGS model and the image database; An updated pose estimation module, which is used for the cloud to render the image under the current pose estimate through 3DGS, and obtain the updated pose estimate through feature matching and PNP optimization iteration; A positioning implementation module, which is used to transmit the updated pose estimate calculated by the cloud to the UAV terminal to obtain UAV positioning.

Citation Information

Patent Citations

  • Unmanned aerial vehicle image visual positioning method supporting scene apparent difference

    CN117893600A

  • Unmanned aerial vehicle visual positioning method and device based on combined guide model image matching

    CN118887280A

  • Visual repositioning method and system based on 3D Gaussian scene and storage medium

    CN118941629A

  • Laser enhanced vision three-dimensional reconstruction method and system based on Gaussian splashing

    CN119180908A

  • Sparse input scene reconstruction method based on 3D Gaussian sputtering and computer program product

    CN119206051A

Cited By

  • Unmanned aerial vehicle cluster identification method and system based on multi-target pointing end point distribution modeling

    CN122044216A