Positioning method and related device

By using 3D point cloud data and Gaussian primitives in visual SLAM technology, combining target illuminance and historical environment maps, the problem of insufficient positioning accuracy and robustness of visual SLAM in complex environments is solved, and higher positioning accuracy and robustness are achieved.

CN120070568APending Publication Date: 2025-05-30CRRC TECH INNOVATION (BEIJING) CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510007135.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Visual SLAM technology has low positioning accuracy and robustness in complex environments, especially affected by lighting conditions.

Method used

By obtaining the 3D point cloud data of the target environment, initialize the Gaussian primitives, obtain a historical environment map that meets the lighting matching conditions from the scene map based on the target illuminance, build a positioning map, and update the camera position pose.

Benefits of technology

It improves the positioning accuracy and robustness of visual SLAM in complex environments, and reduces the impact of changes in lighting conditions on positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070568A_ABST
    Figure CN120070568A_ABST
Patent Text Reader

Abstract

The invention discloses a positioning method and a related device, and relates to the technical field of computer vision, and the method uses 3D point cloud data of a target environment image to initialize Gaussian primitives to obtain new Gaussian primitives. And based on the target illuminance, acquiring a historical environment map meeting a preset illumination matching condition from the historical environment map with the illuminance label as an illumination matching map, and based on the illumination matching map, acquiring a positioning map. And updating the camera pose based on the new Gaussian primitives and the Gaussian primitives of the positioning map. The target illuminance is the ambient illuminance when the target environment image is shot, the illuminance label of the historical environment map represents the ambient illuminance when the historical environment map is generated, and the illumination matching condition at least comprises that the illumination difference degree between the illuminance label and the target illuminance is smaller than the illumination difference degree threshold value. Therefore, according to the method, positioning is carried out based on the historical environment map with the illumination condition similar to the current target environment image, and the positioning accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular, to a positioning method and related devices. Background Art

[0002] The SLAM (Simultaneous Localization and Mapping) technology is one of the core technologies for devices such as robots, drones, intelligent security devices, and locomotive trains to move and work autonomously in complex environments. Through the SLAM technology, the device can instantaneously locate its own position and attitude based on repeatedly observed environmental features during movement, and at the same time construct an incremental environmental map, thereby realizing autonomous positioning and navigation. Compared with lidar SLAM, visual SLAM has attracted much attention for its advantages such as low cost, light equipment, and low power consumption, especially showing significant advantages in resource-constrained scenarios. Visual SLAM relies on the image information collected by the camera, does not require expensive lidar equipment, greatly reduces the hardware cost, and can be seamlessly integrated with the existing visual perception system, thus having higher applicability.

[0003] However, visual SLAM is easily affected by environmental conditions, which reduces the accuracy and robustness of visual SLAM. Therefore, how to improve the positioning accuracy and robustness of visual SLAM in complex environments has become an important research direction in the field of visual SLAM. Summary of the Invention

[0004] In view of the above problems, this application provides a positioning method and related devices to achieve the purpose of improving the positioning accuracy of positioning achieved by visual SLAM technology. The specific solutions are as follows:

[0005] The first aspect of this application provides a positioning method, including:

[0006] Obtain the 3D point cloud data of the target environmental image;

[0007] Initialize the Gaussian basis element using the 3D point cloud data to obtain a new Gaussian basis element;

[0008] Based on the target illuminance, obtain a historical environmental map that meets the preset illuminance matching condition from the scene map table as the illuminance matching map, where the illuminance matching condition at least includes that the illuminance difference between the illuminance label and the target illuminance is less than the preset illuminance difference threshold; the target illuminance is the environmental illuminance when the target environmental image is captured, the scene map table includes historical environmental maps with illuminance labels, and the illuminance label of the historical environmental map represents the environmental illuminance when the historical environmental map is generated;

[0009] Obtain a positioning map based on the illuminance matching map;

[0010] Update the camera pose based on the new Gaussian basis element and the Gaussian basis element of the positioning map.

[0011] In a possible implementation, obtaining the 3D point cloud data of the target environment image includes:

[0012] Receive the monocular environment image collected by the monocular image acquisition device as the target environment image;

[0013] Obtain the depth map of the target environment image through a pre-trained depth estimation model;

[0014] Based on the preset camera internal parameters, map the depth map of the target environment image into a 3D point cloud to obtain the 3D point cloud data.

[0015] In a possible implementation, obtaining the historical environment map that meets the preset illumination matching condition from the scene map table based on the target illumination includes:

[0016] Obtain the illumination of the target environment image based on the preset illumination feature;

[0017] For each target historical environment map, compare the illumination label of the target historical environment map with the target illumination to obtain the illumination difference degree of the target historical environment map, where the target historical environment map includes at least one historical environment map in the scene map table;

[0018] For each target historical environment map, compare the illumination difference degree with the illumination difference degree threshold. If the illumination difference degree is less than the illumination difference degree threshold, determine that the target historical environment map meets the illumination matching condition; if the illumination difference degree is not less than the illumination difference degree threshold, determine that the target historical environment map does not meet the illumination matching condition.

[0019] In a possible implementation, obtaining the positioning map based on the illumination matching map includes:

[0020] If the number of the illumination matching maps is equal to 1, use the illumination matching map as the positioning map;

[0021] If the number of the illumination matching maps is greater than 1, obtain the two historical environment maps with the smallest illumination difference degree from all the illumination matching maps, and use them as the first illumination matching map and the second illumination matching map respectively;

[0022] For each Gaussian basis element at the same position in the first illumination matching map and the second illumination matching map respectively, perform weighted fusion based on the corresponding weighting coefficients to obtain a set of fused Gaussian basis elements. The weighting coefficient of the Gaussian basis element of the target illumination matching map is inversely correlated with the illumination difference degree of the target illumination matching map. The target illumination matching map includes the first illumination matching map and the second illumination matching map;

[0023] Construct the positioning map based on the set of fused Gaussian basis elements.

[0024] In a possible implementation, updating the camera pose based on the new Gaussian basis element and the Gaussian basis element of the positioning map includes:

[0025] If there is the positioning map, use a preset point cloud-based registration algorithm to register the new Gaussian basis element and the Gaussian basis element of the positioning map to obtain a rotation matrix;

[0026] Update the camera pose using the rotation matrix.

[0027] In a possible implementation, after obtaining the historical environment map that meets the preset illumination matching condition from the scene map table based on the target illumination, the positioning method further includes:

[0028] If there is an illumination matching map that meets the preset update condition, select an illumination matching map that meets the update condition as the reference map. The update condition includes that the illumination difference degree is less than a preset update threshold, and the update threshold is less than the illumination difference degree threshold;

[0029] Update the reference map based on the new Gaussian basis element;

[0030] Update the illumination label of the reference map based on the target illumination.

[0031] In a possible implementation, updating the reference map based on the new Gaussian basis element includes:

[0032] Add the new Gaussian basis element to the reference map to obtain a first candidate map;

[0033] Perform heuristic pruning on the number of Gaussian basis elements in the first candidate map to obtain a second candidate map;

[0034] Based on the second candidate map, use a differentiable rendering path to render a rendered image at the camera pose;

[0035] Compare the rendered image with the target monocular environment image, calculate a loss value, backpropagate the loss value, and calculate the gradient corresponding to each Gaussian basis element in the second candidate map;

[0036] Update each Gaussian basis element based on the gradient corresponding to each Gaussian basis element in the second candidate map to obtain the environmental map of the target environmental image;

[0037] Update the reference map to the environmental map of the target environmental image.

[0038] In a possible implementation, after obtaining, based on the target illuminance, a historical environmental map that meets the preset illuminance matching condition from the scene map table, the positioning method further includes:

[0039] If there is no illuminance matching map that meets the update condition, construct a new environmental map based on the new Gaussian basis element;

[0040] Store the new environmental map in the scene map table with the target illuminance as the illuminance label.

[0041] A second aspect of the present application provides a positioning device, including:

[0042] A point cloud generation unit, configured to obtain 3D point cloud data of a target environmental image;

[0043] A basis element initialization unit, configured to initialize Gaussian basis elements using the 3D point cloud data to obtain new Gaussian basis elements;

[0044] An illuminance matching unit, configured to obtain, based on the target illuminance, a historical environmental map that meets the preset illuminance matching condition from the scene map table as an illuminance matching map, where the illuminance matching condition at least includes that the illuminance difference degree between the illuminance label and the target illuminance is less than a preset illuminance difference degree threshold; the target illuminance is the environmental illuminance when the target environmental image is captured, the scene map table includes historical environmental maps with illuminance labels, and the illuminance label of the historical environmental map represents the environmental illuminance when the historical environmental map is generated;

[0045] A positioning map acquisition unit, configured to obtain a positioning map based on the illuminance matching map;

[0046] A positioning unit, configured to update the camera pose based on the new Gaussian basis element and the Gaussian basis element of the positioning map.

[0047] A third aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, where:

[0048] The memory is used to store a computer program;

[0049] The processor is configured to execute the computer program so that the electronic device can implement the positioning method in the first aspect or any implementation manner of the first aspect.

[0050] With the above technical solution, a positioning method and related device provided by the present application initialize a Gaussian basis element using the 3D point cloud data of the target environment image to obtain a new Gaussian basis element. Based on the target illuminance, a historical environment map that meets the preset illuminance matching condition is obtained from the historical environment map with illuminance labels as the illuminance matching map, and a positioning map is obtained based on the illuminance matching map. The camera pose is updated based on the new Gaussian basis element and the Gaussian basis element of the positioning map. Since the target illuminance is the environmental illuminance when the target environment image is captured, the illuminance label of the historical environment map represents the environmental illuminance when the historical environment map is generated, and the illuminance matching condition at least includes that the illuminance difference between the illuminance label and the target illuminance is less than the illuminance difference threshold. Therefore, this method performs positioning based on the historical environment map with illuminance conditions similar to the current target environment image, avoiding the influence of illuminance condition changes on positioning, thereby improving the positioning accuracy and robustness. Description of the Drawings

[0051] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0052] Figure 1 It is a schematic flowchart of a positioning method provided by an embodiment of the present application;

[0053] Figure 2 It is a schematic flowchart of a SLAM system provided by an embodiment of the present application;

[0054] Figure 3 It is a schematic flowchart of a specific implementation process of a positioning method provided by an embodiment of the present application;

[0055] Figure 4 It is a specific implementation flowchart of obtaining a positioning map and a reference map provided by an embodiment of the present application;

[0056] Figure 5 It is a schematic structural diagram of a positioning device provided by an embodiment of the present application;

[0057] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0058] The following describes the embodiments of the present application in combination with the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only for explaining the specific embodiments of the present application and are not intended to limit the present application.

[0059] The embodiments of the present application will be described below in conjunction with the accompanying drawings. As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0060] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0061] Taking a locomotive train in a closed-scene intelligent transportation system of a mobile device as an example, the locomotive train needs to perform efficient and precise transportation operations in a closed yard.

[0062] To achieve precise control of the start and stop of the locomotive train, the traditional lidar SLAM technology realizes environmental perception and map construction of the closed yard through lidar to determine the position of the locomotive train in the yard in real time, so as to ensure transportation efficiency and safety. However, lidar SLAM relies on expensive lidar sensors. For an intelligent transportation system that needs to be deployed on a large scale, it faces many limitations due to its high cost and poor practicability. In addition, the lidar SLAM technology needs to transmit high-bandwidth three-dimensional point cloud data and relies on high-performance computing units for processing, which further exacerbates the problems of high cost and high power consumption. Therefore, the application scenario of the lidar SLAM technology is limited by the large amount of heat generated by the equipment, the high transmission bandwidth load, and the high overall investment and operating costs.

[0063] To overcome the deficiencies of lidar SLAM, the prior art uses visual SLAM technology to construct an environmental map. Visual SLAM relies on the image information collected by the camera device and realizes environmental map construction and positioning through algorithm processing.

[0064] For example, the visual SLAM technology based on NeRF (Neural Radiance Field) constructs an environmental image by using the NeRF method after collecting multi-angle images through a camera device. However, the NeRF method requires high computing power to learn and infer the implicit representation of the three-dimensional scene, resulting in a slow processing speed and difficulty in meeting the real-time positioning requirements.

[0065] For another example, in the visual SLAM technology based on GS (Gaussian Splatting), environmental images and depth information are obtained through an RGB-D camera, and an environmental map is constructed based on the Gaussian Splatting technology. However, the RGB-D camera has a high hardware cost and poor positioning performance.

[0066] Through research, the inventors of this application found that the sensitivity of the existing visual SLAM technology based on GS to lighting conditions further limits its applicability in practical scenarios. For example, when the existing technology obtains environmental images and depth information through an RGB-D camera, it measures depth by actively emitting light, which is limited by the effective detection distance and lighting conditions in outdoor applications, especially performing poorly in complex lighting and large-scale environments. For another example, changes in lighting conditions will significantly affect the extraction and matching of image features. Specifically, in low-light environments, the extraction of feature points may be incomplete; in strong light or uneven light environments, shadows and overexposed areas will reduce the accuracy of positioning and map construction. This sensitivity to lighting changes makes it difficult for the existing visual SLAM technology to work stably under dynamic and complex lighting conditions. It can be seen that the existing visual SLAM technology still has difficulty overcoming the technical drawback of poor adaptability.

[0067] To solve the above technical drawbacks, the embodiments of this application provide a positioning method. By applying the visual SLAM technology based on GS and through a pre-constructed scene map table and the analysis of the illuminance of the currently collected environmental image, a positioning map for positioning is obtained. To overcome the influence of lighting conditions on positioning accuracy and improve positioning accuracy. The following will introduce the positioning method of the embodiments of this application in detail with reference to the accompanying drawings. This method takes the edge inference device in the SLAM system as the execution entity. Refer to Figure 1 , Figure 1 is a schematic flowchart of a positioning method provided by the embodiments of this application. As Figure 1 shown, a data processing method provided by the embodiments of this application may include steps S101 to S105, and the following will describe these steps in detail.

[0068] S101. Obtain the 3D point cloud data of the target environmental image.

[0069] In this embodiment, the target environmental image is an environmental image collected in real time based on a pre-configured camera. For example, a monocular camera or a multi-camera is configured at the top of an intelligent mobile device, and the monocular camera or the multi-camera is controlled to collect environmental images based on a preset acquisition frequency. The 3D point cloud data includes the spatial coordinates and color information of each point in the 3D point cloud.

[0070] In this embodiment, based on the different types of cameras, the information carried by the target environment image is different. Therefore, the method for obtaining the 3D point cloud data of the target environment image is different.

[0071] For example, the target environment image captured by a monocular camera is a monocular image. The depth information of the target environment image is predicted using a pre-constructed depth estimation model, and 3D point cloud data is obtained through the GS technology.

[0072] For another example, the target environment image captured by a multi-camera is a multi-view image. Based on the depth information carried by the multi-view image, 3D point cloud data is obtained through the GS technology.

[0073] It should be noted that the method for obtaining the 3D point cloud data of the target environment image can refer to the prior art.

[0074] S102. Initialize the Gaussian basis element using the 3D point cloud data to obtain a new Gaussian basis element.

[0075] In this embodiment, the environmental map is represented by a set of Gaussian basis elements, that is, the environmental map is composed of multiple Gaussian basis elements.

[0076] In this embodiment, the Gaussian basis element is a parameterized three-dimensional object. The parameters of the Gaussian basis include the mean vector, covariance matrix, color, and transparency. The Gaussian basis element uses the mean vector and covariance matrix to represent the position and shape. The Gaussian basis element is described by its mean vector describing the three-dimensional space position. For each Gaussian basis element, its mean represents the coordinates of the basis element center in the three-dimensional space, and the shape is described by the covariance matrix In the initialization, it is assumed that each point in the 3D point cloud corresponds to a Gaussian basis element. The mean in the Gaussian basis element is the coordinate of the point, and the covariance is calculated using the k-nearest neighbors of the point.

[0077] S103. Based on the target illuminance, obtain the historical environmental map that meets the preset illuminance matching condition from the scene map table as the illuminance matching map.

[0078] In this embodiment, the illuminance matching condition at least includes that the illuminance difference between the illuminance label and the target illuminance is less than the preset illuminance difference threshold. The target illuminance is the environmental illuminance when the target environment image is captured. The scene map table includes historical environmental maps with illuminance labels, and the illuminance label of the historical environmental map represents the environmental illuminance when the historical environmental map is generated.

[0079] In this embodiment, the illuminance difference is the absolute value of the difference between the illuminance label and the target illuminance.

[0080] S104. Obtain the positioning map based on the illuminance matching map.

[0081] In this embodiment, there are various specific methods for obtaining the positioning map based on the light-matching map. For example, a map is randomly selected from the light-matching map as the positioning map. Another example is that multiple light-matching maps are fused to obtain the positioning map.

[0082] In this embodiment, the illuminance of the positioning map still satisfies the light-matching condition.

[0083] S105. Update the camera pose based on the new Gaussian basis element and the Gaussian basis element of the positioning map.

[0084] In this embodiment, there are various methods for updating the camera pose based on the new Gaussian basis element and the Gaussian basis element of the positioning map. In an optional embodiment, a preset point cloud-based registration algorithm can be used to register the new Gaussian basis element and the Gaussian basis element of the positioning map to obtain a rotation matrix, and the rotation matrix is used to update the camera pose.

[0085] As can be seen from the above technical solutions, a positioning method provided by an embodiment of the present application initializes a Gaussian basis element using 3D point cloud data of a target environment image to obtain a new Gaussian basis element. Based on the target illuminance, a historical environment map that satisfies a preset light-matching condition is obtained from the historical environment map with illuminance labels as the light-matching map. Based on the light-matching map, the positioning map is obtained. The camera pose is updated based on the new Gaussian basis element and the Gaussian basis element of the positioning map. Since the target illuminance is the environmental illuminance when the target environment image is captured, the illuminance label of the historical environment map represents the environmental illuminance when the historical environment map is generated, and the light-matching condition at least includes that the illuminance difference between the illuminance label and the target illuminance is less than the illuminance difference threshold. Therefore, this method performs positioning based on a historical environment map with lighting conditions similar to the current target environment image, avoiding the influence of lighting condition changes on positioning, thereby improving positioning accuracy and robustness.

[0086] Furthermore, an embodiment of the present application provides a SLAM system for implementing the positioning method. Figure 2 As shown in the structural schematic diagram of a SLAM system provided by an embodiment of the present application, Figure 2 the SLAM system includes an edge inference device and a monocular imaging device configured in an intelligent mobile device, and an inference server. Among them, the intelligent mobile device includes a mobile phone, a drone, a robot, a car, a train, etc. The edge inference device is a controller configured with computing units such as computing cards, computing chips, and computing cores that focus on large-scale parallel computing. The monocular image acquisition device includes a monocular camera configured on the intelligent mobile device. For example, the monocular image acquisition device and the edge inference device of self-sensing and controlling mobile devices such as a sweeping robot, a locomotive train, and a drone are both configured in the self-sensing and controlling mobile device. The inference server is a high-performance GPU server.

[0087] In this embodiment, a pre-trained depth estimation model is configured in the edge inference device, and the depth estimation model is obtained by training on a high-performance GPU server based on big data. Optionally, the depth estimation model is pre-configured in the edge inference device and used to train a large-scale depth estimation model. The RTX 6000 is a high-performance GPU that can handle the training process of large-scale models. The host device on the edge inference side is equipped with a Samsung 990 EVO SSD storage device and an Intel i5-1240P processor with 4 cores, which supports system operation and map-related operation calculations. The inference-side computing device uses NVIDIA Jetson AGX Orin to achieve fast inference for depth estimation and Gaussian splatting. NVIDIA Jetson AGX Orin has powerful computing capabilities and energy efficiency. Placing the computing tasks of the positioning method on the edge inference device can reduce latency and improve the efficiency and real-time performance of the system.

[0088] It should be noted that a communication connection is pre-established between the edge inference device and the monocular image acquisition device in the SLAM system. The communication connection method is determined based on the hardware configuration methods of the edge inference device and the monocular image acquisition device. For example, for the edge inference device and the monocular image acquisition device configured proximally (inside a mobile device) with the same configuration, wired communication can be used. For the edge inference device configured distally, a wireless communication method can be used to communicate with the monocular image acquisition device. The specific communication method can refer to the prior art.

[0089] In this embodiment, the monocular image acquisition device is used to collect monocular environment images in real time and send the monocular environment images to the edge inference device in real time according to the time sequence. The edge inference device is used to achieve positioning based on the monocular environment images, combined with monocular depth estimation and GS (Gaussian Splatting) technology. Since this system uses monocular environment images for environmental perception and Gaussian splatting imaging, the limitations of the SLAM technology on the performance of the camera device and the edge inference device are reduced, and the speed of environmental map construction and the applicability of the SLAM technology are improved. A low-cost monocular camera is used to replace expensive lidar sensors and RGB-D cameras, thereby simplifying the hardware configuration, reducing the hardware cost, and improving the applicability of the SLAM technology, especially suitable for mobile intelligent devices that require autonomous decision-making operation.

[0090] In addition, since this system uses monocular environment images for environmental perception, it can also be applied to the monocular mode of a multi-camera. Among them, the monocular mode of the multi-camera can be an active-triggered power-saving working mode for reducing power consumption, or a disaster-tolerant working mode that is passively triggered due to a fault.

[0091] Based on the above SLAM system, an embodiment of the present application provides a specific implementation manner of a positioning method, which is specifically applied to an edge inference device in the SLAM system. Figure 3 It is a specific implementation flowchart of a positioning method provided by an embodiment of the present application. As Figure 3 shown, this method may specifically include:

[0092] S301. Receive the target monocular environment image collected by the monocular image acquisition device.

[0093] In this embodiment, the monocular image acquisition device may adopt an Oak-D-Pro camera. The target monocular environment image is an RGB image with a resolution of 640*480. The monocular image acquisition device captures the target monocular environment image at a preset frame rate to form a scene video stream, and sends the scene video stream to the edge inference device frame by frame.

[0094] Among them, the preset frame rate may be a fixed frame rate or a dynamic frame rate to meet the requirements of real-time performance and scene change information. Optionally, the preset frame rate is 10 frames per second.

[0095] S302. Obtain the depth map of each frame of the target monocular environment image through a pre-trained depth estimation model.

[0096] In this embodiment, the depth map of the target monocular environment image includes the absolute depth information of each pixel in the target monocular environment image, and the absolute depth information includes depth values. Specifically, the collected target monocular environment image is input into the depth estimation model. The depth estimation model performs feature analysis on each pixel in the target monocular environment image, infers the absolute depth information (unit: meter) corresponding to the pixel, and outputs the depth map of the target monocular environment image.

[0097] In this embodiment, the depth estimation model has been pre-trained on a large-scale data set and has the generalization ability of depth estimation, and can adapt to a wide range of scenarios including outdoor and indoor. The depth map provides environmental information for the generation of 3D point clouds and the construction of slam maps of.

[0098] S303. Based on the camera internal parameters, map the depth map of the target monocular environment image into a 3D point cloud to obtain 3D point cloud data.

[0099] In this embodiment, the 3D point cloud data includes the spatial coordinates and color information of each point in the 3D point cloud.

[0100] Specifically, the camera internal parameters are used to project the pixels in the depth map of the target monocular environment image to generate the spatial coordinates of each point in the 3D point cloud. For details, see formula (1):

[0101] (1);

[0102] Among them, represents the image coordinates of the pixel, represents the principal point coordinates of the camera internal parameters, and represents the focal length, represents the pixel corresponding depth value on the depth map.

[0103] In this embodiment, the camera internal parameters are fixed after being initialized and configured in the initial state. For the specific method of initializing the camera internal parameters, reference can be made to the prior art.

[0104] S304. Initialize the Gaussian basis element using the 3D point cloud data to obtain a new Gaussian basis element.

[0105] S305. Based on the target illuminance and the scene map table, obtain the positioning map and the reference map.

[0106] In this embodiment, the target illuminance is the illuminance of the target monocular environment image, and the scene map table includes multiple historical environment maps with illuminance labels. The illuminance label of the historical environment map represents the illuminance of the monocular environment image when the historical environment map is generated.

[0107] In this embodiment, the positioning map is obtained based on the illumination matching map. Among them, the illumination matching map includes at least one historical environment map in the scene map table that matches the target illuminance. Among them, the historical environment map that matches the target illuminance satisfies at least that the illuminance difference degree between the illuminance label and the target illuminance is less than the preset illuminance difference threshold.

[0108] In an alternative embodiment, from the scene map table, select the historical environment map whose illuminance difference degree between the illuminance label and the target illuminance is less than the illuminance difference threshold and whose update timestamp is the closest to the current time as the positioning map.

[0109] In another alternative embodiment, from the scene map table, select multiple historical environment maps whose illuminance difference degree between the illuminance label and the target illuminance is less than the illuminance difference threshold, and fuse the multiple historical environment maps, and use the fusion result as the positioning map.

[0110] In this embodiment, the reference map is the illumination matching map that meets the update conditions. The update conditions include that the illuminance difference degree between the illuminance label and the target illuminance is the smallest and less than the preset update difference threshold, where the update difference threshold is less than the illuminance difference threshold.

[0111] S306. If there is a positioning map, perform GICP registration on the new Gaussian basis element and the Gaussian basis element in the positioning map to obtain a rotation matrix, and use the rotation matrix to update the camera pose.

[0112] Specifically, first, a new Gaussian basis element is created based on the 3D point cloud data. . Taking as the source and the Gaussian basis element in the positioning map as the target , the rotation matrix of the camera is obtained by registration using GICP (Generalized Iterative Closest Point, a point cloud-based registration method), and the camera pose of the camera is updated through the rotation matrix.

[0113] The GICP formula is as shown in formula (2) below:

[0114] (2);

[0115] Where is the transformation matrix.

[0116] S307. If there is a reference map, add the new Gaussian basis element to the reference map to obtain the first candidate map.

[0117] In this embodiment, the newly obtained Gaussian basis element after initialization is added to the reference map, that is, the Gaussian basis element set of the reference map is updated to .

[0118] S308. Perform heuristic pruning on the number of Gaussian basis elements in the first candidate map to obtain the second candidate map.

[0119] In this embodiment, to prevent the reduction of calculation efficiency caused by too many Gaussian basis elements in the first candidate map, based on factors such as usage frequency, transparency, and gradient, the Gaussian basis elements that have not been updated for a long time or have a high transparency are removed from the first candidate map.

[0120] Specifically, it is judged whether the Gaussian basis element in the first candidate map satisfies: the transparency is greater than the transparency threshold or the update time is greater than the update time threshold If so, remove the Gaussian basis element from the first candidate map, and update the first candidate map to the second candidate map, as shown in formula (3) below:

[0121] (3).

[0122] S309. Based on the second candidate map, use the differentiable rendering path to render the rendered image at the camera pose.

[0123] In this embodiment, the method for rendering the environmental image at the camera pose using the differentiable rendering path based on the second candidate map includes:

[0124] B1. Projected Gaussian basis element. Based on the current camera pose, project the Gaussian basis element in three-dimensional space onto the phase plane, and consider the mean value of the Gaussian basis element during the projection process and covariance matrix .

[0125] Specifically, first project the mean position. Use the internal parameter matrix K of the camera to project the three-dimensional mean vector to the pixel position on the two-dimensional plane as shown in Equation (4):

[0126] (4);

[0127] where and are the normalized coordinates on the phase plane. The internal parameter matrix K of the camera is generally expressed as:

[0128] , where and represent the focal lengths, and are the principal point coordinates. The obtained is the pixel position of the mean value of the Gaussian basis element on the phase plane.

[0129] Then, project the covariance matrix: Let the rotation matrix of the monocular camera be R and the translation vector be t. Then the projection of the three-dimensional covariance matrix is , where the Jacobian matrix J is expressed as Equation (5):

[0130] (5);

[0131] where z is the depth value corresponding to the Gaussian basis element with a mean value of . The depth value of the Gaussian basis element is its projection in the direction of the camera viewpoint (usually the z-axis direction). J describes the local change of the projection relationship, that is, the change in the two-dimensional plane position (u, v) corresponding to the small change in the three-dimensional space coordinates (x, y, z).

[0132] B2. At each pixel position on the phase plane, sort all the projected Gaussian basis elements covering this pixel position in descending order according to the depth value (distance from the camera) to obtain the Gaussian basis element sequence at the pixel position.

[0133] In this embodiment, the Gaussian basis element sequence includes the target Gaussian basis elements sorted in descending order according to the depth value. The target Gaussian basis elements are the Gaussian basis elements covering this pixel after projection. The pixel position The Gaussian basis element sequence is represented as: .

[0134] In this embodiment, the depth values of the Gaussian basis elements are arranged in descending order. So that during rendering, the Gaussian basis elements with larger depth (far away) are processed first to ensure that the Gaussian basis elements in the front can cover the pixels in the back.

[0135] B3. At the position of each pixel in the phase plane, use the transparency and color of each Gaussian basis element in the Gaussian basis element sequence corresponding to this pixel to perform weighted summation to obtain the rendering color of the pixel.

[0136] In this embodiment, the specific method of the rendering color can be seen in formula (6):

[0137] (6);

[0138] Where, is the transparency of the i-th Gaussian basis element in the Gaussian basis element sequence, is the color of the i-th Gaussian basis element in the Gaussian basis element sequence, is the transparency of the j-th Gaussian basis element in the Gaussian basis element sequence.

[0139] B6. Render the rendering image based on the rendering colors of all pixels to obtain the rendered image.

[0140] S310. Compare the rendered image with the target monocular environment image, calculate the loss value, and backpropagate the loss value to calculate the gradient corresponding to each Gaussian basis element in the second candidate map.

[0141] In this embodiment, the parameters of each Gaussian basis element include the mean position , the covariance matrix Σ, the transparency , the color c, and the gradient is .

[0142] Specifically, the specific method of calculating the loss value and backpropagating the loss value to calculate the parameters of each Gaussian basis element in the second candidate map includes:

[0143] C1. Based on the loss function, calculate the difference value between the rendered image and the target monocular environment image to obtain the loss value.

[0144] In this embodiment, the difference value represents the degree of difference between the rendered image and the target monocular environment image. The greater the difference, the greater the difference between the map and the real scene, and the greater the loss value.

[0145] The loss function formula is as formula (7):

[0146] (7);

[0147] Where, is the k-th pixel in the input target monocular environment image, is the k-th pixel in the rendered image, and the total number of pixels is K.

[0148] C2. Backpropagate the loss value along the differentiable rendering path to calculate the gradients corresponding to the parameters of each Gaussian basis element in the second candidate map.

[0149] S311. Update each Gaussian basis element based on the gradients corresponding to the Gaussian basis elements in the second candidate map to obtain the environment map of the target monocular environment image.

[0150] In this embodiment, the optimization formula is formula (8):

[0151] (8);

[0152] where 𝜃 represents the parameters of the Gaussian basis elements in, and 𝜂 is a preset learning rate.

[0153] S312. Update the reference map to the environment map of the target monocular environment image, and update the illuminance label of the reference map to the target illuminance.

[0154] It should be noted that when the SLAM system runs, based on the collected monocular environment maps, the above S301 - S312 are repeated, and the positioning task (S301~306) and the update task (S307~S312) are repeatedly executed based on each monocular environment map.

[0155] It can be seen from the above technical solutions that this solution uses a monocular depth estimation method to replace RGB-D cameras and lidar, reducing hardware dependence, reducing system complexity, and improving the applicability of the solution. Using the Gaussian sputtering method to construct the map improves the accuracy and efficiency of map construction and realizes efficient 3D reconstruction. It can achieve a positioning rate of 5 - 10 frames per second, meeting the real-time requirements of the intelligent transportation system in a closed scenario. The low-cost and high-efficiency solution provides an economical and efficient SLAM solution through the combination of a monocular camera and Gaussian sputtering technology, improving the economy of the solution.

[0156] Specifically, in a positioning method provided by an embodiment of the present application, the scene map table includes historical environment maps under different illuminations. By continuously optimizing the reference map in the scene map table through the update task, the historical environment maps under each illumination condition can more accurately represent the actual scene. Through the positioning task, the camera pose parameters are updated based on the illumination condition to obtain the positioning map, solving the influence of different illuminations on visual slam and improving the positioning accuracy.

[0157] Further, through the method for obtaining the positioning map and the reference map shown in S305, the positioning map for the positioning task and the reference map for the update task are isolated, and the positioning map is discarded after use and not saved, thereby preventing the map from being updated frequently and stabilizing the operation of visual slam.

[0158] Further, through the method for executing the update task shown in S307~S312, in the presence of a reference map, based on the target monocular environmental image, the reference map and the illuminance label are updated to realize the update and optimization of the historical environmental map under the current illumination condition, improve the accuracy and integrity of the historical environmental map in the scene map table, and thereby improve the accuracy of subsequent positioning tasks.

[0159] It should be noted that Figure 3 The corresponding embodiment only provides a specific implementation process of an optional positioning method provided by the embodiments of the present application, and the present application can also be implemented through other specific processes.

[0160] For example, in an optional embodiment, the present application further includes the following steps A1 and A2:

[0161] A1. If there is no positioning map, initialize the Gaussian basis element using 3D point cloud data to obtain the target environmental map, and initialize the camera pose based on the target environmental map.

[0162] In this embodiment, the situation where there is no positioning map includes: there is no illumination matching map in the scene map table, that is, there is no historical environmental map matching the target illuminance. It should be noted that in the initial state, the absence of a historical environmental map in the scene map table also belongs to the situation where there is no positioning map.

[0163] In this embodiment, the camera pose is initialized using a two-frame based initialization method, and the specific initialization method can refer to the prior art.

[0164] In this embodiment, when creating a new environmental map, the Gaussian basis element set therein is initialized using the Gaussian sputtering technique to ensure a high-precision representation of the new environmental map. The illuminance of the target monocular environmental image will be used as the illuminance label of the new environmental map.

[0165] Specifically, initialize the Gaussian basis element , the Gaussian basis element is a parameterized three-dimensional object, using the mean , covariance matrix , and the color c represents the position, shape, and color of the Gaussian basis element. It is assumed that each point in the point cloud corresponds to a Gaussian basis element during initialization, and the mean is the coordinate of the point , during the initialization stage Calculations are performed using the k nearest neighbors of each point, and c uses the corresponding pixel color. The SLAM method of the present invention represents the environmental map using a set of Gaussian basis elements. Set the new environmental map as the reference map .

[0166] It can be seen that in the case where there is no positioning map, the camera pose is initialized, and after the camera pose is initialized, a positioning map is obtained based on subsequent frame monocular environmental images, and the camera pose is continuously updated based on the positioning map to achieve positioning.

[0167] A2. If there is no reference map, after initializing the Gaussian basis elements using 3D point cloud data to obtain new Gaussian basis elements, a new environmental map is constructed based on the set of new Gaussian basis elements, and the new environmental map is stored in the scene map table with the target illuminance as the illuminance label.

[0168] In this embodiment, the situation where there is no reference map includes: there is no illumination matching map in the scene map table that satisfies the update condition, where the update condition includes that the illumination difference degree from the target illuminance is less than the update difference threshold. It can be understood that in the case where there is no positioning map, there must be no reference map.

[0169] In this embodiment, a new environmental map is generated using the new Gaussian basis elements, and the new environmental map is stored in the scene map table with the target illuminance as the illuminance label. After collecting subsequent monocular environmental images, since the illuminance does not change within a short period of time, this new environmental map will be used as the reference map to continuously update this new environmental map.

[0170] For another example, S306 is an optional method for executing a positioning task. In the case where there is a positioning map, the camera pose is updated based on the GICP registration method, that is, positioning. In other alternative embodiments, the camera pose can also be updated by other specific update methods, which are not limited in this embodiment.

[0171] For another example, in an optional embodiment, in order to overcome the influence of illumination changes in the scene, the embodiment of the present application provides a specific implementation manner of S305, obtaining a positioning map and an update map based on a multi-scale illuminance perception Gaussian map fusion scheme, improving the robustness of the system. During the SLAM optimization and positioning process, the positioning map and the reference map are decoupled to solve the problem of system instability caused by simultaneous positioning and optimization.

[0172] For another example, S305 for obtaining the reference map and S307~S312 for executing the update task are optional steps for updating the historical environmental map and its illuminance label in the scene map table based on the currently collected target environmental map.

[0173] For another example, the present application is not limited to specific application scenarios. Specifically, it is very important for applications such as robots, drones, and intelligent security devices that need to move and work in complex environments. The following are several optional application scenario examples:

[0174] For a household floor cleaning robot, images of the home environment are captured by a single camera installed on the robot. Using monocular depth estimation technology, the distance (depth) of each image pixel is calculated to form three-dimensional point cloud data. Through Gaussian sputtering technology, these point cloud data are optimized into a high-precision home map. The robot can not only know the layout of the room but also know its own position in the room in real time, so as to intelligently plan the cleaning route, avoid obstacles, and perform efficient cleaning.

[0175] In the field of drone navigation, during flight, the drone uses the installed camera to capture environmental images in real time, obtains three-dimensional information of the environment through monocular depth estimation technology, and constructs a three-dimensional map using Gaussian sputtering technology. The drone can autonomously navigate in complex indoor or outdoor environments, avoid collisions, and plan the best flight path according to the real-time map to perform tasks such as express delivery and environmental monitoring.

[0176] In the field of intelligent security, the cameras installed in the security system can capture images of the monitored area in real time, generate three-dimensional point cloud data through monocular depth estimation technology, construct an environmental map through Gaussian sputtering technology, understand the layout and dynamic changes of the monitored area in real time, perform personnel positioning and behavior analysis, and improve the intelligent level of security monitoring.

[0177] It can be seen that the present invention combines monocular depth estimation technology and Gaussian sputtering technology to achieve environmental perception and autonomous positioning and navigation using a single RGB camera, providing a low-cost and effective visual SLAM solution for various application scenarios.

[0178] For another example, there are various specific methods for S305 to obtain the positioning map and the reference map. Figure 4 This application provides a specific implementation flowchart for obtaining the positioning map and the reference map. As Figure 4 shown, this method includes:

[0179] S401. Obtain the illuminance of the target monocular environmental image.

[0180] In this embodiment, there are various methods for analyzing the illuminance of the monocular environmental map to obtain the illuminance. For example, the illuminance can be obtained by arranging illuminance sensors in the scene, or the illuminance can be estimated by analyzing the illuminance characteristics of the target monocular environmental image. Among them, the specific method for estimating the illuminance by analyzing the target monocular environmental image is:

[0181] D1. Extract the luminance channel 𝐿 from the target monocular environmental image.

[0182] Specifically, convert the target monocular environmental image into a luminance image, and calculate the luminance channel 𝐿 using formula (9):

[0183] (9).

[0184] Calculate the mean value of the luminance channel using formula (10) :

[0185] (10).

[0186] Calculate the variance of the luminance channel using formula (11) :

[0187] (11);

[0188] where 𝐻 and 𝑊 are the height and width of the target monocular environmental image respectively.

[0189] D2. Convert the target monocular environmental image into the HSV color space .

[0190] Calculate the mean value of the hue channel using formula (12) :

[0191] (12).

[0192] Calculate the variance of the hue channel using formula (13) :

[0193] (13).

[0194] D3. Use formula (14) to perform weighted summation on each feature to obtain the illuminance.

[0195] (14);

[0196] where:

[0197] Contrast The calculation formula is formula (15):

[0198]

[0199] The calculation formula of the exposure value 𝐸𝑉 is formula (16):

[0200] (16);

[0201] Where: N is the aperture value, t is the shutter speed, and ISO is the sensitivity

[0202] Ratio of the shadow area The calculation formula is Formula (17):

[0203] (17).

[0204] Ratio of the highlight area The calculation formula is Formula (18):

[0205] (18);

[0206] Indicates the weight coefficient of each feature. The weight coefficient of each feature is adjusted according to the actual situation and experimental data to optimize the illuminance Accuracy.

[0207] In summary, after reading the monocular environmental map in this step, through the multi-scale analysis method, various features such as the luminance histogram, luminance mean and variance, hue and saturation, color balance, color distribution, shadow and highlight areas, exposure value (EV), and image contrast in the monocular environmental map are obtained, and the quantified illuminance value after synthesizing various features, that is, the illuminance, is calculated. The analysis of various features at different scales ensures the accurate perception and processing of complex lighting environments and improves the accuracy of illuminance.

[0208] S402. Compare the illuminance of the target monocular environmental image with the illuminance label of the target historical environmental map to obtain the illumination difference degree, and obtain the illumination matching map according to the illumination difference degree.

[0209] In this embodiment, the target historical environmental map includes at least one historical environmental map in the scene map table. For example, the target historical environmental map includes n environmental maps in the scene map table whose time interval between the time stamp and the current time is less than the preset time interval threshold. For another example, ignoring the time stamp, the historical environmental maps in the scene map table are directly used as the target historical environmental map.

[0210] In this embodiment, if the illumination difference degree is not less than the preset illumination difference degree threshold, the illumination matching result is unmatched; if the illumination difference degree is less than the preset illumination difference degree threshold, the illumination matching result is matched.

[0211] Specifically, the specific implementation method of this step includes:

[0212] E1. Obtain the illuminance difference threshold.

[0213] E2. Calculate the illumination difference degree.

[0214] Specifically, for each target historical environment map, the formula (19) is used to calculate the target illuminance of the target monocular environment image and the illuminance label of the target historical environment map to obtain the absolute value of the difference , which is used as the illumination difference degree.

[0215] (19).

[0216] E3. Obtain the minimum value of the illumination difference degrees between each target historical environment map and the target monocular environment image , and compare the minimum value with the illumination difference degree threshold. If , it is determined that there is an illumination matching map. Compare the illumination difference degrees of each target historical environment map with the illumination difference degree threshold, and use the target historical environment maps with illumination difference degrees less than the illumination difference degree threshold as the illumination matching maps.

[0217] It should be noted that if , it is determined that there is no illumination matching map, that is, there is no positioning map.

[0218] Furthermore, it should be noted that the illumination difference degree threshold is dynamically adjusted according to the severity of the current environmental illumination change , for example, when the system detects a large illumination change, the threshold can be appropriately increased; in a scene with relatively stable illumination, the threshold can be decreased to improve the accuracy of illumination matching.

[0219] S403. If the number of illumination matching maps is equal to 1, use this illumination matching map as the positioning map.

[0220] S404. If the number of illumination matching maps is greater than 1, obtain the two illumination matching maps with the smallest illumination difference degrees from the target monocular environment image, and use them as the first illumination matching map and the second illumination matching map respectively.

[0221] In this embodiment, based on the illumination difference degree result of E2, select the two illumination matching maps with the smallest illumination difference degrees from the target monocular environment image, that is, the first illumination matching map with the illuminance label and the second illumination matching map with the illuminance label . .

[0222] S405. Perform weighted fusion on the Gaussian basis elements at the same positions of the first illumination matching map and the second illumination matching map to obtain a set of fused Gaussian basis elements.

[0223] In this embodiment, the set of fused Gaussian basis elements includes the fused Gaussian basis elements at each position, and the first illumination matching map​ The weighting coefficient α and the second illumination matching map The weighting coefficient (1 - α) are both related to the illumination difference degree. The greater the illumination difference degree between the illumination matching map and the target monocular environmental image, the smaller the weighting coefficient.

[0224] Use formula (20) to determine the weighting coefficient:

[0225] (20);

[0226] Where , , represents the illumination label of the first illumination matching map of, represents the illumination label of the second illumination matching map of.

[0227] Use formula (21) to perform weighted fusion on the Gaussian basis element sets of the first illumination matching map and the second illumination matching map to calculate the Gaussian basis element set of the fusion map, that is, the fused Gaussian basis element set :

[0228] (21);

[0229] Where, and are respectively the Gaussian basis element sets of the first illumination matching map and the second illumination matching map.

[0230] S406. Construct a positioning map based on the fused Gaussian basis element set.

[0231] S407. Select an illumination matching map that meets the update conditions as the reference map.

[0232] In this embodiment, the update conditions include that the illumination difference degree is less than the update threshold, where the update threshold is less than the illumination difference threshold. Optionally, select the illumination matching map with the smallest illumination difference degree from the illumination matching maps with the illumination difference degree less than the update threshold as the reference map.

[0233] As can be seen from the above technical solution, this method uses the illuminance of the collected target monocular environmental image, that is, the target illuminance, as a benchmark to determine whether there is an environmental image in the historical environmental map that meets the illumination matching condition, that is, the illumination matching map. The illumination matching map is an environmental image with illumination conditions similar to those of the target monocular environmental image. Based on the illumination matching map, the positioning image is determined, avoiding low positioning accuracy caused by environmental illumination changes. Further, in the case of multiple illumination matching maps, the positioning image is determined by fusing multiple illumination matching maps (only two illumination matching maps are taken as examples in this embodiment), and the Gaussian sputtering characteristic is used to perform smooth transition and fine adjustment on the positioning image.

[0234] Thus, by fusing to obtain the positioning map, the system will select the two maps closest to the current illuminance for fusion. The fusion process includes weighted summation of the Gaussian basis element sets based on the weight coefficients to adjust and merge the Gaussian basis elements of the historical environmental maps to be fused (that is, the first illumination matching map and the second illumination matching map). Among them, the weight coefficients are determined based on the illumination difference degree. That is, the operation of fusing the Gaussian basis elements will be adjusted using the illuminance of the current target environmental map to implement the fusion operation of the illumination matching map of the current target environmental map, and the Gaussian sputtering characteristic is used for smooth transition and fine adjustment to generate a fine map reflecting the current scene. The fused map is used as the positioning map for the SLAM positioning task. Improve the adaptability of the positioning image used for positioning (camera pose update) to the environmental conditions of the current scene. Thus, improve the positioning accuracy and accuracy under the current illumination conditions.

[0235] Further, in the case of the existence of the illumination matching map, this method continues to select a historical environmental map as the reference image based on the update condition and enters the subsequent environmental image update task. That is, the reference image is updated based on the currently input target monocular environmental image to optimize the environmental image under the current illumination conditions. Thus, improve the accuracy of the subsequent positioning task.

[0236] Further, the positioning map and the reference map are obtained by different methods. That is, the acquisition methods of the positioning map and the reference map are decoupled to isolate the positioning map and the reference map, ensuring the stability of the slam algorithm.

[0237] In summary, the present application proposes a visual SLAM method combining monocular camera and GS technology. First, by establishing multiple environmental maps adapted to different lighting conditions, the positioning accuracy of visual SLAM under complex lighting conditions is significantly improved. Aiming at the impact of lighting changes on the positioning accuracy, the present invention further proposes a mechanism for separating the positioning map from the reference map, dynamically selects the map most matching the current lighting condition for positioning during the positioning process, and incrementally updates the environmental map in the background, thereby significantly improving the stability and reliability of visual SLAM during all-weather operation.

[0238] This method uses 3D Gaussian splatting as the representation of the scene map. Compared with traditional dense point clouds or mesh representations, 3D Gaussian splatting has significant advantages in terms of data storage and computational efficiency. Gaussian splatting more naturally describes the environmental structure through a continuous probability distribution, and at the same time supports efficient rendering and update operations, which is suitable for SLAM scenarios with high real-time requirements. Combining with the depth information generated by monocular depth estimation technology, 3D Gaussian splatting can accurately and efficiently express scene details, further enhancing the system's performance in complex environments. This method does not require an RGB-D camera, significantly reducing the hardware cost. At the same time, by optimizing the algorithm and data representation, it improves the lighting adaptation ability and can achieve accurate positioning and efficient map construction in dynamic and complex lighting environments.

[0239] The present invention can be widely applied in the field of artificial intelligence, especially suitable for intelligent mobile devices. By collecting monocular environmental images through a camera device and combining monocular depth estimation and Gaussian splatting technology, it constructs an environmental map in real time based on monocular images. While ensuring the accuracy and stability of visual SLAM, it realizes the functions of efficient positioning and environmental perception of the device during all-weather operation. The purpose of this invention is to provide a low-cost and highly applicable visual SLAM solution through the combination of monocular depth estimation and Gaussian splatting technology. Using monocular depth estimation technology, the depth information of each image pixel is inferred to generate three-dimensional point cloud data. Using Gaussian splatting technology to optimize the generated point cloud data instead of NeRF, it improves the processing speed of the system and ensures that the SLAM system can run at a rate of 5-10 frames per second, meeting the real-time positioning requirements of intelligent devices.

[0240] The above introduced a positioning method provided by an embodiment of the present application. Next, the device for executing the above positioning method will be introduced.

[0241] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a positioning device provided by an embodiment of the present application. As Figure 5 shown, the positioning device 500 includes:

[0242] A point cloud generation unit 501, configured to obtain 3D point cloud data of a target environmental image;

[0243] The primitive initialization unit 502 is configured to initialize the Gaussian primitives using the 3D point cloud data to obtain new Gaussian primitives;

[0244] The illumination matching unit 503 is configured to obtain, based on the target illumination intensity, a historical environment map that meets a preset illumination matching condition from the scene map table as an illumination matching map, where the illumination matching condition at least includes that the illumination difference degree between the illumination label and the target illumination intensity is less than a preset illumination difference degree threshold; the target illumination intensity is the environmental illumination intensity when the target environmental image is captured, the scene map table includes historical environment maps with illumination labels, and the illumination label of the historical environment map represents the environmental illumination intensity when the historical environment map is generated;

[0245] The positioning map acquisition unit 504 is configured to obtain a positioning map based on the illumination matching map;

[0246] The positioning unit 505 is configured to update the camera pose based on the new Gaussian primitives and the Gaussian primitives of the positioning map.

[0247] In a possible implementation, when the point cloud generation unit is configured to obtain the 3D point cloud data of the target environmental image, it is specifically configured to:

[0248] Receive a monocular environmental image collected by a monocular image acquisition device as the target environmental image;

[0249] Obtain a depth map of the target environmental image through a pre-trained depth estimation model;

[0250] Map the depth map of the target environmental image to a 3D point cloud based on preset camera internal parameters to obtain the 3D point cloud data.

[0251] In a possible implementation, when the illumination matching unit is configured to, it is specifically configured to:

[0252] Obtain, based on the target illumination intensity, a historical environment map that meets a preset illumination matching condition from the scene map table, including:

[0253] Obtain the illumination intensity of the target environmental image based on a preset illumination feature;

[0254] For each target historical environment map, compare the illumination label of the target historical environment map with the target illumination intensity to obtain the illumination difference degree of the target historical environment map, where the target historical environment map includes at least one historical environment map in the scene map table;

[0255] For each target historical environment map, compare the light difference degree with the light difference degree threshold. If the light difference degree is less than the light difference degree threshold, determine that the target historical environment map meets the light matching condition; if the light difference degree is not less than the light difference degree threshold, determine that the target historical environment map does not meet the light matching condition.

[0256] In a possible implementation, the positioning map acquisition unit is used to obtain a positioning map based on the light matching map, and specifically is used for:

[0257] If the number of the light matching maps is equal to 1, use the light matching map as the positioning map;

[0258] If the number of the light matching maps is greater than 1, obtain two historical environment maps with the smallest light difference degree from all the light matching maps, and use them as the first light matching map and the second light matching map respectively;

[0259] Perform weighted fusion on the Gaussian basis elements at the same positions of the first light matching map and the second light matching map respectively based on the corresponding weighted coefficients to obtain a set of fused Gaussian basis elements. The weighted coefficient of the Gaussian basis element of the target light matching map is inversely correlated with the light difference degree of the target light matching map. The target light matching map includes the first light matching map and the second light matching map;

[0260] Construct the positioning map based on the set of fused Gaussian basis elements.

[0261] In a possible implementation, when the positioning unit is used to update the camera pose based on the new Gaussian basis element and the Gaussian basis element of the positioning map, it is specifically used for:

[0262] If there is a positioning map, use a preset point cloud-based registration algorithm to register the new Gaussian basis element and the Gaussian basis element of the positioning map to obtain a rotation matrix;

[0263] Use the rotation matrix to update the camera pose.

[0264] In a possible implementation, the positioning device further includes an update unit, which is used to, after obtaining a historical environment map that meets the preset light matching condition from the scene map table based on the target illuminance, if there is a light matching map that meets the preset update condition, select a light matching map that meets the update condition as a reference map. The update condition includes that the light difference degree is less than a preset update threshold, and the update threshold is less than the light difference degree threshold; update the reference map based on the new Gaussian basis element; update the illuminance label of the reference map based on the target illuminance.

[0265] In a possible implementation, when the updating unit is used to update the reference map based on the new Gaussian basis elements, it is specifically configured to: add the new Gaussian basis elements to the reference map to obtain a first candidate map; perform heuristic pruning on the number of Gaussian basis elements in the first candidate map to obtain a second candidate map; based on the second candidate map, use a differentiable rendering path to render a rendered image in the camera pose; compare the rendered image with the target monocular environment image, calculate a loss value, backpropagate the loss value, and calculate the gradient corresponding to each Gaussian basis element in the second candidate map; update each Gaussian basis element based on the gradient corresponding to each Gaussian basis element in the second candidate map to obtain an environment map of the target environment image; and update the reference map to the environment map of the target environment image.

[0266] In a possible implementation, the initializing map unit is configured to, after obtaining a historical environment map that meets a preset illumination matching condition from a scene map table based on the target illumination, if there is no illumination matching map that meets the update condition, construct a new environment map based on the new Gaussian basis elements; and store the new environment map in the scene map table with the target illumination as the illumination label.

[0267] An electronic device is also provided in an embodiment of the present application. Refer to Figure 6 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 6 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiment of the present application.

[0268] As Figure 6 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0269] Typically, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0270] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the positioning methods provided by the embodiments of the present application.

[0271] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement any one of the positioning methods provided by the embodiments of the present application.

[0272] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0273] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits or dedicated circuits, etc. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0274] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0275] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device or data center to another website, computer, training device or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A positioning method, characterized in that: include: Obtain 3D point cloud data of the target environment image; Initialize the Gaussian primitives using the 3D point cloud data to obtain new Gaussian primitives; Based on the target illumination, a historical environment map satisfying a preset illumination matching condition is obtained from a scene map table as an illumination matching map, wherein the illumination matching condition at least includes that an illumination difference between an illumination label and a target illumination is less than a preset illumination difference threshold; the target illumination is the ambient illumination when the target environment image is captured, the scene map table includes a historical environment map with an illumination label, and the illumination label of the historical environment map indicates the ambient illumination when the historical environment map is generated; Based on the illumination matching map, obtaining a positioning map; The camera pose is updated based on the new Gaussian primitives and the Gaussian primitives of the positioning map.

2. The positioning method according to claim 1, characterized in that: The 3D point cloud data of the target environment image is obtained by: Receiving a monocular environment image acquired by a monocular image acquisition device as a target environment image; Obtain a depth map of the target environment image through a pre-trained depth estimation model; Based on preset camera intrinsic parameters, the depth map of the target environment image is mapped into a 3D point cloud to obtain the 3D point cloud data.

3. The positioning method according to claim 1, characterized in that: The method of obtaining a historical environment map that meets a preset illumination matching condition from a scene map table based on the target illumination includes: Based on preset lighting characteristics, obtaining the illumination of the target environment image; For each target historical environment map, comparing the illumination label of the target historical environment map with the target illumination to obtain the illumination difference of the target historical environment map, wherein the target historical environment map includes at least one historical environment map in the scene map table; For each target historical environment map, the illumination difference is compared with the illumination difference threshold. If the illumination difference is less than the illumination difference threshold, it is determined that the target historical environment map meets the illumination matching condition; if the illumination difference is not less than the illumination difference threshold, it is determined that the target historical environment map does not meet the illumination matching condition.

4. The positioning method according to claim 1, characterized in that: The acquiring of a positioning map based on the illumination matching map comprises: If the number of the illumination matching maps is equal to 1, the illumination matching map is used as the positioning map; If the number of the illumination matching maps is greater than 1, two historical environment maps with the smallest illumination difference are obtained from all the illumination matching maps, and the two maps are used as the first illumination matching map and the second illumination matching map respectively; performing weighted fusion on Gaussian primitives at the same position of the first illumination matching map and the second illumination matching map based on corresponding weighted coefficients to obtain a fused Gaussian primitive set, wherein the weighted coefficient of the Gaussian primitive of the target illumination matching map is inversely correlated with the illumination difference of the target illumination matching map, and the target illumination matching map includes the first illumination matching map and the second illumination matching map; The positioning map is constructed based on the fused Gaussian primitive set.

5. The positioning method according to claim 1, characterized in that: The updating of the camera pose based on the new Gaussian primitives and the Gaussian primitives of the positioning map comprises: If the positioning map exists, using a preset point cloud-based registration algorithm, the new Gaussian basis element is registered with the Gaussian basis element of the positioning map to obtain a rotation matrix; Update the camera pose using the rotation matrix.

6. The positioning method according to claim 5, characterized in that: After acquiring the historical environment map that meets the preset illumination matching condition from the scene map table based on the target illumination, the positioning method further includes: If there is an illumination matching map that satisfies a preset update condition, selecting an illumination matching map that satisfies the update condition as a reference map, wherein the update condition includes that the illumination difference is less than a preset update threshold, and the update threshold is less than the illumination difference threshold; Based on the new Gaussian primitive, updating the reference map; Based on the target illuminance, an illuminance tag of the reference map is updated.

7. The positioning method according to claim 6, characterized in that: The updating of the reference map based on the new Gaussian primitives comprises: adding the new high-base primitive to the reference map to obtain a first candidate map; Heuristically pruning the number of Gaussian primitives in the first candidate map to obtain a second candidate map; Based on the second candidate map, using a differentiable rendering path to render and obtain a rendered image at the camera pose; Comparing the rendered image with the target monocular environment image, calculating a loss value, back-propagating the loss value, and calculating a gradient corresponding to each Gaussian basis element in the second candidate map; updating each Gaussian primitive based on the gradient corresponding to each Gaussian primitive in the second candidate map to obtain an environment map of the target environment image; The reference map is updated to be an environment map of the target environment image.

8. The positioning method according to claim 6, characterized in that: After acquiring the historical environment map that meets the preset illumination matching condition from the scene map table based on the target illumination, the positioning method further includes: If there is no illumination matching map that meets the update condition, construct a new environment map based on the new Gaussian primitive; The new environment map is stored in the scene map table with the target illuminance as the illuminance tag.

9. A positioning device, characterized in that: include: A point cloud generation unit, used to obtain 3D point cloud data of a target environment image; A primitive initialization unit, used for initializing Gaussian primitives using the 3D point cloud data to obtain new Gaussian primitives; An illumination matching unit is used to obtain, based on the target illumination, a historical environment map that meets a preset illumination matching condition from a scene map table as the illumination matching map, wherein the illumination matching condition at least includes that an illumination difference between an illumination label and a target illumination is less than a preset illumination difference threshold; the target illumination is the ambient illumination when the target environment image is captured, the scene map table includes a historical environment map with an illumination label, and the illumination label of the historical environment map indicates the ambient illumination when the historical environment map is generated; A positioning map acquisition unit, configured to acquire a positioning map based on the illumination matching map; A positioning unit is used to update the camera pose based on the new Gaussian primitives and the Gaussian primitives of the positioning map.

10. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the positioning method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Indoor mobile robot adaptive fusion positioning mapping voice interaction method and system

    CN121702384A