Robot indoor navigation method and system
By combining sparse visual odometers and neural radiation fields and jointly optimizing, the problem of high computing resource consumption of robot navigation systems during three-dimensional reconstruction is solved, and high-quality three-dimensional scene reconstruction and real-time robot navigation are achieved.
Patent Information
- Application Number
- CN202510049836.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-16
AI Technical Summary
The existing robot navigation system consumes high computing resources during three-dimensional reconstruction and cannot meet the real-time requirements.
Combining sparse visual odometer and neural radiation field, dense geometric enhancement is performed through a deep neural network of Transformer architecture, and neural radiation field parameters and camera poses are jointly optimized to achieve high-quality three-dimensional scene reconstruction and real-time robot navigation.
It realizes accurate camera pose estimation and high-quality three-dimensional scene reconstruction, reduces computing overhead and memory consumption, ensures real-time visual positioning and map construction, and is suitable for hardware platforms with limited resources.
Smart Images

Figure CN120014159A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and in particular relates to a robot indoor navigation method and system. Background Art
[0002] With the widespread application of intelligent robots in complex indoor environments, how to accurately locate and map the environment through visual information has become one of the key technologies. Existing visual odometry systems mainly rely on sparse feature matching, estimate the camera pose through image sequences, and generate environmental maps. However, these methods usually face problems such as insufficient sparse feature points, large influence of occlusion and illumination changes. In order to solve these problems, dense visual SLAM technology came into being, which generates more accurate three-dimensional scenes through dense depth maps and normal information. However, existing dense visual SLAM methods rely on a lot of computing resources and perform poorly in real-time and memory consumption.
[0003] At the same time, NeRF, as an efficient scene representation method, can retain high-fidelity photometric and geometric information when generating images, and achieve high-quality 3D scene reconstruction. However, NeRF has a large amount of computation and slow rendering speed, which makes it difficult to meet the needs of real-time robot navigation. Currently, although NeRF is combined with SLAM for 3D scene reconstruction, most of them use RGB-D sensors for depth information estimation, which consumes a lot of actual computing resources and cannot meet real-time requirements. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a robot indoor navigation method and system, which are used to solve the problem that the current robot navigation three-dimensional reconstruction computing resource consumption is high and cannot meet the real-time requirements.
[0005] In a first aspect of an embodiment of the present invention, a robot indoor navigation method is provided, comprising: Obtain the RGB image frame sequence collected by the camera, and estimate the camera pose and sparse 3D landmark points of the image frame through the depth block visual odometer; Based on the camera pose and sparse three-dimensional landmark points, dense geometric enhancement is performed on the image frame through a deep neural network with a transformer architecture to obtain dense geometric prior information, wherein the dense geometric prior information includes a dense depth map and a normal map; The neural radiance field is trained based on the initial camera pose, dense geometric prior information and image frame data, and the neural radiance field parameters and camera pose are jointly optimized based on photometric consistency, geometric consistency and camera pose constraints. Based on the image frames acquired in real time, novel views are synthesized through the optimized neural radiation field, and the robot navigates based on the novel views.
[0006] In a second aspect of an embodiment of the present invention, a robot indoor navigation system is provided, comprising: The visual tracking module is used to obtain the RGB image frame sequence collected by the camera and estimate the camera pose and sparse 3D landmark points of the image frame through the depth block visual odometer; A geometric enhancement module is used to perform dense geometric enhancement on image frames based on camera pose and sparse three-dimensional landmark points through a deep neural network with a transformer architecture to obtain dense geometric prior information, wherein the dense geometric prior information includes a dense depth map and a normal map; A training optimization module is used to train the neural radiance field based on the initial camera pose, dense geometric prior information and image frame data, and jointly optimize the neural radiance field parameters and camera pose based on photometric consistency, geometric consistency and camera pose constraints; The visual navigation module is used to synthesize novel views based on the image frames acquired in real time through the optimized neural radiation field, and perform robot navigation based on the novel views.
[0007] In a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect of the embodiment of the present invention when executing the computer program.
[0008] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method provided in the first aspect of the embodiment of the present invention are implemented.
[0009] In an embodiment of the present invention, by combining sparse visual odometry and neural radiation field and jointly optimizing the neural radiation field, accurate camera pose estimation and high-quality three-dimensional scene reconstruction can be achieved. This not only has low computational overhead and low memory consumption, but also can ensure the real-time performance of visual positioning and mapping. At the same time, it can effectively cope with complex factors in dynamic indoor environments, such as lighting changes, occlusions, and low-texture areas, and provide stable and reliable positioning and mapping support for robot navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0011] Figure 1A schematic diagram of a flow chart of a robot indoor navigation method provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a robot indoor navigation system provided by one embodiment of the present invention; Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0012] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0013] It should be understood that the term "including" and other similar expressions in the specification or claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions, such as a process, method, system, or device including a series of steps or units is not limited to the listed steps or units. In addition, "first" and "second" are used to distinguish different objects, not to describe a specific order.
[0014] See also Figure 1 , a schematic flow chart of a robot indoor navigation method provided by an embodiment of the present invention includes: S101, obtaining a sequence of RGB image frames collected by a camera, and estimating the camera pose and sparse three-dimensional landmark points of the image frames through a depth block visual odometer; Obtain a sequence of RGB image frames captured by a monocular camera on the robot, and process each input frame as a potential key frame. It can be understood that this embodiment is applicable to a robot with a monocular RGB camera, and can achieve high-precision pose estimation and environment reconstruction without additional depth or inertial information.
[0015] Optionally, in each input frame, randomly select k The size is s × s square patches; The patches are added to the patch map as sparse features, where i The first frame in the image input k A patch can be represented as , where u and v are pixel coordinates, d T is the inverse depth value with respect to u and v.
[0016] For keyframes iFor each patch of j The reprojection error is calculated by the following formula: ; Where K represents the camera intrinsic parameter matrix, T i Yes Frame i The transformation matrix to the world coordinate system, is the patch location; Predicting optical flow correction vectors on patches via a recurrent neural network , and optimize it by the least squares method to update the camera pose and depth value of each patch; Each sliding window keeps the three most recent frames as key frames, and if the optical flow intensity between the fourth frame and the previous frame is lower than the set threshold, the frame is discarded.
[0017] The Deep Patch Visual Odometry (DPVO) is an odometer method that combines deep learning and traditional visual odometry technology. It uses a deep learning model to predict the camera's pose change, thereby estimating the camera's motion trajectory. The Deep Patch Visual Odometry can calculate the camera motion through optical flow estimation between consecutive image frames and extract sparse depth information, thereby providing a preliminary pose estimate for subsequent neural radiance field optimization.
[0018] Among them, the motion information between consecutive image frames is extracted through optical flow estimation, and the initial camera pose of each frame is generated through sparse depth estimation and graph optimization based on the motion information; the sparse depth information of the image frame is extracted, and the sparse three-dimensional landmark points are estimated based on the sparse depth information and the geometric constraints between frames.
[0019] Inter-frame geometric constraints refer to the geometric constraint relationship between geometric feature points in adjacent frames, which can be determined based on the projection position of the baseline and the point on the polar line in polar geometry. Sparse 3D landmark points are a set of discrete feature points in 3D space used for positioning, mapping or visual positioning, which can be selected by visual positioning and other methods.
[0020] S102, based on the camera pose and the sparse three-dimensional landmark points, performing dense geometric enhancement on the image frame through a deep neural network of a transformer architecture to obtain dense geometric prior information, wherein the dense geometric prior information includes a dense depth map and a normal map; Dense geometry enhancement refers to improving the density and accuracy of geometric information in 3D reconstruction or visual odometer through algorithms. In this embodiment, dense geometry enhancement is performed by using a deep neural network with a Transformer architecture to obtain dense geometric prior information.
[0021] Among them, the dense depth map corresponding to the monocular image frame is predicted through the deep neural network of the Transformer architecture; based on the predicted dense depth map and sparse three-dimensional landmark points, the depth distribution of the depth map is adjusted through the scale alignment mechanism; the normal information is inferred according to the adjusted dense depth map, and the normal is used as the scene geometry constraint to obtain the normal map.
[0022] The scale alignment mechanism is to adjust the feature maps to the same scale for processing when processing feature maps of different resolutions or scales, such as depth adjustment.
[0023] In some embodiments, a Dense Prediction Transformer (DPT) network is used to predict and generate a dense depth map to represent the depth information of the entire scene.
[0024] To align the dense depth with the sparse depth, assume that the sparse depth D s and dense depth D d All follow Gaussian distribution , , the scale factor is calculated by α and offset β : ; The dense depth Output the alignment depth by aligning with the formula: .
[0025] S103, training the neural radiation field according to the initial camera pose, dense geometric prior information and image frame data, and jointly optimizing the neural radiation field parameters and the camera pose based on photometric consistency, geometric consistency and camera pose constraints; Neural radiance field is a fully connected neural network that can generate 3D scene models based on partial 2D images through learning. It can use rendering loss to train the network to reproduce the perspective of the scene. Based on the initial camera pose and RGB image data of the image frame, the neural radiance field is used to establish an implicit 3D scene representation, and the neural radiance field is parameterized based on dense geometric prior information.
[0026] Specifically, the three-dimensional space coordinates and viewing angle information are positionally encoded; based on the position encoding, the volume density and color information are modeled through a multi-layer perceptron; through light sampling and volume rendering, an implicit representation of the three-dimensional scene is generated; based on the input image frames, as well as the supervisory signals of the dense depth map and normal map, the model parameters of the neural radiation field are gradually optimized.
[0027] In some embodiments, NeRF uses a neural radiation field representation based on the Nerfacto model (including multi-resolution hash coding, spherical harmonic coding, etc.), which can represent the geometry and lighting information of the scene. The input of the neural radiation field is the three-dimensional space coordinates (x, y, z) and the viewing direction. θ The output is the volume density and color value of each point.
[0028] Photometric consistency means that the luminosity (i.e. grayscale value and RGB) of the same point or patch between two frames of images hardly changes; geometric consistency means that the scale (i.e. size) of the same static point hardly changes between adjacent frames; camera pose constraints refer to the geometric constraints that exist between different camera poses.
[0029] Specifically, based on the photometric consistency loss between the rendered image and the input image, the color representation of the scene is optimized; based on the geometric consistency loss between the rendered depth map and the dense depth map, the depth representation of the scene is optimized; based on the geometric constraints of the initial camera pose and the pose optimization loss, the camera pose is corrected; according to the photometric consistency loss, geometric consistency loss and pose optimization loss, each loss function is weightedly combined to minimize the total loss, and the neural radiation field parameters and the camera pose are jointly optimized.
[0030] Through volume rendering of neural radiance fields, the generated image of the scene at a given camera pose is calculated, compared with the real image, and jointly optimized using the following loss function:
[0031] In the formula, color loss L rgb The depth loss is obtained by calculating the mean square error of the color values between the generated image and the real image. L d Defined as the KL divergence between the depth map of the climate and the generated depth map, the normal loss L n is the cosine similarity between the predicted normal and the true normal, L reg is the regularization loss. The MLP parameters of the neural radiance field and the camera pose are iteratively updated through the Adam optimizer.
[0032] S104. Based on the image frames acquired in real time, novel views are synthesized through the optimized neural radiation field, and robot navigation is performed based on the novel views.
[0033] The neural radiance field after training and optimization can output views of different perspectives for robot navigation based on the camera pose corresponding to the image frame. Novel View Synthesis (NVS) is an image-based rendering method that can transform or interpolate the perspective of a partial view of a given scene and synthesize images of other new perspectives of the scene. The neural radiance field can achieve high-quality view synthesis by learning the implicit neural representation of the three-dimensional scene. According to the real-time position and orientation of the camera, it generates a scene representation image from that perspective, that is, it generates a novel view for robot navigation.
[0034] Specifically, the camera pose of the image frame is updated based on the depth block visual odometry; a novel view is generated based on the implicit three-dimensional scene representation of the optimized neural radiation field for inference of unobserved areas and environment completion; and the robot is path planned and actively avoids obstacles based on the novel view and the camera pose.
[0035] In this embodiment, by combining sparse visual odometry with neural radiation fields, high-precision pose estimation and high-fidelity three-dimensional environment mapping can be achieved in complex environments. By jointly optimizing the camera pose and NeRF representation, the influence of factors such as illumination changes and texture sparsity can be effectively overcome. At the same time, by optimizing computing resource consumption and memory usage, real-time robot navigation and mapping can be achieved under lower hardware requirements. Compared with traditional NeRF-based mapping methods, the camera tracking frequency can be significantly improved to meet real-time requirements. By jointly optimizing the neural radiation field, memory consumption can be reduced, so that the system can run efficiently on conventional GPUs or embedded devices, and is suitable for hardware platforms with limited resources.
[0036] The combination of NeRF-VO can effectively cope with complex factors in dynamic indoor environments, such as lighting changes, occlusions, and low-texture areas, thereby providing stable and reliable positioning and mapping support for robots. Based on sparse visual odometry and deep neural network architecture, it can reduce computational overhead, and significantly reduce computing resource consumption compared to traditional SLAM and visual odometry methods.
[0037] It should be understood that the serial numbers of the steps in the above embodiments do not imply a sequence of execution, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0038] Figure 2 A schematic diagram of the structure of a robot indoor navigation system provided by an embodiment of the present invention, the system comprising: The visual tracking module 210 is used to obtain the RGB image frame sequence collected by the camera and estimate the camera pose and sparse 3D landmark points of the image frame through the depth block visual odometer; Wherein, the visual tracking module 210 includes: A pose estimation unit is used to extract motion information between consecutive image frames through optical flow estimation, and generate the initial camera pose of each frame of image through sparse depth estimation and graph optimization based on the motion information; The landmark point estimation unit is used to extract sparse depth information of the image frame and estimate sparse three-dimensional landmark points based on the sparse depth information and inter-frame geometric constraints.
[0039] A geometry enhancement module 220 is used to perform dense geometry enhancement on the image frame based on the camera pose and sparse three-dimensional landmark points through a deep neural network of a transformer architecture to obtain dense geometry prior information, wherein the dense geometry prior information includes a dense depth map and a normal map; Wherein, the geometric enhancement module 220 includes: The depth prediction unit is used to predict the dense depth map corresponding to the monocular image frame through the deep neural network of the Transformer architecture; A depth alignment unit, used to adjust the depth distribution of the depth map through a scale alignment mechanism based on the predicted dense depth map and sparse 3D landmark points; The normal inference unit is used to infer normal information according to the adjusted dense depth map, and use the normal as a scene geometry constraint to obtain a normal map.
[0040] A training optimization module 230, for training the neural radiance field according to the initial camera pose, dense geometric prior information and image frame data, and jointly optimizing the neural radiance field parameters and the camera pose based on photometric consistency, geometric consistency and camera pose constraints; Specifically, the training of the neural radiation field according to the initial camera pose, dense geometric prior information and image frame data includes: Position encoding of three-dimensional space coordinates and viewing angle information; Based on position encoding, volume density and color information are modeled through multi-layer perceptron; Generate an implicit representation of the 3D scene through ray sampling and volume rendering; The model parameters of the neural radiance field are gradually optimized based on the input image frames and the supervision signals of the dense depth map and normal map.
[0041] Among them, the color representation of the scene is optimized based on the photometric consistency loss between the rendered image and the input image; Optimize the depth representation of the scene based on the geometric consistency loss between the rendered depth map and the dense depth map; Correct the camera pose based on the geometric constraints of the initial camera pose and the pose optimization loss; According to the photometric consistency loss, geometric consistency loss and pose optimization loss, each loss function is weightedly combined to minimize the total loss and jointly optimize the neural radiation field parameters and camera pose.
[0042] The visual navigation module 240 is used to synthesize novel views based on the image frames collected in real time through the optimized neural radiation field, and perform robot navigation based on the novel views.
[0043] Wherein, the visual navigation module includes: A pose updating unit, configured to update a camera pose of an image frame based on the depth block visual odometer; A view generation unit for generating novel views based on the implicit 3D scene representation of the optimized neural radiance field for inference of unobserved areas and environment completion; The visual navigation unit is used to perform path planning and active obstacle avoidance for the robot based on novel views and camera poses.
[0044] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and modules can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0045] Figure 3 1 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. The electronic device is used for indoor navigation of a robot. Figure 3 As shown, the electronic device 3 of this embodiment includes: a memory 310, a processor 320 and a system bus 330, wherein the memory 310 includes an executable program 3101 stored thereon, and those skilled in the art can understand that Figure 3 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0046] Combine the following Figure 3 A detailed introduction to the various components of electronic equipment: The memory 310 can be used to store software programs and modules. The processor 320 executes various functional applications and data processing of the electronic device by running the software programs and modules stored in the memory 310. The memory 310 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device (such as cache data), etc. In addition, the memory 310 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0047] The memory 310 contains an executable program 3101 of a robot navigation method, and the executable program 3101 can be divided into one or more modules / units, which are stored in the memory 310 and executed by the processor 320 to implement robot indoor navigation, etc. The one or more modules / units can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 3101 in the electronic device 5. For example, the computer program 3101 can be divided into functional modules such as a visual tracking module, a geometric enhancement module, a training optimization module, and a visual navigation module.
[0048] The processor 320 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 310, and calling data stored in the memory 310, it performs various functions of the electronic device and processes data, thereby monitoring the overall status of the electronic device. Optionally, the processor 320 may include one or more processing units; preferably, the processor 320 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 320.
[0049] The system bus 330 is used to connect the various functional components inside the computer, and can transmit data information, address information, and control information. Its types can be, for example, PCI bus, ISA bus, CAN bus, etc. The instructions of the processor 320 are transmitted to the memory 310 through the bus, and the memory 310 feeds back data to the processor 320. The system bus 330 is responsible for the data and instruction exchange between the processor 320 and the memory 310. Of course, the system bus 330 can also be connected to other devices, such as network interfaces, display devices, etc.
[0050] In the embodiment of the present invention, the executable program executed by the processor 320 included in the electronic device includes: Obtain the RGB image frame sequence collected by the camera, and estimate the camera pose and sparse 3D landmark points of the image frame through the depth block visual odometer; Based on the camera pose and sparse three-dimensional landmark points, dense geometric enhancement is performed on the image frame through a deep neural network with a transformer architecture to obtain dense geometric prior information, wherein the dense geometric prior information includes a dense depth map and a normal map; The neural radiance field is trained based on the initial camera pose, dense geometric prior information and image frame data, and the neural radiance field parameters and camera pose are jointly optimized based on photometric consistency, geometric consistency and camera pose constraints. Based on the image frames acquired in real time, novel views are synthesized through the optimized neural radiation field, and the robot navigates based on the novel views.
[0051] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0052] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0053] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A robot indoor navigation method, characterized in that: include: Obtain the RGB image frame sequence collected by the camera, and estimate the camera pose and sparse 3D landmark points of the image frame through the depth block visual odometer; Based on the camera pose and sparse three-dimensional landmark points, dense geometric enhancement is performed on the image frame through a deep neural network with a transformer architecture to obtain dense geometric prior information, wherein the dense geometric prior information includes a dense depth map and a normal map; The neural radiance field is trained based on the initial camera pose, dense geometric prior information and image frame data, and the neural radiance field parameters and camera pose are jointly optimized based on photometric consistency, geometric consistency and camera pose constraints. Based on the image frames acquired in real time, novel views are synthesized through the optimized neural radiation field, and the robot navigates based on the novel views.
2. The method according to claim 1, characterized in that The method of estimating the camera pose and sparse three-dimensional landmark points of the image frame by using the depth block visual odometer includes: The motion information between consecutive image frames is extracted through optical flow estimation, and the initial camera pose of each frame is generated through sparse depth estimation and graph optimization based on the motion information; The sparse depth information of the image frames is extracted, and the sparse 3D landmark points are estimated based on the sparse depth information and the geometric constraints between frames.
3. The method according to claim 1, characterized in that The dense geometric enhancement of the image frame by the deep neural network of the transformer architecture to obtain the dense geometric prior information includes: Predict the dense depth map corresponding to the monocular image frame through the deep neural network of the Transformer architecture; Based on the predicted dense depth map and sparse 3D landmark points, the depth distribution of the depth map is adjusted through a scale alignment mechanism; The normal information is inferred according to the adjusted dense depth map, and the normal is used as the scene geometry constraint to obtain the normal map.
4. The method according to claim 1, characterized in that: The training of the neural radiation field according to the initial camera pose, dense geometric prior information and image frame data includes: Position encoding of three-dimensional space coordinates and viewing angle information; Based on position encoding, volume density and color information are modeled through multi-layer perceptron; Generate an implicit representation of the 3D scene through ray sampling and volume rendering; The model parameters of the neural radiance field are gradually optimized based on the input image frames and the supervision signals of the dense depth map and normal map.
5. The method according to claim 1, characterized in that The joint optimization of neural radiation field parameters and camera pose based on photometric consistency, geometric consistency and camera pose constraints includes: Optimize the color representation of the scene based on the photometric consistency loss between the rendered image and the input image; Optimize the depth representation of the scene based on the geometric consistency loss between the rendered depth map and the dense depth map; Correct the camera pose based on the geometric constraints of the initial camera pose and the pose optimization loss; According to the photometric consistency loss, geometric consistency loss and pose optimization loss, each loss function is weightedly combined to minimize the total loss, and the neural radiation field parameters and camera pose are jointly optimized.
6. The method according to claim 1, characterized in that The method of synthesizing novel views based on the real-time acquired image frames through the optimized neural radiation field and performing robot navigation based on the novel views includes: Updating a camera pose of an image frame based on the depth block visual odometry; Generate novel views based on the implicit 3D scene representation of the optimized neural radiance field for inference and environment completion of unobserved areas; Based on novel views and camera poses, the robot can plan its path and avoid obstacles actively.
7. A robot indoor navigation system, characterized in that: include: The visual tracking module is used to obtain the RGB image frame sequence collected by the camera and estimate the camera pose and sparse 3D landmark points of the image frame through the depth block visual odometer; A geometric enhancement module is used to perform dense geometric enhancement on image frames based on camera pose and sparse three-dimensional landmark points through a deep neural network with a transformer architecture to obtain dense geometric prior information, wherein the dense geometric prior information includes a dense depth map and a normal map; A training optimization module is used to train the neural radiance field based on the initial camera pose, dense geometric prior information and image frame data, and jointly optimize the neural radiance field parameters and camera pose based on photometric consistency, geometric consistency and camera pose constraints; The visual navigation module is used to synthesize novel views based on the image frames acquired in real time through the optimized neural radiation field, and perform robot navigation based on the novel views.
8. The system according to claim 7, characterized in that The visual tracking module comprises: A pose estimation unit is used to extract motion information between consecutive image frames through optical flow estimation, and generate the initial camera pose of each frame of image through sparse depth estimation and graph optimization based on the motion information; The landmark point estimation unit is used to extract sparse depth information of the image frame and estimate sparse three-dimensional landmark points based on the sparse depth information and inter-frame geometric constraints.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of a robot indoor navigation method as described in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed, the steps of a robot indoor navigation method as claimed in any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Distributed multi-agent cooperation method based on mixed implicit neural field
CN120562462A
Physical object three-dimensional positioning method and system based on video data
CN121236170A
Method and system for three-dimensional positioning of physical objects based on video data
CN121236170B