An autonomous navigation method and related devices based on the fusion of laser and vision

By using laser and vision fusion method in the autonomous navigation system, multi-sensor data is processed and semantic maps are generated, the problem of route planning in complex dynamic environments is solved, and high-precision dynamic obstacle recognition and obstacle avoidance capabilities are achieved.

CN119780955BActive Publication Date: 2025-06-13SHENZHEN GREAT WORKER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510287594.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-13
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing autonomous navigation technology cannot effectively plan routes in complex dynamic environments, and faces the problem of insufficient spatial and temporal alignment accuracy of multi-sensor data and difficult to guarantee the real-time and consistency of semantic maps under dynamic object interference.

Method used

Using an autonomous navigation method based on the fusion of laser and vision, multi-sensor space-time joint synchronization processing is performed to generate multi-modal data streams with space-time aligned multi-modal data streams by acquiring laser point cloud data of lidar, binocular image data of binocular cameras, and IMU clock signals. Then, through the fusion process of cross-modal feature pyramid and attention-guided attention, dense environmental characterization with semantic labels and spatial and temporal masks of dynamic objects are generated. Next, a multi-layer semantic map is determined through hierarchical factor graph optimization and fused with the neural radiation field to generate a mixed map of the explicit semantic map and the implicit radiation field. Finally, the movement trajectory of the target device is determined by combining space-time constraints on the space-time masks of mixed maps and dynamic objects.

Benefits of technology

Achieve high-precision real-time distinction between dynamic obstacles and static structures in dynamic and complex environments, significantly improving the obstacle avoidance robustness of the motion trajectory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119780955B_ABST
    Figure CN119780955B_ABST
Patent Text Reader

Abstract

The present application provides an autonomous navigation method and related devices based on the fusion of laser and vision. The method includes: performing multi-sensor spatio-temporal joint synchronization processing on laser point cloud data, binocular image data, and clock signals to generate spatio-temporally aligned multi-modal data streams; processing the multi-modal data streams through a cross-modal feature pyramid and attention-guided fusion to generate a dense environment representation and a spatio-temporal mask of dynamic objects; determining a multi-layer semantic map corresponding to the dense environment representation through hierarchical factor graph optimization; fusing the multi-layer semantic map with a neural radiance field to generate a hybrid map of an explicit semantic map and an implicit radiance field; and determining the motion trajectory of the target device through spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of dynamic objects. By implementing the solution of the present application, it is possible to achieve real-time distinction between dynamic obstacles and static structures in a dynamic and complex environment, thereby significantly improving the obstacle avoidance robustness of the motion trajectory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data fusion technology, and in particular, to an autonomous navigation method and related devices based on the fusion of laser and vision. Background Art

[0002] With the rapid development of robot and autonomous driving technologies, the accuracy and stability of autonomous navigation systems have become increasingly important. Traditional autonomous navigation methods mainly rely on single-sensor data, such as pure lidar or vision cameras. These methods often show obvious limitations in complex and changing environments.

[0003] In complex dynamic environments, existing autonomous navigation technologies face key challenges such as insufficient spatio-temporal alignment accuracy of multi-sensor data and difficulties in ensuring the real-time performance and consistency of semantic maps under the interference of dynamic objects. Traditional methods rely on single-modal sensors (such as lidar or vision) to build environmental models, resulting in serious fragmentation of geometric information and semantic understanding in scenarios such as sudden changes in lighting and frequent movement of dynamic obstacles. Moreover, explicit maps are difficult to represent implicit environmental attributes (such as material reflection characteristics and lighting distribution). In addition, the mixed interference of dynamic objects and static structures makes it impossible for the motion planning module to effectively distinguish passable areas from instantaneous risk areas. Especially in high-frequency dynamic scenarios such as warehousing logistics and human-robot coexistence, existing technologies often cause trajectory jitter or even collisions due to lagging map updates or semantic ambiguities. Summary of the Invention

[0004] This application provides an autonomous navigation method and related devices based on the fusion of laser and vision, which are used to solve the problem that related technologies cannot effectively perform route planning in complex dynamic environments.

[0005] In the first aspect of this application, an autonomous navigation method based on the fusion of laser and vision is provided. The autonomous navigation method based on the fusion of laser and vision includes:

[0006] Obtain the laser point cloud data of the lidar, the binocular image data of the binocular camera, and the IMU clock signal;

[0007] Perform multi-sensor spatio-temporal joint synchronization processing on the laser point cloud data, the binocular image data, and the IMU clock signal to generate a spatio-temporally aligned multi-modal data stream;

[0008] Process the multi-modal data stream through cross-modal feature pyramids and attention guidance to generate a dense environmental representation with semantic labels and a spatio-temporal mask for dynamic objects;

[0009] Determine a multi-layer semantic map corresponding to the dense environmental representation through hierarchical factor graph optimization;

[0010] Fuse the multi-layer semantic map with the neural radiance field to generate a hybrid map of an explicit semantic map and an implicit radiance field;

[0011] Determine the motion trajectory of the target device by performing spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of the dynamic object.

[0012] Optionally, in the first implementation manner of the first aspect of this application, the step of performing multi-sensor spatio-temporal joint synchronization processing on the lidar point cloud data, the binocular image data, and the IMU clock signal to generate a spatio-temporally aligned multi-modal data stream includes:

[0013] Construct a hardware trigger network based on the precise clock protocol;

[0014] Generate a sensor data pool with timestamp-aligned lidar, binocular camera, and IMU clock signal through the hardware trigger network;

[0015] Perform interpolation processing on the unsynchronized lidar point cloud data and binocular image data in the sensor data pool to generate a spatio-temporally continuous lidar point cloud sequence and image frame sequence;

[0016] Determine the external parameter matrix corresponding to the lidar point cloud sequence and the image frame sequence based on deep reinforcement learning;

[0017] Generate a spatio-temporally aligned multi-modal data stream through the spatio-temporal joint synchronization of the sensor data pool and the external parameter matrix.

[0018] Optionally, in the second implementation manner of the first aspect of this application, the step of generating a dense environment representation with semantic labels and a spatio-temporal mask of a dynamic object by performing cross-modal feature pyramid and attention-guided fusion processing on the multi-modal data stream includes:

[0019] Generate an edge feature map and a height feature encoding map by extracting geometric features of the spatio-temporally aligned lidar point cloud data;

[0020] Generate a semantic feature map and an instance mask map by performing semantic segmentation on the spatio-temporally aligned binocular image data;

[0021] Determine a cross-modal feature tensor corresponding to the edge feature map and the semantic feature map through cross-modal attention-guided feature fusion;

[0022] Generate a spatio-temporal mask of the dynamic object by performing dynamic object detection on the cross-modal feature tensor;

[0023] Generate a dense environment representation with semantic labels through multi-scale feature pyramid fusion of the height encoded feature map and the instance mask map.

[0024] Optionally, in the third implementation manner of the first aspect of the present application, the step of determining the multi-layer semantic map corresponding to the dense environment representation by optimizing the hierarchical factor graph includes:

[0025] Jointly model the ICP matching residuals of the laser point cloud, the visual reprojection error, and the IMU pre-integration constraint of the dense environment representation as factor nodes, and generate a corresponding local factor graph;

[0026] The spatio-temporal density distribution of the local factor graph divides the global map corresponding to the dense environment representation into several sub-map units;

[0027] Based on loop detection, perform incremental semantic processing on the several sub-map units, and generate corresponding sub-map semantic maps;

[0028] Fuse geometric and semantic information according to the overlap region confidence of all the sub-map semantic maps, and generate a globally consistent multi-layer semantic map.

[0029] Optionally, in the fourth implementation manner of the first aspect of the present application, the step of fusing the multi-layer semantic map with the neural radiance field to generate a hybrid map of the explicit semantic map and the implicit radiance field includes:

[0030] Map the semantic labels and geometric features at different levels in the multi-layer semantic map to a sparse voxel hash table to generate an explicit semantic voxel field;

[0031] Remove the implicit voxel regions in the neural radiance field that conflict with the multi-layer semantic map according to the geometric boundaries of the explicit semantic voxel field to generate a geometrically constrained radiance field;

[0032] Perform residual learning on the color rendering branch of the geometrically constrained radiance field according to the semantic labels of the explicit semantic voxel field to generate an implicit radiance field that fuses the explicit semantics;

[0033] Perform spatio-temporal mask filtering on the implicit radiance field according to the spatio-temporal mask of the dynamic object to generate a hybrid map of the explicit semantic map and the implicit radiance field.

[0034] Optionally, in the fifth implementation manner of the first aspect of the present application, the step of determining the motion trajectory of the target device by spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of the dynamic object includes:

[0035] Generate a spatio-temporal risk field according to the semantic labels and the radiance field density distribution of the implicit radiance field;

[0036] Determine the boundary of the feasible region of the starting point and the target point of the target device according to the spatio-temporal risk field and the spatio-temporal mask of the dynamic object;

[0037] Generate an initial motion trajectory within the boundary of the feasible region through multi-objective trajectory optimization;

[0038] Based on the spatio-temporal occupancy probability distribution of the spatio-temporal mask of the dynamic object, adjust the timestamps and spatial coordinates of the trajectory points corresponding to the initial motion trajectory in real time;

[0039] According to the timestamps and spatial coordinates adjusted in real time, update the motion trajectory of the target device in real time.

[0040] Optionally, in the sixth implementation manner of the first aspect of the present application, the method further includes:

[0041] Compare the real-time sensor data of the hybrid map with the density distribution of the historical implicit radiation field to determine the environmental structure change area;

[0042] According to the environmental structure change area, perform local semantic map update on the explicit semantic map;

[0043] Perform radiation field reprojection optimization on the updated local semantic map and the implicit radiation field;

[0044] Correct the density distribution of the optimized radiation field through backpropagation of the spatio-temporal mask of the dynamic object;

[0045] Generate an environment-adaptive hybrid map based on the updated explicit semantic map and the implicit radiation field.

[0046] The second aspect of the present application provides an autonomous navigation device based on laser and vision fusion, and the autonomous navigation device based on laser and vision fusion includes:

[0047] An acquisition module, configured to acquire lidar point cloud data, binocular image data of a binocular camera, and an IMU clock signal;

[0048] A first generation module, configured to perform multi-sensor spatio-temporal joint synchronization processing on the lidar point cloud data, the binocular image data, and the IMU clock signal to generate a spatio-temporally aligned multi-modal data stream;

[0049] A second generation module, configured to generate a dense environment representation with semantic labels and a spatio-temporal mask of a dynamic object by processing the multi-modal data stream through cross-modal feature pyramids and attention guidance;

[0050] A first determination module, configured to determine a multi-layer semantic map corresponding to the dense environment representation through hierarchical factor graph optimization;

[0051] A third generation module, configured to fuse the multi-layer semantic map with a neural radiation field to generate a hybrid map of an explicit semantic map and an implicit radiation field;

[0052] A second determination module, configured to determine the motion trajectory of the target device by performing spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of the dynamic object.

[0053] In a third aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor. The processor is configured to execute a computer program stored on the memory. When the processor executes the computer program, the steps in the autonomous navigation method based on laser and vision fusion provided in the first aspect of the embodiments of the present application are implemented.

[0054] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the autonomous navigation method based on laser and vision fusion provided in the first aspect of the embodiments of the present application are implemented.

[0055] In summary, according to an autonomous navigation method and related devices provided by the solution of the present application, lidar point cloud data, binocular image data of a binocular camera, and an IMU clock signal are acquired; the lidar point cloud data, the binocular image data, and the IMU clock signal are subjected to multi-sensor spatio-temporal joint synchronization processing to generate spatio-temporally aligned multi-modal data streams; the multi-modal data streams are processed through a cross-modal feature pyramid and attention guidance to generate a dense environment representation with semantic labels and a spatio-temporal mask of dynamic objects; a multi-layer semantic map corresponding to the dense environment representation is determined through hierarchical factor graph optimization; the multi-layer semantic map is fused with a neural radiance field to generate a hybrid map of an explicit semantic map and an implicit radiance field; the motion trajectory of the target device is determined by performing spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of the dynamic object. Through the implementation of the solution of the present application, in a dynamic and complex environment, through the fusion of the explicit semantic map and the implicit radiance field and the joint constraints of the spatio-temporal mask of the dynamic object, high-precision real-time differentiation between dynamic obstacles and static structures can be achieved, thereby significantly improving the obstacle avoidance robustness of the motion trajectory. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 It is a schematic flowchart of the autonomous navigation method based on laser and vision fusion provided by the embodiments of the present application;

[0057] Figure 2 It is a schematic diagram of program modules of the autonomous navigation device based on laser and vision fusion provided by the embodiments of the present application;

[0058] Figure 3 It is a schematic structural diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0059] In order to make the invention objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0060] To solve the problem that the related technology cannot effectively perform route planning in a complex dynamic environment, the embodiments of the present application provide an autonomous navigation method based on the fusion of laser and vision, as Figure 1 is a schematic flowchart of the autonomous navigation method based on the fusion of laser and vision provided in this embodiment. The autonomous navigation method based on the fusion of laser and vision includes the following steps:

[0061] Step 110: Obtain the laser point cloud data of the lidar, the binocular image data of the binocular camera, and the IMU clock signal.

[0062] Specifically, in this embodiment, obtaining the laser point cloud data of the lidar, the binocular image data of the binocular camera, and the IMU clock signal is to collect multi-modal information of the environment and the device. The lidar generates point cloud data by emitting laser light and receiving reflected light, providing high-precision depth information of the environment; the binocular camera obtains two-dimensional images through the difference in the perspectives of the two cameras, and calculates the disparity between the images to infer three-dimensional space information; the IMU clock signal provides timestamp information to help synchronize the data of different sensors and ensure the spatio-temporal consistency between multi-sensor data.

[0063] Step 120: Perform multi-sensor spatio-temporal joint synchronization processing on the laser point cloud data, the binocular image data, and the IMU clock signal to generate a spatio-temporally aligned multi-modal data stream.

[0064] Specifically, in this embodiment, performing multi-sensor spatio-temporal joint synchronization processing on the laser point cloud data, the binocular image data, and the IMU clock signal aims to ensure the spatio-temporal consistency between the data by accurately matching the timestamps of the data of different sensors. Through this synchronization processing, the data collected by the lidar, vision, and inertial measurement unit sensors can be aligned in time and space, thereby realizing the fusion of data from different sources and ensuring the accuracy and consistency of subsequent processing.

[0065] In an optional real-time mode of this embodiment, the steps of performing multi-sensor spatio-temporal joint synchronization processing on lidar point cloud data, binocular image data, and IMU clock signals to generate a spatio-temporally aligned multi-modal data stream include: constructing a hardware trigger network based on the Precision Clock Protocol; generating a sensor data pool with timestamp alignment for lidar, binocular cameras, and IMU clock signals through the hardware trigger network; performing interpolation processing on the unsynchronized lidar point cloud data and binocular image data in the sensor data pool to generate a spatio-temporally continuous lidar point cloud sequence and image frame sequence; determining an external parameter matrix corresponding to the lidar point cloud sequence and image frame sequence based on deep reinforcement learning; and generating a spatio-temporally aligned multi-modal data stream through the spatio-temporal joint synchronization of the sensor data pool and the external parameter matrix.

[0066] Specifically, in this embodiment, in order to implement the construction of a hardware trigger network based on the Precision Clock Protocol, it is first necessary to introduce a precision clock synchronization protocol and adopt the IEEE 1588 Precision Time Protocol (PTP). The PTP protocol transmits time information between different devices through the master-slave clock mechanism to ensure that the clocks of sensors can be synchronized, achieving a synchronization accuracy of up to microseconds or even sub-microseconds. Specifically, when implementing, install PTP clock sources between sensors, and the sensor devices are connected through a network to synchronize their clocks regularly according to the PTP protocol. Whenever a sensor collects data, a timestamp is attached to record the current system time. Through this precise synchronization, the problem of clock drift between different sensors can be effectively solved, enabling the data of sensors such as lidar, binocular cameras, and IMUs to be processed under the same time reference.

[0067] Based on the above hardware trigger network, through the data acquisition of sensors and the synchronization of the Precision Clock Protocol, a sensor data pool with spatio-temporal alignment can be generated. The data of each sensor will have a timestamp added during acquisition, indicating the exact time of data acquisition. Since the sampling frequencies of lidar, binocular cameras, and IMUs are different, data may be collected from different sensors within the same time window, resulting in data misalignment. In this case, the generated sensor data pool can ensure that each sensor's data has a timestamp for subsequent alignment and fusion.

[0068] However, despite the basis of spatio-temporal alignment, due to the sampling frequencies and data transmission delays of different sensors, there are still cases where some sensor data is unsynchronized, especially the laser point cloud data and image frame data. At this time, interpolation processing needs to be performed on the unsynchronized laser point cloud data and image data. For each frame of unsynchronized laser point cloud data, linear interpolation can be performed using the laser data of adjacent frames before and after, so as to infer the point cloud data of the intermediate frame. Similarly, for binocular image data, through the matching features between images, interpolation is performed using the image data of the front and back frames to generate a spatio-temporally continuous image sequence and laser point cloud sequence.

[0069] The extrinsic matrix corresponding to the laser point cloud sequence and image frame sequence is determined through deep reinforcement learning. The extrinsic matrix includes a rotation matrix and a translation vector, which describe the relative position and attitude between sensors. To solve the problem of dynamic adjustment of the extrinsic parameters, an agent based on DRL is designed. In this agent, the action space is the 6 degrees of freedom (6-DoF) of the extrinsic parameters, including the adjustment amounts of rotation and translation. By defining the reward function as the laser-vision feature matching residual, the agent can continuously optimize the extrinsic matrix through interaction with the environment, maximizing the feature matching degree between the laser point cloud and the image. For example, assume there is a laser point cloud P and an image frame I, and the matching error between them can be expressed as:

[0070] ,

[0071] where, and respectively represent the corresponding feature points on the laser point cloud and the image, represents the three-dimensional point at the corresponding position extracted from the image frame. By optimizing this matching error, the agent gradually learns the optimal extrinsic matrix. The optimization algorithm in reinforcement learning is used to continuously adjust the extrinsic matrix during this process, thereby improving the calibration accuracy.

[0072] After the optimization of the extrinsic matrix, these extrinsic matrices are applied to the sensor data pool to achieve spatio-temporal joint synchronization. Specifically, the laser point cloud data and image data will be coordinate-transformed according to the extrinsic matrix and converted to the same coordinate system, thereby ensuring the spatial alignment of the data. At the same time, combined with the timestamp information in the sensor data pool, the consistency of the data of different sensors in time is ensured. Through this spatio-temporal joint synchronization, a multi-modal data stream is generated, and this multi-modal data stream includes precisely aligned laser point clouds, image frames, and IMU data.

[0073] Step 130: Process the multi-modal data stream through cross-modal feature pyramid and attention-guided fusion to generate a dense environmental representation with semantic labels and a spatio-temporal mask for dynamic objects.

[0074] Specifically, in this embodiment, the multi-modal data stream is processed through cross-modal feature pyramid and attention guidance fusion to extract and integrate the key information in each sensor data. The feature pyramid structure can extract deep features at different scales, while the attention mechanism dynamically adjusts the fusion process according to the importance of the data, highlighting the information in the key regions. The dense environment representation and dynamic object spatio-temporal mask generated by this process can effectively represent the environmental structure and the spatio-temporal positions of dynamic objects, providing an accurate basis for subsequent spatial modeling and path planning.

[0075] In an optional real-time manner in this embodiment, the steps of generating a dense environment representation with semantic labels and a dynamic object spatio-temporal mask through cross-modal feature pyramid and attention guidance fusion processing of the multi-modal data stream include: generating an edge feature map and a height feature encoding map by extracting the geometric features of the spatio-temporally aligned lidar point cloud data; generating a semantic feature map and an instance mask map by semantic segmentation of the spatio-temporally aligned binocular image data; determining a cross-modal feature tensor corresponding to the edge feature map and the semantic feature map through cross-modal attention-guided feature fusion; generating a dynamic object spatio-temporal mask by dynamic object detection of the cross-modal feature tensor; and generating a dense environment representation with semantic labels through multi-scale feature pyramid fusion of the height encoding feature map and the instance mask map.

[0076] Specifically, in this embodiment, to extract the geometric features of the spatio-temporally aligned lidar point cloud data, it is first necessary to extract important geometric features from the point cloud data, mainly including edge and height change information. The point cloud data generated by lidar can reveal the three-dimensional geometric structure of the scene. The edge feature represents the contour of an object or a place with a sharp change, which can help the system identify the significant structures or boundaries in the scene. The height feature encoding map captures the changes in the terrain in the scene by analyzing the height information of the point cloud, such as identifying the relative heights of buildings, roads, or obstacles. In the processing of binocular image data, semantic segmentation technology is first used to decompose the image into different semantic regions, such as roads, buildings, or pedestrians. Semantic segmentation classifies each pixel in the image through a convolutional neural network (CNN), so that each pixel is assigned to a specific category. For example, if the value of a certain pixel in the image is I(x,y) and this pixel belongs to category c, the goal of semantic segmentation is to output a label map L(x,y)=c through the network, where c represents the category of this pixel. In this way, a semantic feature map can be generated, which maps each pixel to a specific semantic label. In addition, on the basis of semantic segmentation, the instance mask map further generates a mask for each object in the image. For example, for multiple objects in the same category, the instance mask map can distinguish the regions of each object and generate specific object masks by calculating the overlap or separation between different objects.

[0077] After that, a cross-modal attention-guided feature fusion method is adopted to fuse the edge feature map extracted from the lidar point cloud with the semantic feature map extracted from the image. The cross-modal attention mechanism uses the self-attention mechanism (Self-Attention) to align the feature maps of the lidar point cloud and the image at each scale, and adjusts the weights of the features by calculating the correlation between different modalities, so that the information in the important regions is enhanced. For example, assuming that the edge feature map and the semantic feature map are E and S respectively, a weighted fusion feature tensor F can be calculated through the attention mechanism, and the formula is:

[0078] ,

[0079] where, and are the learned weight matrices, is the activation function (such as ReLU). Through this fusion, the edge features and semantic features can complement each other, thus enhancing the expression ability of the fusion features. By performing dynamic object detection on the cross-modal feature tensor, dynamic objects in the scene can be identified. Dynamic objects are manifested as objects whose positions change between multiple time frames. By analyzing the spatio-temporal information in the cross-modal feature tensor, dynamic objects can be effectively detected and distinguished from the static background. Temporal Convolutional Networks (TCN) can be used for dynamic object detection to capture temporal information. By processing the features at each time step, a spatio-temporal mask of dynamic objects is generated. For example, the mask M(t) generated by the TCN model represents the dynamic state of an object at a certain moment, where M(t)=1 indicates that the object is dynamic, and M(t)=0 indicates that the object is static.

[0080] Finally, the multi-scale feature pyramid fusion of the height feature encoding map and the instance mask map can generate a dense environmental representation with semantic labels. The Feature Pyramid Networks (FPN) help the system capture the detailed information of the environment at different resolutions by extracting and fusing feature maps at different scales. Assuming that the feature maps extracted from multiple scales s1, s2, …, sn are Fs1, Fs2, …, Fsn respectively, these features are fused through the pyramid structure to generate the final environmental representation :

[0081] ,

[0082] where, is the weight of each scale feature, and the finally generated It contains the global information and local details of the environment, and can provide an accurate environmental representation for subsequent path planning, navigation, and other perception tasks. This representation not only includes the semantic information in the image (such as roads, buildings, etc.), but also combines the geometric information of the laser point cloud (such as the edges and heights of objects), enabling the system to accurately understand and reconstruct the environmental structure. The semantic label refers to assigning a specific meaning or category to each element or region in the environment, which is achieved through technologies such as semantic segmentation, object detection, and classification.

[0083] Step 140: Determine a multi-layer semantic map corresponding to the dense environmental representation through hierarchical factor graph optimization.

[0084] Specifically, in this embodiment, a multi-layer semantic map corresponding to the dense environmental representation is determined through hierarchical factor graph optimization. Using the factor graph model, semantic information at different levels can be integrated and optimized, thereby generating a semantic map with a hierarchical structure. This process combines the constraints of different perception information, improves the accuracy and reliability of the environmental representation, and enables the multi-layer semantic map to better reflect the dynamic changes and complex structure of the environment.

[0085] In an optional real-time manner in this embodiment, the steps of determining a multi-layer semantic map corresponding to the dense environmental representation through hierarchical factor graph optimization include: jointly modeling the ICP matching residuals of the laser point cloud, visual reprojection errors, and IMU pre-integration constraints of the dense environmental representation as factor nodes, and generating a corresponding local factor graph; the spatio-temporal density distribution of the local factor graph divides the global map corresponding to the dense environmental representation into several sub-graph units; based on loop detection, incremental semantic processing is performed on the several sub-graph units, and corresponding sub-graph semantic maps are generated; according to the confidence weighted fusion of geometric and semantic information in the overlapping regions of all sub-graph semantic maps, a globally consistent multi-layer semantic map is generated.

[0086] Specifically, in this embodiment, the ICP (Iterative Closest Point) matching residuals of the laser point cloud estimate the relative position of adjacent laser scans by minimizing the spatial distance between point clouds; the visual reprojection error constrains the camera pose in a SLAM (Simultaneous Localization and Mapping) system through the difference between the projection of feature points of an image on the image plane and the actual observation; the IMU pre-integration constraint estimates the pose change during movement based on the acceleration and angular velocity data provided by the inertial measurement unit through temporal integration. To achieve the effective fusion of these constraints, they need to be modeled as factor nodes in a factor graph to form a unified optimization framework. A factor graph is a graph structure used to represent various constraint conditions, where nodes represent variables and edges represent constraints. In such a factor graph, each factor corresponds to a constraint condition, and the edges connect the variables to be optimized. By minimizing the cost functions of all factors in the factor graph, the optimal variable values can be obtained. In this solution, the ICP matching residuals of the laser point cloud, the visual reprojection error, and the IMU pre-integration constraint will be added as factors to the factor graph, and each factor will have a corresponding cost function, which can be optimized by the least squares method. After constructing the factor graph, the next step is to divide the global map according to the spatio-temporal density distribution. The spatio-temporal density distribution reflects the variation law of sensor data (including laser point cloud, image, and IMU data) in time and space. By analyzing the distribution of these data, the entire environment can be divided into several sub-map units, and each sub-map unit represents the environmental characteristics of a local area. By analyzing the density of different sub-maps, the appropriate sub-map size can be selected according to requirements to ensure that each sub-map can contain sufficient semantic information and geometric structure while avoiding computational redundancy caused by excessive subdivision. The division process of the sub-maps is based on loop detection. Loop detection is a technique used to determine whether one has returned to a previously visited place, which is achieved by comparing the current sensor observations with historical observations. During the construction of sub-maps, loop detection can determine whether they belong to the same environmental area by calculating the overlapping area between two sub-maps. If a loop is detected, incremental semantic processing can be triggered to update the semantic information within the sub-map. The incremental semantic processing method fuses new observations with historical observations and gradually optimizes the semantic expression of each sub-map during the continuous accumulation of data. Each sub-map is optimized through a local factor graph, and finally, a semantic map related to its position and time is generated.

[0087] During the generation of the sub - map semantic map, the integration of geometric information and semantic information is considered. The geometric information comes from the structural data of the laser point cloud, while the semantic information comes from the semantic segmentation results of the images. Therefore, a weighted fusion mechanism needs to be designed to combine the geometric information and semantic information in each sub - map. When fusing, weighting can be performed according to the confidence of the overlapping area between sub - maps to ensure that in the overlapping area, the semantic information and geometric information between multiple sub - maps can be accurately aligned and fused. Suppose the overlapping area of two sub - maps G1 and G2 is , is the confidence of the overlapping area, then their fusion cost function can be expressed as:

[0088] ,

[0089] where, is the cost of geometric information, is the cost of semantic information. The cost of geometric information mainly reflects the matching error of the spatial structure or shape, and the cost of semantic information mainly reflects the matching error of object categories, labels or semantic identifications in images or other sensor data. Through weighted fusion, it can be ensured that the geometric information and semantic information of each sub - map are consistent in the overlapping area, and finally a globally consistent multi - layer semantic map is generated.

[0090] Step 150: Integrate the multi - layer semantic map with the neural radiance field to generate a hybrid map of the explicit semantic map and the implicit radiance field.

[0091] Specifically, in this embodiment, to break through the limitations of traditional explicit maps, the system introduces a neural radiance field (an implicit rendering model) for hybrid modeling. The voxel structure (3D grid cells) of the explicit semantic map serves as the spatial constraint of the neural radiance field, restricting its reconstruction range to improve computational efficiency; the neural radiance field models implicit properties (such as material reflectivity, lighting and shadows) through ray tracing, filling in the missing details of the explicit map. The spatio - temporal mask of dynamic objects filters the radiance field data in the moving areas at this stage, avoiding dynamic interference from affecting the static environment modeling, and finally forming a hybrid map that complements the explicit and implicit parts: the explicit part supports real - time path planning, and the implicit part enhances environmental understanding.

[0092] In an optional implementation manner of this embodiment, the steps of fusing the multi-layer semantic map with the neural radiance field to generate a hybrid map of the explicit semantic map and the implicit radiance field include: mapping semantic labels and geometric features at different levels in the multi-layer semantic map to a sparse voxel hash table to generate an explicit semantic voxel field; removing the implicit voxel regions in the neural radiance field that conflict with the multi-layer semantic map according to the geometric boundaries of the explicit semantic voxel field to generate a geometrically constrained radiance field; performing residual learning on the color rendering branch of the geometrically constrained radiance field according to the semantic labels of the explicit semantic voxel field to generate an implicit radiance field that fuses the explicit semantics; and performing spatio-temporal mask filtering processing on the implicit radiance field according to the spatio-temporal mask of the dynamic object to generate a hybrid map of the explicit semantic map and the implicit radiance field.

[0093] Specifically, in this embodiment, semantic labels and geometric features at different levels in the multi-layer semantic map are mapped to a sparse voxel hash table to generate an explicit semantic voxel field. A voxel represents a small volume unit in space and is used to construct a 3D map. In the sparse voxel hash table, only the regions where there are objects are explicitly stored, while the blank regions do not occupy storage space, which can effectively reduce the computational and storage overheads. The explicit semantic voxel field can assign a unique semantic label and geometric attribute to each voxel in space by combining geometric features and semantic labels from the lidar point cloud and images. This mapping process is implemented through a voxelization algorithm, where the features of each voxel not only reflect its geometric information (such as depth, boundaries) but also contain semantic information (such as roads, buildings, obstacles, etc.).

[0094] The neural radiance field uses an implicit representation method to model the lighting, texture, and structural features in 3D space, and it estimates the color and density at different viewpoints by training a neural network. In this process, the implicit voxels refer to the spatial regions learned through the neural network, and the physical information (such as color and density) they contain is not explicitly stored but is deduced through the radiance field. Since the neural radiance field is not directly restricted by geometric shape constraints, the generated implicit voxels may not be consistent with the semantic labels and geometric features in the multi-layer semantic map. For example, in an urban scene, the neural radiance field may generate some regions that cannot match the roads or buildings. To ensure the generated Figure 1To ensure consistency, the geometric boundaries of the explicit semantic voxel field will be used to contrast and eliminate these conflicting regions, so that the physical information in the radiation field is consistent with the geometric features of the actual environment. On this basis, by learning the semantic labels of the explicit semantic voxel field, the color rendering branch of the geometrically constrained radiation field is further optimized. Residual learning can reduce rendering errors and generate an implicit radiation field that incorporates explicit semantics. By training the neural network, the output of color rendering is made more consistent with the label information in the explicit semantic voxel field, which can improve the accuracy of the final rendering results, especially in object recognition and texture reconstruction in complex scenes.

[0095] Finally, spatio-temporal filtering is performed based on the spatio-temporal mask of dynamic objects to further optimize the representation of the implicit radiation field. This process helps to remove the interference of dynamic objects on the modeling of the static environment, ensuring that the final hybrid map only contains static geometric structures. The spatio-temporal mask of dynamic objects is detected during the multi-modal data fusion process by analyzing lidar point clouds, image frames, and IMU (Inertial Measurement Unit) data to identify moving objects such as pedestrians and vehicles. In this step, the spatio-temporal mask is used to filter out the influence of these dynamic objects, thereby obtaining a more stable and accurate representation of the static environment. The generated hybrid map is a combination of the explicit semantic map and the implicit radiation field, which not only retains the geometric structure and object classification information of the environment but also enables high-quality 3D rendering through the implicit radiation field.

[0096] Step 160: Determine the motion trajectory of the target device through the spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of dynamic objects.

[0097] Specifically, in this embodiment, determining the motion trajectory of the target device through the spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of dynamic objects is to ensure more accurate device positioning and path planning. By constraining the spatio-temporal mask of dynamic objects with the environmental map, the motion trajectory of the target device can be tracked in real time, avoiding the interference of dynamic objects while maintaining a high-precision perception of the environment. Such constraints help to improve the stability and reliability of the autonomous navigation system, ensuring that the device can achieve precise positioning and navigation in complex environments.

[0098] In an optional implementation manner of this embodiment, the steps of determining the motion trajectory of the target device by means of spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of dynamic objects include: generating a spatio-temporal risk field according to the semantic tags and the radiation field density distribution of the implicit radiation field; determining the boundary of the feasible region of the starting point and the target point of the target device according to the spatio-temporal risk field and the spatio-temporal mask of dynamic objects; generating an initial motion trajectory within the boundary of the feasible region through multi-objective trajectory optimization; adjusting the timestamps and spatial coordinates of the trajectory points corresponding to the initial motion trajectory in real time based on the spatio-temporal occupancy probability distribution of the spatio-temporal mask of dynamic objects; and updating the motion trajectory of the target device in real time according to the adjusted timestamps and spatial coordinates.

[0099] Specifically, in this embodiment, when constructing the spatio-temporal risk field, it is first necessary to combine the semantic tags and the radiation field density distribution of the implicit radiation field. This process mainly estimates the risk levels of different regions in the environment by analyzing the geometric features in the environment and the semantic information of objects. By combining the semantic tags with the radiation field density distribution of the implicit radiation field, the spatio-temporal risk distribution of each voxel or spatial region can be obtained. This spatio-temporal risk field reflects the potential risks of different regions in the environment to the target device at a certain moment or during a certain period of time. For example, dynamic objects may make some regions more dangerous, while static obstacles will result in a high density of collision risks in that region. According to the generated spatio-temporal risk field and the spatio-temporal mask of dynamic objects, the boundary of the feasible region of the target device can be further determined. The spatio-temporal mask of dynamic objects is obtained by detecting and calibrating dynamic objects in the environment. These dynamic objects may affect the path planning of the target device. By combining with the spatio-temporal risk field, it can be determined which regions are feasible at a specific moment, that is, regions with no risk or low risk. It can be understood that the starting point and the target point of the target device (including but not limited to autonomous vehicles, drones, mobile robots, AGVs, etc.) also need to be mapped into the spatio-temporal risk field to determine a suitable boundary of the feasible region. This boundary is a region composed of multiple sub-regions, including all the spaces that can be safely passed through. After determining the boundary of the feasible region, an initial motion trajectory is generated next through multi-objective trajectory optimization. Using the spatial information within the boundary of the feasible region and the spatio-temporal mask of dynamic objects, the trajectory planning of the target device is optimized. Multi-objective trajectory optimization considers the balance of multiple objectives, such as the shortest path, the minimum risk path, the obstacle avoidance path, etc., and ensures that the target device can avoid collisions and complete the task within the given time window. Optimization algorithms such as the A* algorithm, the Dijkstra algorithm, or optimization methods based on deep reinforcement learning can be used to generate the preliminary motion trajectory.

[0100] Based on the spatio-temporal occupancy probability distribution of dynamic objects, the initial motion trajectory will be adjusted in real time. During this process, the spatio-temporal occupancy probability distribution represents the probability that a dynamic object occupies a given position at a given time point and spatial location. This probability distribution is generated by continuously monitoring the motion trajectory of the dynamic object and can be updated dynamically. By analyzing the spatio-temporal occupancy probability distribution, the target device can adjust its trajectory in a timely manner to avoid collisions with dynamic objects. When adjusting the trajectory, the time stamp and spatial coordinates of the real-time adjustment will be adjusted according to the position of the dynamic object to ensure that the target device can respond to environmental changes in real time. The real-time updated motion trajectory is based on the continuous adjustment of time stamps and spatial coordinates. The time stamp represents the state of the target device at different time points, while the spatial coordinates represent the position of the device in three-dimensional space. By continuously optimizing the time stamps and spatial coordinates in the trajectory, the target device can ensure that it moves in the optimal way throughout the entire path. At each trajectory update, the device recalculates the possible path according to the current spatio-temporal occupancy probability distribution to ensure that the path is always the safest. This method of updating the motion trajectory based on real-time adjustment not only ensures that the device can avoid collisions with dynamic objects, but also can adapt to complex environmental changes and achieve efficient and safe path planning.

[0101] It should be noted that to ensure that the target device can dynamically respond to environmental changes, the trajectory planning not only considers the current spatial position, but also takes into account the time factor. Therefore, a spatio-temporal graph can be used to represent each time step and spatial position in the path. The trajectory planning of the target device can evaluate each spatio-temporal coordinate point, calculate the collision risk with dynamic objects, and dynamically adjust the path. For example, the trajectory can be optimized by combining the following objective function:

[0102] ,

[0103] where, represents the objective function for trajectory optimization, is the spatial position of the target device at time t , is the spatio-temporal occupancy probability at the corresponding position, is an indicator function indicating whether the position is safe (i.e., no dynamic object occupies the position), is the travel distance of the target device between adjacent time steps, is the safety weight, which determines the importance of whether the path is safe. If a certain position is occupied by a dynamic object at time t , the safety index will penalize the trajectory. The higher the safety, the stronger the ability of the target device to avoid obstacles. It is the path length weight, indicating that the target device hopes to complete the target in the shortest possible time. The larger the weight, the more the device tends to choose a shorter path. It is the collision risk weight, which controls the path selection to prevent the target device from approaching the position of dynamic objects.

[0104] When the target device encounters dynamic objects during real-time operation, the trajectory will be adjusted in real time according to the new environmental state. In this process, the real-time update of timestamps and spatial coordinates is very crucial. For this purpose, spatio-temporal dynamic optimization algorithms are used, such as those based on Kalman Filter or Particle Filter, to update the trajectory of the target device in real time. Through these methods, the state of the target device at the next time step can be estimated, and the best position of the target device can be recalculated based on the latest position and occupancy probability of the dynamic objects. Through the following Kalman filter update formula, the position and velocity estimates of the target device can be updated:

[0105] ,

[0106] where, is the device state (such as position, velocity) at time k, is the state transition matrix, is the control input (such as acceleration), is the process noise, is the control input matrix, whose function is to convert the control input into an impact on the system state. By estimating and real-time updating the motion state of the device, accurate position prediction of the target device in a dynamic environment can be achieved.

[0107] In an alternative implementation of this embodiment, the real-time sensor data of the hybrid map is compared with the density distribution of the historical implicit radiation field to determine the environmental structure change area; the explicit semantic map is updated locally according to the environmental structure change area; the radiation field re-projection optimization is performed on the updated local semantic map and the implicit radiation field; the density distribution of the optimized radiation field is corrected through the backpropagation of the dynamic object spatio-temporal mask; an environment-adaptive hybrid map is generated based on the updated explicit semantic map and the implicit radiation field.

[0108] Specifically, in this embodiment, to determine the changed area of the environmental structure, it is necessary to compare the real-time sensor data with the density distribution of the historical implicit radiation field. The implicit radiation field (such as the Neural Radiance Field - NeRF) models the radiation field of the scene through a multi-layer MLP (Multi-Layer Perceptron) network. These models provide information about illumination, color, and transparency for each spatial point. By comparing the real-time lidar point cloud or image data (provided by sensors such as lidar and cameras) with the density data in the historical implicit radiation field, the areas that have changed compared to the original environment can be identified. After determining the changed area of the environmental structure, local updates need to be made to the explicit semantic map. The explicit semantic map is constructed by extracting semantic information from sensor data (such as images or lidar point clouds), including information such as the category and boundaries of objects. Local semantic map update means only modifying the areas affected by environmental changes, rather than the entire map, which can significantly improve computational efficiency. During the local update process, by comparing real-time data and historical data, objects that have changed or disappeared can be corrected, such as adding new obstacles or deleting objects that no longer exist. The relationship between the updated local semantic map and the implicit radiation field also needs to be further optimized, especially the reprojection optimization of the radiation field. Since the explicit semantic map and the implicit radiation field are generated by different modeling methods, there may be geometric or semantic inconsistencies between them. Through reprojection optimization, the spatial information of the explicit semantic map can be more precisely aligned with the voxel positions in the implicit radiation field, and the parameters of the implicit radiation field can be adjusted to better match the updated structure and object positions in the environment. After the radiation field reprojection optimization, it is necessary to further correct the density distribution of the optimized radiation field through the backpropagation of the dynamic object spatio-temporal mask. The spatio-temporal mask is generated based on the detection results of dynamic objects and is used to mark objects in the scene that do not belong to the static environment, such as moving vehicles or walking people. Through backpropagation, the influence of these dynamic objects is eliminated, thus avoiding misleading the static environment model (such as the implicit radiation field). The core idea of backpropagation is to adjust the parameters in the model by comparing the error between the predicted radiation field and the real data, so that the radiation field better matches the results of the dynamic object spatio-temporal mask, thereby reducing the influence of dynamic objects on the environmental model. Finally, based on the updated explicit semantic map and the implicit radiation field, an adaptive hybrid map can be generated. This hybrid map not only retains the clear structural information provided by the explicit semantic labels but also integrates the high-precision details of the implicit radiation field. The adaptive hybrid map can automatically update with the changes in the environment and make adjustments according to the dynamic changes in the environment. By comprehensively applying the advantages of the explicit semantic map and the implicit radiation field, this hybrid map can provide a more comprehensive and accurate environmental representation in complex scenes.

[0109] It should be noted that when performing reprojection optimization, the calculation formula can be expressed as:

[0110] ,

[0111] where \(E\) is the total optimization error, represents a spatial point, and represent the projection positions in the map and the data respectively, and represent the density values respectively, is the weight for balancing the geometric error and the density error. Through this formula, the error can be minimized by adjusting the model parameters, thereby optimizing the alignment between the implicit radiation field and the explicit semantic map.

[0112] According to an autonomous navigation method based on laser and vision fusion provided by the solution of the present application, the laser point cloud data of the lidar, the binocular image data of the binocular camera, and the IMU clock signal are acquired; the laser point cloud data, the binocular image data, and the IMU clock signal are subjected to multi-sensor spatio-temporal joint synchronization processing to generate a spatio-temporally aligned multi-modal data stream; the multi-modal data stream is processed by cross-modal feature pyramid and attention-guided fusion to generate a dense environmental representation with semantic labels and a spatio-temporal mask of dynamic objects; the multi-layer semantic map corresponding to the dense environmental representation is determined by hierarchical factor graph optimization; the multi-layer semantic map is fused with the neural radiation field to generate a hybrid map of the explicit semantic map and the implicit radiation field; the motion trajectory of the target device is determined by spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of dynamic objects. By implementing the solution of the present application, in a dynamic and complex environment, through the fusion of the explicit semantic map and the implicit radiation field and the joint constraints of the spatio-temporal mask of dynamic objects, high-precision real-time discrimination between dynamic obstacles and static structures can be achieved, thereby significantly improving the obstacle avoidance robustness of the motion trajectory.

[0113] Figure 2 An autonomous navigation device based on laser and vision fusion provided by an embodiment of the present application, which can be used to implement the autonomous navigation method based on laser and vision fusion in the foregoing embodiment. As Figure 2 shown, the autonomous navigation device based on laser and vision fusion mainly includes:

[0114] An acquisition module 10, configured to acquire the laser point cloud data of the lidar, the binocular image data of the binocular camera, and the IMU clock signal;

[0115] A first generation module 20, configured to perform multi-sensor spatio-temporal joint synchronization processing on the laser point cloud data, the binocular image data, and the IMU clock signal to generate a spatio-temporally aligned multi-modal data stream;

[0116] The second generation module 30 is used to generate a dense environment representation with semantic labels and a spatio-temporal mask of dynamic objects through cross-modal feature pyramid and attention-guided fusion processing of the multi-modal data stream;

[0117] The first determination module 40 is used to determine a multi-layer semantic map corresponding to the dense environment representation through hierarchical factor graph optimization;

[0118] The third generation module 50 is used to fuse the multi-layer semantic map with the neural radiance field to generate a hybrid map of an explicit semantic map and an implicit radiance field;

[0119] The second determination module 60 is used to determine the motion trajectory of the target device through spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of dynamic objects.

[0120] In an optional real-time manner of this embodiment, the first generation module is specifically used to: construct a hardware trigger network based on the Precision Time Protocol; generate a sensor data pool with timestamp alignment of lidar, binocular camera, and IMU clock signals through the hardware trigger network; perform interpolation processing on the unsynchronized lidar point cloud data and binocular image data in the sensor data pool to generate a spatio-temporally continuous lidar point cloud sequence and image frame sequence; determine an external parameter matrix corresponding to the lidar point cloud sequence and the image frame sequence based on deep reinforcement learning; generate a spatio-temporally aligned multi-modal data stream through spatio-temporal joint synchronization of the sensor data pool and the external parameter matrix.

[0121] In an optional real-time manner of this embodiment, the second generation module is specifically used to: generate an edge feature map and a height feature encoding map by extracting geometric features of the spatio-temporally aligned lidar point cloud data; generate a semantic feature map and an instance mask map by semantic segmentation of the spatio-temporally aligned binocular image data; determine a cross-modal feature tensor corresponding to the edge feature map and the semantic feature map through cross-modal attention-guided feature fusion; generate a spatio-temporal mask of dynamic objects by detecting dynamic objects in the cross-modal feature tensor; generate a dense environment representation with semantic labels through multi-scale feature pyramid fusion of the height encoding feature map and the instance mask map.

[0122] In an optional real-time manner of this embodiment, the first determination module is specifically used to: jointly model the lidar point cloud ICP matching residual, visual reprojection error, and IMU pre-integration constraint of the dense environment representation as factor nodes, and generate a corresponding local factor graph; the spatio-temporal density distribution of the local factor graph divides the global map corresponding to the dense environment representation into several sub-graph units; perform incremental semantic processing on the several sub-graph units based on loop detection, and generate corresponding sub-graph semantic maps; generate a globally consistent multi-layer semantic map by fusing geometric and semantic information with confidence weighting of the overlapping regions of all sub-graph semantic maps.

[0123] In an optional real-time manner of this embodiment, the third generation module is specifically configured to: map semantic labels and geometric features at different levels in the multi-layer semantic map to a sparse voxel hash table to generate an explicit semantic voxel field; eliminate the implicit voxel regions in the neural radiance field that conflict with the multi-layer semantic map according to the geometric boundary of the explicit semantic voxel field to generate a geometrically constrained radiance field; perform residual learning on the color rendering branch of the geometrically constrained radiance field according to the semantic labels of the explicit semantic voxel field to generate an implicit radiance field that fuses explicit semantics; perform spatio-temporal mask filtering on the implicit radiance field according to the spatio-temporal mask of the dynamic object to generate a hybrid map of the explicit semantic map and the implicit radiance field.

[0124] In an optional real-time manner of this embodiment, the second determination module is specifically configured to: generate a spatio-temporal risk field according to the semantic label and the radiance field density distribution of the implicit radiance field; determine the boundary of the feasible region of the starting point and the target point of the target device according to the spatio-temporal risk field and the spatio-temporal mask of the dynamic object; generate an initial motion trajectory within the boundary of the feasible region through multi-objective trajectory optimization; adjust the timestamps and spatial coordinates of the trajectory points corresponding to the initial motion trajectory in real time based on the spatio-temporal occupancy probability distribution of the spatio-temporal mask of the dynamic object; update the motion trajectory of the target device in real time according to the timestamps and spatial coordinates adjusted in real time.

[0125] In an optional real-time manner of this embodiment, the third generation module is further configured to: compare the real-time sensor data of the hybrid map with the density distribution of the historical implicit radiance field to determine the environmental structure change region; update the local semantic map of the explicit semantic map according to the environmental structure change region; perform radiance field reprojection optimization on the updated local semantic map and the implicit radiance field; correct the density distribution of the optimized radiance field through backpropagation of the spatio-temporal mask of the dynamic object; generate an environment-adaptive hybrid map based on the updated explicit semantic map and the implicit radiance field.

[0126] An autonomous navigation device based on the fusion of laser and vision provided by the solution of the present application acquires the laser point cloud data of a lidar, the binocular image data of a binocular camera, and the IMU clock signal; performs multi-sensor spatio-temporal joint synchronization processing on the laser point cloud data, the binocular image data, and the IMU clock signal to generate a spatio-temporally aligned multi-modal data stream; processes the multi-modal data stream through a cross-modal feature pyramid and attention guidance to generate a dense environmental representation with semantic labels and a spatio-temporal mask of dynamic objects; determines a multi-layer semantic map corresponding to the dense environmental representation through hierarchical factor graph optimization; fuses the multi-layer semantic map with a neural radiance field to generate a hybrid map of an explicit semantic map and an implicit radiance field; determines the motion trajectory of the target device through spatio-temporal joint constraints on the hybrid map and the spatio-temporal mask of dynamic objects. Through the implementation of the solution of the present application, in a dynamic and complex environment, through the fusion of the explicit semantic map and the implicit radiance field and the joint constraint of the spatio-temporal mask of dynamic objects, high-precision real-time differentiation between dynamic obstacles and static structures can be achieved, thereby significantly improving the obstacle avoidance robustness of the motion trajectory.

[0127] Provided by the solution of the present application Figure 3 This is an electronic device provided by an embodiment of the present application. This electronic device can be used to implement the autonomous navigation method based on the fusion of laser and vision in the foregoing embodiments, and mainly includes:

[0128] A memory 301, a processor 302, and a computer program 303 stored on the memory 301 and executable on the processor 302. The memory 301 and the processor 302 are communicatively connected. When the processor 302 executes the computer program 303, the autonomous navigation method based on the fusion of laser and vision in the foregoing embodiments is implemented. Among them, the number of processors can be one or more.

[0129] The memory 301 can be a high-speed random access memory (RAM, Random Access Memory) or a non-volatile memory, such as a disk memory. The memory 301 is used to store executable program code, and the processor 302 is coupled to the memory 301.

[0130] Furthermore, an embodiment of the present application also provides a computer-readable storage medium, which can be disposed in the electronic device in the foregoing embodiments. The computer-readable storage medium can be the memory in the foregoing Figure 3 illustrated embodiments.

[0131] A computer program is stored on the computer-readable storage medium. When the program is executed by a processor, it implements the autonomous navigation method based on laser and vision fusion in the foregoing embodiments. Further, the computer-readable storage medium may also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a RAM, a magnetic disk, or an optical disc.

[0132] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0134] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application.

Claims

1. An autonomous navigation method based on laser and vision fusion, characterized in that: include: Obtain laser point cloud data from the laser radar, binocular image data from the binocular camera, and IMU clock signal; Performing multi-sensor spatiotemporal joint synchronization processing on the laser point cloud data, the binocular image data, and the IMU clock signal to generate a spatiotemporal aligned multimodal data stream; Processing the multimodal data stream by cross-modal feature pyramid and attention-guided fusion to generate dense environment representation with semantic labels and spatiotemporal masks of dynamic objects; Determining a multi-layer semantic map corresponding to the dense environment representation through hierarchical factor graph optimization; fusing the multi-layer semantic map with the neural radiation field to generate a hybrid map of the explicit semantic map and the implicit radiation field; The motion trajectory of the target device is determined by the spatiotemporal joint constraints on the hybrid map and the spatiotemporal mask of the dynamic object.

2. The autonomous navigation method based on laser and vision fusion according to claim 1 is characterized in that: The step of performing multi-sensor spatiotemporal joint synchronous processing on the laser point cloud data, the binocular image data and the IMU clock signal to generate a spatiotemporal aligned multimodal data stream includes: Build a hardware trigger network based on the precise clock protocol; Generate a sensor data pool with timestamp alignment of the laser radar, the binocular camera and the IMU clock signal through the hardware trigger network; Performing interpolation processing on the unsynchronized laser point cloud data and the binocular image data in the sensor data pool to generate a spatiotemporally continuous laser point cloud sequence and image frame sequence; Determine an extrinsic parameter matrix corresponding to the laser point cloud sequence and the image frame sequence based on deep reinforcement learning; A multimodal data stream aligned in time and space is generated through the spatiotemporal joint synchronization of the sensor data pool and the extrinsic parameter matrix.

3. The autonomous navigation method based on laser and vision fusion according to claim 2 is characterized in that: The step of processing the multimodal data stream by cross-modal feature pyramid and attention-guided fusion to generate dense environment representation with semantic labels and spatiotemporal masks of dynamic objects includes: Generate edge feature map and height feature encoding map by extracting geometric features of spatiotemporally aligned laser point cloud data; Generate semantic feature maps and instance mask maps through semantic segmentation of spatiotemporally aligned binocular image data; Determining a cross-modal feature tensor corresponding to the edge feature map and the semantic feature map through feature fusion guided by cross-modal attention; Generating a dynamic object spatiotemporal mask by detecting dynamic objects from the cross-modal feature tensor; A dense environment representation with semantic labels is generated according to the multi-scale feature pyramid fusion of the height feature encoding map and the instance mask map.

4. The autonomous navigation method based on laser and vision fusion according to claim 1 is characterized in that: The step of determining a multi-layer semantic map corresponding to the dense environment representation by hierarchical factor graph optimization comprises: The laser point cloud ICP matching residual, visual reprojection error and IMU pre-integration constraint represented by the dense environment are jointly modeled as factor nodes, and a corresponding local factor graph is generated; The spatiotemporal density distribution of the local factor graph divides the global map corresponding to the dense environment representation into a plurality of sub-graph units; Performing incremental semantic processing on the plurality of sub-graph units based on loop closure detection, and generating corresponding sub-graph semantic maps; The geometric and semantic information are weightedly fused according to the confidence of the overlapping areas of all the sub-graph semantic maps to generate a globally consistent multi-layer semantic map.

5. The autonomous navigation method based on laser and vision fusion according to claim 4 is characterized in that: The step of fusing the multi-layer semantic map with the neural radiation field to generate a hybrid map of the explicit semantic map and the implicit radiation field comprises: Mapping semantic labels and geometric features at different levels in the multi-layer semantic map into a sparse voxel hash table to generate an explicit semantic voxel field; Eliminating implicit voxel regions in the neural radiation field that conflict with the multi-layer semantic map according to the geometric boundary of the explicit semantic voxel field to generate a geometrically constrained radiation field; Performing residual learning on the color rendering branch of the geometrically constrained radiation field according to the semantic label of the explicit semantic voxel field to generate an implicit radiation field that integrates explicit semantics; The implicit radiation field is subjected to spatiotemporal mask filtering processing according to the spatiotemporal mask of the dynamic object to generate a hybrid map of the explicit semantic map and the implicit radiation field.

6. The autonomous navigation method based on laser and vision fusion according to claim 5 is characterized in that: The step of determining the motion trajectory of the target device by applying spatiotemporal joint constraints to the hybrid map and the spatiotemporal mask of the dynamic object comprises: generating a spatiotemporal risk field according to the semantic label and the radiation field density distribution of the implicit radiation field; Determine the feasible area boundary of the starting point and the target point of the target device according to the spatiotemporal risk field and the spatiotemporal mask of the dynamic object; Generate an initial motion trajectory within the boundary of the feasible region through multi-objective trajectory optimization; Adjusting the timestamps and spatial coordinates of the trajectory points corresponding to the initial motion trajectory in real time based on the spatiotemporal occupancy probability distribution of the spatiotemporal mask of the dynamic object; The motion trajectory of the target device is updated in real time according to the timestamp and the spatial coordinates adjusted in real time.

7. The autonomous navigation method based on laser and vision fusion according to claim 5 is characterized in that: The method further comprises: Comparing the real-time sensor data of the hybrid map with the density distribution of the historical implicit radiation field to determine the area of ​​environmental structure change; Performing a local semantic map update on the explicit semantic map according to the environmental structure change area; Perform radiation field reprojection optimization on the updated local semantic map and implicit radiation field; Correcting the density distribution of the optimized radiation field by back-propagation of the spatiotemporal mask of the dynamic object; An environment-adaptive hybrid map is generated based on the updated explicit semantic map and implicit radiation field.

8. An autonomous navigation device based on laser and vision fusion, characterized in that: The autonomous navigation device based on laser and vision fusion includes: An acquisition module is used to acquire laser point cloud data from the laser radar, binocular image data from the binocular camera, and IMU clock signal; A first generating module is used to perform multi-sensor spatiotemporal joint synchronous processing on the laser point cloud data, the binocular image data and the IMU clock signal to generate a spatiotemporal aligned multimodal data stream; A second generation module is used to process the multimodal data stream through cross-modal feature pyramid and attention-guided fusion to generate dense environment representation with semantic labels and spatiotemporal masks of dynamic objects; A first determination module, configured to determine a multi-layer semantic map corresponding to the dense environment representation through hierarchical factor graph optimization; A third generation module is used to fuse the multi-layer semantic map with the neural radiation field to generate a hybrid map of the explicit semantic map and the implicit radiation field; The second determination module is used to determine the motion trajectory of the target device by applying spatiotemporal joint constraints to the hybrid map and the spatiotemporal mask of the dynamic object.

9. An electronic device, characterized in that: The device comprises a memory and a processor, wherein: The processor is used to execute the computer program stored in the memory; When the processor executes the computer program, the steps of the autonomous navigation method based on laser and vision fusion described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the autonomous navigation method based on laser and vision fusion described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-source information hierarchical fusion method and device and storage medium

    CN114429432A

  • Three-dimensional point cloud semantic map construction method based on neural radiation field

    CN118379450A