Laser vision fusion navigation positioning method, device, equipment and storage medium
By aligning and fusion processing of point cloud data and image data, target recognition combined with a preset neural network model, and generating navigation and positioning maps, the problem of insufficient navigation and positioning accuracy in complex scenarios in the existing technology is solved, and the accuracy and reliability of the navigation system are improved.
Patent Information
- Application Number
- CN202510629041.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing autonomous driving and intelligent navigation technologies are difficult to meet the needs of high-precision navigation and positioning in complex scenarios. A single sensor solution has problems such as low data alignment efficiency, insufficient feature fusion, and insufficient model generalization capabilities, resulting in large errors in navigation map construction and poor system robustness.
By acquiring point cloud data and image data, data alignment is performed and fusion conversion is performed to generate a fusion map. The target recognition of the fusion map is used to identify the target, obtain the passable area, obstacle category and location information, and the multiple target recognition results are fused through formula (1) to generate a navigation positioning map.
The accuracy and reliability of navigation positioning in complex scenarios are improved, and the problems of large errors in navigation map construction and poor system robustness are solved. The generated navigation positioning map can more accurately support the path planning and obstacle avoidance of the navigation system.
Smart Images

Figure CN120141510A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of positioning and navigation, and particularly relates to a laser-vision fusion navigation and positioning method, device, equipment, and storage medium. Background Art
[0002] With the rapid development of autonomous driving and intelligent navigation technologies, high-precision environmental perception has become the core requirement for reliable navigation and positioning. Current autonomous driving and intelligent navigation technologies are all implemented through pure laser solutions or pure vision solutions. However, single-sensor solutions generally have defects. For example, although lidar can provide accurate three-dimensional coordinates, it lacks semantic information. Although vision sensors have rich texture and semantic recognition capabilities, they are easily affected by environmental factors such as lighting and occlusion, resulting in incomplete obstacle recognition and fuzzy division of passable areas, making it difficult to meet the navigation and positioning accuracy requirements in complex scenarios. Existing fusion solutions have tried to combine multi-modal data, but there are problems such as low data alignment efficiency, insufficient feature fusion, and insufficient model generalization ability, resulting in large errors in navigation map construction and poor system robustness.
[0003] Therefore, how to improve the accuracy and reliability of navigation and positioning in complex scenarios is an urgent problem to be solved. Summary of the Invention
[0004] In view of this, a laser-vision fusion navigation and positioning method, device, equipment, and storage medium provided by the embodiments of the present application can improve the accuracy and reliability of navigation and positioning in complex scenarios. The laser-vision fusion navigation and positioning method, device, equipment, and storage medium provided by the embodiments of the present application are implemented as follows: Obtain the point cloud data and image data of the current area, and perform data alignment on the point cloud data and the image data; Perform fusion conversion processing on the aligned point cloud data and image data to obtain a fusion map; According to the fusion map and a preset neural network model, obtain multiple target recognition results, where the multiple target recognition results include passable areas, obstacle categories, and position information. The preset neural network model is obtained by training an initial neural network model with sample data collected in historical navigation and positioning scenarios. The sample data includes multiple historical fusion maps collected in the historical navigation and positioning scenarios, as well as multiple passable areas, multiple historical obstacle categories, and multiple position information in the historical navigation and positioning scenarios; Perform fusion processing on the multiple target recognition results according to formula (1) to obtain a navigation and positioning map, where the navigation and positioning map is used for navigation and positioning; (1) Wherein, is the navigation and positioning map, is the quantity of the target recognition results, is the weight coefficient of the i-th target recognition result, is the i-th target recognition result, and Ci is the confidence of the i-th target recognition result.
[0005] In some embodiments, obtaining multiple target recognition results according to the fused map and a preset neural network model includes: Performing normalization processing on the fused map to obtain a normalized fused map; Performing feature extraction on the normalized fused map to obtain a multi-scale feature map; Performing multi-task branch processing on the multi-scale feature map to obtain a passable area, obstacle categories, and position information.
[0006] In some embodiments, the multi-task branch includes a passable area segmentation branch, an obstacle category recognition branch, and a position information extraction branch. Performing multi-task branch processing on the multi-scale feature map to obtain a passable area, obstacle categories, and position information includes: Performing upsampling processing on the multi-scale feature map through the passable area segmentation branch, and inputting the upsampled multi-scale feature map into a preset probability prediction function for calculation to obtain a passable area; Performing mapping processing on the multi-scale feature map through the obstacle category recognition branch to obtain the possibility scores of each obstacle category in the multi-scale feature map, and determining the obstacle category with the highest possibility score as the current obstacle category; Performing information extraction processing on the multi-scale feature map through the position information extraction branch to obtain position information, where the position information includes at least one of spatial coordinates, relative position relationships, and positioning confidence.
[0007] In some embodiments, obtaining the point cloud data and image data of the current area and performing data alignment on the point cloud data and the image data includes: Obtaining the timestamp of the point cloud data of the current area, the timestamp of the image data, and the external parameter matrix between the point cloud data and the image data, where the external parameter matrix includes rotation parameters and translation vectors; Aligning the timestamp of the point cloud data and the timestamp of the image data according to the time interpolation method; Performing rotation alignment on the point cloud data and the image data according to the rotation parameters, and performing translation alignment on the point cloud data and the image data according to the translation vectors.
[0008] In some embodiments, the fusion map includes a dynamic area and a static area. After performing fusion conversion processing on the aligned point cloud data and the image data to obtain the fusion map, the method further includes: Segment the fusion map according to a preset area segmentation method to obtain a segmented fusion map, and the segmented fusion map includes multiple different areas; Perform semantic recognition on the segmented fusion map according to a visual semantic algorithm to obtain the semantic categories of multiple different areas in the fusion map, and the semantic categories include dynamic objects and static objects; Calculate the motion vectors of the point cloud data in multiple different areas of the segmented fusion map, and mark the areas with semantic categories of dynamic objects and motion vectors greater than a preset motion threshold as dynamic areas; Perform rejection processing on the point cloud data and image data corresponding to the dynamic areas to obtain a processed fusion map.
[0009] In some embodiments, the method for performing fusion processing on the multiple target recognition results to obtain a navigation and positioning map includes: Calculate the multiple result confidence levels of different target recognition results in the multiple target recognition results according to the evidence theory algorithm; Perform probability assignment calculation on the multiple result confidence levels of different target recognition results to obtain the recognition accuracy rates of different target recognition results; Determine the target category of the target recognition result corresponding to the current recognition accuracy rate according to the recognition accuracy rate, and generate a navigation and positioning map according to the target category.
[0010] In some embodiments, before performing normalization processing on the fusion map to obtain a normalized fusion map, the method further includes: Perform bilateral filtering denoising processing on the fusion map to obtain a processed fusion map.
[0011] A laser-vision fusion navigation and positioning device provided by an embodiment of the present application includes: An acquisition module, configured to acquire point cloud data and image data of a current area, and perform data alignment on the point cloud data and the image data; A processing module, configured to perform fusion conversion processing on the aligned point cloud data and the image data to obtain a fusion map; An identification module, configured to obtain multiple target identification results according to the fused map and a preset neural network model. The multiple target identification results include passable areas, obstacle categories, and location information. The preset neural network model is obtained by training an initial neural network model with sample data collected in a historical navigation and positioning scenario. The sample data includes multiple historical fused maps collected in the historical navigation and positioning scenario, multiple passable areas, multiple historical obstacle categories, and multiple location information in the historical navigation and positioning scenario; A fusion module, configured to perform fusion processing on the multiple target identification results according to formula (1) to obtain a navigation and positioning map for navigation and positioning; (1) Where, is the navigation and positioning map, is the number of the target identification results, is the weight coefficient of the i-th target identification result, is the i-th target identification result, and Ci is the confidence level of the i-th target identification result.
[0012] The computer device provided by the embodiment of the present application includes a memory and a processor. The memory stores a computer program that can run on the processor, and the processor implements the method described in the embodiment of the present application when executing the program.
[0013] The computer-readable storage medium provided by the embodiment of the present application stores a computer program thereon, and the computer program implements the method provided by the embodiment of the present application when executed by a processor.
[0014] A laser vision fusion navigation and positioning method, device, equipment, and storage medium provided by the embodiment of the present application include: obtaining point cloud data and image data of the current area, and performing calibration through data alignment; performing fusion conversion processing on the aligned data to generate a fused map including geometric structure and semantic information; based on the fused map and a preset neural network model, identifying passable areas, obstacle categories, and location information from the fused map based on the preset neural network model; fusing the above multi-target identification results to generate a navigation and positioning map to provide an environmental model support for the navigation system. In this way, navigation and positioning are performed through the obtained navigation and positioning map, improving the accuracy and reliability of navigation and positioning in complex scenarios and solving the technical problems proposed in the background technology. Description of the Drawings
[0015] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a schematic diagram of the implementation process of a laser vision fusion navigation and positioning method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the implementation process of another laser vision fusion navigation and positioning method provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the implementation process of another laser vision fusion navigation and positioning method provided by an embodiment of the present application; Figure 4 It is a laser vision fusion navigation and positioning device provided by an embodiment of the present application. Detailed implementation manners
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0019] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments and can be combined with each other without conflict.
[0020] It should be noted that the terms "first / second / third" involved in the embodiments of the present application are used to distinguish similar or different objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0021] Figure 1 It is a schematic diagram of the implementation process of a laser vision fusion navigation and positioning method provided by an embodiment of the present application. As Figure 1 shown, the method may include the following steps 101 to 104: Step 101: Obtain the point cloud data and image data of the current area, and align the point cloud data and image data.
[0022] In the embodiment of the present application, a lidar device can be used to collect the point cloud data of the current area. The lidar emits laser beams at a certain frequency, and calculates the distance between the object and the sensor by measuring the time when the laser beams are reflected back, so as to obtain the three-dimensional coordinate information of the object surface and form the point cloud data.
[0023] At the same time, a vision sensor such as a camera is used to collect the image data of the current area. The image data contains the texture, color and semantic information of the scene. Since there may be a time difference in the data collection of the lidar and the camera, it is necessary to synchronize the data of the two. The data of the two can be ensured to be collected at the same moment by means of hardware triggering; or through software algorithms, the data is interpolated or matched according to the timestamps to align the point cloud data and the image data in time.
[0024] Since the installation positions and angles of the lidar and the camera are different, spatial calibration is also required. Tools such as calibration boards can be used to determine the relative position and attitude relationship between the two through feature matching and coordinate transformation, and project the point cloud data and the image data into the same coordinate system.
[0025] Step 102: Perform fusion conversion processing on the aligned point cloud data and image data to obtain a fusion map.
[0026] In the embodiment of the present application, geometric features are extracted from the point cloud data, such as the density, curvature, surface normal, etc. of the point cloud. Algorithms based on machine learning, such as principal component analysis, can be used to calculate the surface normal and curvature information of the point cloud.
[0027] Semantic features are extracted from the image data, such as the category, edge, texture, etc. of the object. Deep learning models, such as convolutional neural networks, can be used to perform semantic segmentation and object detection on the image to extract useful feature information.
[0028] The extracted point cloud geometric features and image semantic features are fused. The feature splicing method can be adopted to splice the two features in the feature dimension; or the feature weighted fusion method can be adopted, different weights are assigned according to the importance of the features, and then fusion is performed.
[0029] The fused features are processed to generate a fusion map. The fusion map can be in the form of a grid map, an octree map, etc., where each grid or node contains the geometric information of the point cloud and the semantic information of the image.
[0030] Step 103: Obtain multiple target recognition results according to the fusion map and the preset neural network model.
[0031] In an embodiment of the present application, the current fused map is input into a trained preset neural network model, and the model outputs multiple target recognition results through forward propagation calculation.
[0032] For the passable area, the result output by the model can be a pixel-level classification probability map, and the probability map is converted into a binary passable area mask through threshold processing.
[0033] For the obstacle category, the model outputs the probability that each area belongs to different obstacle categories, and selects the category with the highest probability as the obstacle category of that area.
[0034] For the position information, the model can directly output the three-dimensional coordinates or pixel coordinates of the obstacle or the passable area.
[0035] The process of training the preset neural network model includes collecting sample data in historical navigation and positioning scenarios, including multiple historical fused maps and the corresponding annotation data of the passable area, obstacle category, and position information. An initial neural network model is constructed, and the model can adopt structures such as convolutional neural networks and recurrent neural networks.
[0036] The initial neural network model is trained using the sample data, and the parameters of the model are continuously adjusted using the backpropagation algorithm and an optimizer (such as stochastic gradient descent) to minimize the error between the output result of the model and the annotation data.
[0037] Step 104, perform fusion processing on multiple target recognition results to obtain a navigation and positioning map, which is used for navigation and positioning.
[0038] In an embodiment of the present application, the recognition results of the passable area, obstacle category, and position information are fused. Methods such as voting method and weighted average method can be used to comprehensively consider the confidence levels of different results to obtain the final fusion result.
[0039] Post-process the fusion result, such as removing noise, filling holes, smoothing boundaries, etc., to improve the accuracy and reliability of the result.
[0040] According to the result after fusion processing, a navigation and positioning map is generated. Specifically, the fusion processing is performed according to formula (1) to obtain the navigation and positioning map.
[0041] (1) where is the navigation and positioning map, which is a two-dimensional or three-dimensional map data structure that combines multiple target recognition results. Its specific form can be a grid map, a feature map, etc., and is used for path planning, obstacle avoidance, etc. in the navigation and positioning system. is the number of the target recognition results. Different target recognition results may come from the outputs of the neural network model for different features or the target recognition results at different times. is the weight coefficient of the i-th target recognition result, which reflects the importance of this target recognition result in the final fusion. is the i-th target recognition result, which can be a vector or a matrix. The specific form depends on the content of the target recognition. Ci is the confidence of the i-th target recognition result, and its value range is a numerical value between 0 and 1, which is used to measure the reliability of this target recognition result.
[0042] The navigation and positioning map can be a two-dimensional map or a three-dimensional map, which contains information such as passable areas, positions and categories of obstacles.
[0043] Apply the navigation and positioning map to the navigation and positioning system to provide path planning and obstacle avoidance guidance for robots, autonomous driving vehicles, etc.
[0044] In the embodiments of the present application, by fusing the point cloud data of the lidar and the image data of the camera, the advantages of both are fully utilized, and the accuracy and reliability of environmental perception are improved. Using the preset neural network model for target recognition can quickly and accurately obtain the passable area, obstacle category and position information. Through the fusion processing of multiple target recognition results, the generated navigation and positioning map is more accurate, providing strong support for navigation and positioning.
[0045] Based on the Figure 1 schematic diagram of the implementation process of a laser-vision fusion navigation and positioning method shown above, the present application also provides a schematic diagram of the implementation process of a laser-vision fusion navigation and positioning method, as shown in Figure 2 shown. This method may include the following steps 201 to step 203: Step 201, perform normalization processing on the fusion map to obtain the normalized fusion map.
[0046] In the embodiments of the present application, the collected fusion map contains the point cloud data of the lidar and the image data of the camera. First, perform linear normalization processing on the fusion map to normalize the pixel values to the [0, 1] interval. For example, for the height information in the lidar point cloud data, find its minimum and maximum values and process them according to the linear normalization formula.
[0047] Step 202, perform feature extraction on the normalized fusion map to obtain a multi-scale feature map.
[0048] In the embodiments of the present application, a convolutional neural network is used to extract features from the normalized fusion map. Feature maps of different scales are output in different convolutional layers. For example, high-resolution feature maps are obtained in shallow convolutional layers to capture detailed information; low-resolution but semantically rich feature maps are obtained in deep convolutional layers. At the same time, a feature pyramid network is constructed to fuse feature maps of different scales to obtain multi-scale feature maps.
[0049] Step 203: Perform multi-task branching processing on the multi-scale feature map to obtain the passable area, obstacle category, and position information.
[0050] In the embodiments of the present application, the multi-scale feature map is upsampled through a transposed convolutional layer to make its size the same as the original fusion map. Then, a convolutional layer and a Sigmoid function are used to output the probability map of each pixel belonging to the passable area. The threshold is set to 0.5. Pixels with a probability greater than 0.5 are marked as the passable area, and pixels less than 0.5 are marked as the non-passable area to obtain a binary passable area mask.
[0051] The multi-scale feature map is input into a fully connected layer, and the probability of each area belonging to different obstacle categories is output through a Softmax activation function. For example, obstacle categories such as vehicles and columns are recognized, and the category with the highest probability is selected as the obstacle category of the area.
[0052] A convolutional layer and a regression layer are used to process the multi-scale feature map to output the three-dimensional coordinate information of the obstacle. For example, through the recognition of a vehicle, the X, Y, and Z axis coordinates of the vehicle in the parking lot coordinate system are output.
[0053] In the embodiments of the present application, normalization processing is used to eliminate the scale differences of different sensor data in the fusion map, unify the data distribution, improve the stability and convergence efficiency of neural network training, and enhance the adaptability of the model to diverse scenarios; multi-scale feature maps are extracted by means of a convolutional neural network, etc., taking into account fine-grained geometric details and high-level semantic information, solving the problem of incomplete recognition of single-scale features in complex environments, and improving the comprehensiveness of target recognition; multi-task branching processing enables each target recognition task to focus on optimizing exclusive features, realizing feature complementarity while reducing computational redundancy, accurately dividing the passable area, reducing the misclassification rate of obstacles, improving the accuracy of position information, enhancing the inference efficiency and robustness of the model, and adapting to complex dynamic scenarios.
[0054] In the above Figure 2 Based on the schematic implementation flow diagram of a laser-vision fusion navigation and positioning method shown, the present application also provides a schematic implementation flow diagram of a laser-vision fusion navigation and positioning method, as Figure 3 shown. This method may include the following steps 301 to step 303: Step 301: Upsample the multi-scale feature map through the passable area segmentation branch, and input the upsampled multi-scale feature map into a preset probability prediction function for calculation to obtain the passable area.
[0055] In the embodiment of the present application, a deconvolution layer (such as transposed convolution) or an interpolation algorithm (such as bilinear interpolation) is used to upsample the multi-scale feature map, gradually restoring the resolution of the feature map to 1 / 2 of the original fused map or the same size. For example, for a deep feature map with a size of H / 4×W / 4×C, it is upsampled to H×W×C / 4 through 2 times of deconvolution (convolution kernel size 2×2, stride 2), where H is the height of the feature map, W is the width of the feature map, and C is the number of channels of the feature map, retaining spatial details.
[0056] Input the upsampled feature map into a preset probability prediction function (such as Sigmoid function or Softmax function), and output the probability value of each pixel belonging to the passable area. 。
[0057] Set a probability threshold T (such as 0.5), and mark the pixels >T as the passable area to generate a binary mask. 。
[0058] Step 302: Map the multi-scale feature map through the obstacle category recognition branch to obtain the likelihood scores of each obstacle category in the multi-scale feature map, and determine the obstacle category with the highest likelihood score as the current obstacle category.
[0059] In the embodiment of the present application, the multi-scale feature map is mapped to the category dimension through global average pooling or a fully connected layer to obtain a likelihood score vector for each obstacle category. where N is the preset number of obstacle categories (such as pedestrians, vehicles, traffic cones, etc.).
[0060] Use the Softmax function to normalize the score vector so that where zi is the original score output by the branch network.
[0061] Select the category with the highest likelihood score as the current obstacle category, and at the same time retain the score vector for localization confidence evaluation.
[0062] Step 303: Extract information from the multi-scale feature map through the position information extraction branch to obtain position information, where the position information includes at least one of spatial coordinates, relative position relationships, and localization confidence.
[0063] In the embodiments of the present application, the three-dimensional spatial coordinates (x, y, z) or two-dimensional pixel coordinates (u, v) of the obstacle are directly output through a regression network. For the point cloud fusion feature, the PointNet++ network is used to extract the local point cloud coordinates, and the three-dimensional coordinates in the global coordinate system are regressed through a fully connected layer; for the visual feature, the two-dimensional coordinates of the image detection frame are converted into three-dimensional world coordinates through the external camera parameters.
[0064] Taking the sensor coordinate system as the origin, the distance, azimuth angle, and height of the obstacle relative to the sensor are calculated to form a relative position vector.
[0065] The uncertainty of the position estimate is represented by the predicted variance or probability distribution of the output coordinates, which is used for the subsequent navigation decision-making module to adjust the path planning weight.
[0066] In the embodiments of the present application, the exclusive upsampling branch is used to retain the spatial details, combined with pixel-level probability prediction, to avoid the boundary blurring problem caused by excessive downsampling in traditional single-task models, making the edge division between the passable area and the obstacle clearer. At the same time, the spatial coordinates, relative position, and positioning confidence are output, providing hierarchical positioning information for the navigation system. The accurate coordinates are used for global path planning, the relative position supports real-time obstacle avoidance, and the positioning confidence assists the decision-making system to dynamically adjust the control strategy, overall improving the positioning reliability and navigation safety.
[0067] In some embodiments, the point cloud data and image data of the current area are acquired, and the data alignment of the point cloud data and image data is performed, including: acquiring the timestamp of the point cloud data of the current area, the timestamp of the image data, and the external parameter matrix between the point cloud data and the image data, where the external parameter matrix includes rotation parameters and translation vectors.
[0068] In the embodiments of the present application, when the lidar and the camera collect data, their respective timestamp information is synchronously recorded to identify the acquisition time of the data. For example, the lidar collects point cloud data at a frequency of 10Hz and generates corresponding timestamps, and the camera collects image data at a frequency of 30Hz and records the synchronous timestamps.
[0069] The relative position relationship between the lidar and the camera is determined through a joint calibration method to obtain the external parameter matrix. The external parameter matrix is used to describe the relative position relationship between the lidar coordinate system and the camera coordinate system, and is specifically represented as a homogeneous transformation matrix: , where is a 3X3 rotation matrix, which is composed of the rotation parameters around the three axes of the camera coordinate system and is used to describe the spatial rotation of the lidar coordinate system relative to the camera coordinate system. Specifically, it is represented as: , where is the rotation angle (pitch angle) of the lidar coordinate system around the x-axis of the camera coordinate system, and the unit is radians. is the rotation angle around the y-axis (roll angle), in radians, is the rotation angle around the z-axis (yaw angle), in radians, are the basic rotation matrices around the three axes respectively, following the right-hand screw rule, is a 3×1 translation vector, representing the three-dimensional coordinate offset of the origin of the lidar coordinate system in the camera coordinate system. Specifically, the meanings are as follows: is the offset distance of the lidar in the horizontal direction (x-axis) of the camera coordinate system, in meters; is the offset distance in the vertical direction (y-axis), in meters; is the offset distance in the depth direction (z-axis), in meters.
[0070] The extrinsic parameter matrix includes rotation parameters and a translation vector. The rotation parameters are used to describe the rotation angles of the lidar coordinate system relative to the camera coordinate system (such as rotations around the x-axis, y-axis, and z-axis); the translation vector is used to describe the three-dimensional coordinate offset of the origin of the lidar coordinate system in the camera coordinate system (such as distances in the horizontal, vertical, and depth directions).
[0071] Specifically, multiple groups of point cloud and image data from different perspectives can be collected through a calibration tool (such as a checkerboard calibration board combined with special software), and an accurate extrinsic parameter matrix can be calculated.
[0072] Specifically, the timestamps of the point cloud data and the image data are aligned according to the time interpolation method.
[0073] In the embodiments of this application, when the timestamps of the lidar point cloud and the image are inconsistent, the time interpolation method is used for synchronization. For example, for the timestamp of a certain image data, two adjacent timestamps are found in the point cloud data, and these two timestamps are respectively earlier and later than the current image timestamp.
[0074] According to the time difference between these two adjacent point cloud data and the corresponding point cloud data, the interpolated point cloud data corresponding to the current image timestamp is calculated by means of linear interpolation or curve interpolation, so that the point cloud and the image are matched in time.
[0075] To ensure data validity, a time synchronization threshold (such as 50 milliseconds) is set, and only the point cloud and image pairs with a time difference within the threshold are processed subsequently, and the invalid data pairs with too large time deviations are excluded.
[0076] Specifically, the point cloud data and the image data are rotationally aligned according to the rotation parameters, and the point cloud data and the image data are translationally aligned according to the translation vector.
[0077] In the embodiments of the present application, the spatial angle of the lidar point cloud data is adjusted by using the rotation parameters in the extrinsic matrix. Specifically, the three-dimensional coordinates of each lidar point are spatially transformed according to the rotation parameters to eliminate the direction deviation caused by the difference in the sensor installation angle, so that the spatial orientation of the point cloud data is consistent with the image data.
[0078] Based on the translation vector in the extrinsic matrix, position offset compensation is performed on the rotated point cloud data. The origin of the lidar point cloud coordinate system is adjusted to the origin position of the camera coordinate system, so that the two types of data are in the same reference coordinate system in space, ensuring that the three-dimensional coordinates of the point cloud can accurately correspond to the two-dimensional pixel points of the image.
[0079] In the embodiments of the present application, the timestamp error between the point cloud and the image is controlled within an extremely small range by the time interpolation method, effectively solving the problem of object position misalignment caused by different sensor sampling frequencies in dynamic scenes, ensuring the time consistency of the data. The accurate calibration of the extrinsic matrix and the rotation and translation transformation significantly reduce the position error of the lidar point cloud projected onto the image plane, ensuring the accurate correspondence between the three-dimensional coordinates of the point cloud and the semantic information of the image, providing a high-precision spatial reference for the subsequent generation of the fused map.
[0080] In some embodiments, the fused map includes a dynamic area and a static area. After performing fusion conversion processing on the aligned point cloud data and image data to obtain the fused map, it further includes: dividing the fused map according to a preset area segmentation method to obtain the segmented fused map, and the segmented fused map includes multiple different areas.
[0081] In some embodiments, the fused map generated from the aligned point cloud data and image data is divided into multiple sub-areas according to the spatial division rule. Specifically, any of the following methods can be used: Rasterization segmentation: Divide the fused map into uniform grids (such as grids of 0.5 m × 0.5 m), and each grid contains the point cloud data (such as density, height) and image semantic features (such as color, texture) in the area.
[0082] Superpixel segmentation: Based on the color and texture similarity of the image, divide the continuous areas in the fused map into superpixel blocks, and each superpixel block corresponds to a semantically coherent area in the environment (such as road surface, wall, vehicle surface).
[0083] Segmentation output: Obtain multiple sub-areas containing geometric information (point cloud coordinates) and semantic features (image pixel values), providing basic units for subsequent processing.
[0084] Specifically, semantic recognition is performed on the segmented fused map according to the visual semantic algorithm to obtain the semantic categories of multiple different areas in the fused map, and the semantic categories include dynamic objects and static objects.
[0085] In some embodiments, a deep learning model or a traditional image algorithm is used to perform semantic analysis on each segmented sub-region to determine its category.
[0086] The categories of dynamic objects mainly include object categories with the ability of autonomous movement such as pedestrians, vehicles, cyclists, etc.; The categories of static objects mainly include fixed environmental elements such as buildings, roads, trees, traffic signs, etc.
[0087] Semantic labels (such as "vehicle", "pedestrian", "road surface", "wall") are assigned to each sub-region to establish the correspondence between semantic categories and regions.
[0088] Specifically, the motion vectors of the point cloud data in multiple different regions in the segmented fused map are calculated, and the regions with the semantic category of dynamic objects and the motion vectors greater than the preset motion threshold are marked as dynamic regions.
[0089] In some embodiments, the motion information of each region is obtained by any of the following methods: Point cloud time series analysis: By comparing the changes in the point cloud coordinates of the same region in two adjacent frames of the fused map, the average displacement is calculated as the motion vector (such as the displacement in the x, y, and z directions); Visual optical flow method: Based on the image sequence, the optical flow field of the sub-region is calculated to obtain the pixel-level motion direction and speed, and then converted into the motion vector of the whole region.
[0090] A preset motion threshold is set (such as 0.1 m / frame). For the regions with the semantic category of dynamic objects (such as "vehicle", "pedestrian") and the motion vectors greater than the threshold, they are marked as dynamic regions. For example, if a region is recognized as a "vehicle" and the displacement exceeds 0.1 m / frame, it is determined as a region of a dynamically moving vehicle.
[0091] Specifically, the point cloud data and image data corresponding to the dynamic regions are removed to obtain the processed fused map.
[0092] In some embodiments, the point cloud data (such as vehicle point cloud) and image data (such as vehicle pixel region) marked as dynamic regions are removed from the fused map, and only the region data with the semantic category of static objects (such as "road surface", "wall") or dynamic objects but the motion vectors not exceeding the threshold (such as a stationary vehicle) are retained.
[0093] The processed fused map is generated, which only contains static environmental structures and stationary object information, providing stable input data for subsequent target recognition and navigation planning.
[0094] In the embodiments of the present application, through double screening of semantic recognition and motion analysis, the areas corresponding to dynamic objects are accurately removed, avoiding interference with the static map, making the fused map more conform to the fixed structure of the actual environment, and improving the reliability of long-term use of the map. The combination of region segmentation and semantic recognition converts global dynamic detection into local analysis of sub-regions, reducing the computational complexity; the motion vector calculation is based on the temporal correlation between point clouds and images, avoiding the high time-consuming problem of full-scene matching of traditional three-dimensional point clouds.
[0095] In some embodiments, fusion processing is performed on multiple target recognition results to obtain a navigation and positioning map, including: calculating multiple target recognition results according to the evidence theory algorithm to obtain multiple result confidence levels of different target recognition results among the multiple target recognition results.
[0096] In the embodiments of the present application, the evidence theory algorithm can comprehensively consider the reliability and uncertainty of different target recognition results. The multiple target recognition results are used as evidence and input into the algorithm.
[0097] Through the relevant calculation steps of the evidence theory algorithm, the characteristics and credibility of each target recognition result are analyzed to obtain multiple result confidence levels of different target recognition results. These result confidence levels reflect the reliable degree of each recognition result. For example, for a target result identified by a lidar and a camera respectively, the confidence level of the lidar recognition result and the confidence level of the camera recognition result can be evaluated through the evidence theory algorithm.
[0098] Specifically, probability assignment calculation is performed on multiple result confidence levels of different target recognition results to obtain the recognition accuracy rates of different target recognition results.
[0099] In the embodiments of the present application, according to the obtained multiple result confidence levels, probability assignment is performed on different target recognition results. It can more reasonably measure the contribution of each recognition result to the final decision.
[0100] Through the results of probability assignment, the recognition accuracy rates of different target recognition results are calculated. The recognition accuracy rate can intuitively reflect the accurate degree of each target recognition result. For example, if the recognition accuracy rate of a certain target recognition result is high, it indicates that this result is more credible.
[0101] Specifically, according to the recognition accuracy rate, the target category of the target recognition result corresponding to the current recognition accuracy rate is determined, and a navigation and positioning map is generated according to the target category.
[0102] In the embodiments of the present application, according to the calculated recognition accuracy rate, the target category of the target recognition result corresponding to the current recognition accuracy rate is determined. The target category with the highest recognition accuracy rate is selected as the final target category. For example, if the recognition accuracy rate of the lidar identifying the target as "car" is higher than the recognition accuracy rate of the camera identifying the target as "truck", then the final target category is determined to be "car".
[0103] Based on the determined target category, combined with other relevant information (such as the position, size, etc. of the target), a navigation and positioning map is generated. The navigation and positioning map will accurately reflect the category and distribution of targets in the environment, providing an accurate navigation basis for the navigation system.
[0104] In the embodiments of the present application, the evidence theory algorithm is used to comprehensively process multiple target recognition results, fully considering the uncertainty of different recognition results, and can calculate the recognition accuracy rate more accurately, thereby improving the accuracy of target recognition and reducing the possibility of misjudgment. The navigation and positioning map generated based on the accurate target recognition results can more truly reflect the environmental information, improving the reliability and usability of the map. This helps the navigation system make more reasonable and accurate navigation decisions.
[0105] In some embodiments, before obtaining the normalized fused map by normalizing the fused map, it further includes: performing bilateral filtering denoising processing on the fused map to obtain the processed fused map.
[0106] In the embodiments of the present application, the fused map contains the three-dimensional coordinate data of the laser point cloud and the two-dimensional image data of the camera. The possible types of noise include: Gaussian noise: manifested as random offsets of the point cloud coordinates or random perturbations of the image pixel values (such as lidar measurement errors, camera sensor thermal noise); Salt-and-pepper noise: manifested as isolated outliers (such as distant mis-measured points in the point cloud) or black and white noise points in the image; Mixed noise: the superposition of two or more types of noise in a complex environment, resulting in local data anomalies in the fused map.
[0107] According to the resolution and noise intensity of the fused map, set the spatial radius (such as 3×3, 5×5 pixels) and pixel value difference threshold (such as the gray value difference not exceeding 20) of the bilateral filtering.
[0108] For each pixel in the fused map, calculate the spatial distance weight and pixel value similarity weight of all pixels in its neighborhood, and use the weighted average as the new value of this pixel.
[0109] If the fused map contains multi-channel data (such as the X, Y, Z coordinates of the point cloud and the RGB color values of the image), bilateral filtering is performed on each channel separately to ensure that the noise in each dimension of the data is effectively suppressed.
[0110] In the embodiments of the present application, bilateral filtering can specifically remove noise while retaining the important structural features in the fused map, avoiding the edge blurring problem caused by traditional filtering methods. When the denoised fused map is normalized, the data distribution is more stable.
[0111] Although the present application provides method operation steps such as in the embodiments or flowcharts, based on routine or non-creative labor, there can be more or fewer operation steps. The step order listed in this embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in the order shown in this embodiment or the drawings or executed in parallel (for example, in an environment with parallel processors or multi-threaded processing).
[0112] As Figure 4 shown, the embodiments of the present application also provide a laser vision fusion navigation and positioning device 400. The device includes: An acquisition module 401, configured to acquire the point cloud data and image data of the current area, and perform data alignment on the point cloud data and image data; A processing module 402, configured to perform fusion conversion processing on the aligned point cloud data and image data to obtain a fused map; An identification module 403, configured to obtain multiple target identification results according to the fused map and a preset neural network model. The multiple target identification results include passable areas, obstacle categories, and position information. The preset neural network model is obtained by training an initial neural network model according to the sample data collected in the historical navigation and positioning scenarios. The sample data includes multiple historical fused maps collected in the historical navigation and positioning scenarios, multiple passable areas, multiple historical obstacle categories, and multiple position information in the historical navigation and positioning scenarios; A fusion module 404, configured to perform fusion processing on the multiple target identification results according to formula (1) to obtain a navigation and positioning map, where the navigation and positioning map is used for navigation and positioning; (1) Wherein, is the navigation and positioning map, is the number of target identification results, is the weight coefficient of the i-th target identification result, is the i-th target identification result, and Ci is the confidence level of the i-th target identification result.
[0113] In some embodiments, the processing module 402 is further configured to normalize the fused map to obtain a normalized fused map; The processing module 402 is further configured to extract features from the normalized fused map to obtain a multi-scale feature map; The processing module 402 is further configured to perform multi-task branch processing on the multi-scale feature map to obtain a passable area, an obstacle category, and position information.
[0114] In some embodiments, the processing module 402 is further configured to perform upsampling processing on the multi-scale feature map through a passable area segmentation branch, input the upsampled multi-scale feature map into a preset probability prediction function for calculation, and obtain a passable area; The processing module 402 is further configured to perform mapping processing on the multi-scale feature map through an obstacle category recognition branch to obtain the possibility scores of each obstacle category in the multi-scale feature map, and determine the obstacle category with the highest possibility score as the current obstacle category; The processing module 402 is further configured to perform information extraction processing on the multi-scale feature map through a position information extraction branch to obtain position information, where the position information includes at least one of spatial coordinates, relative position relationships, and positioning confidence.
[0115] In some embodiments, the acquisition module 401 is further configured to acquire the timestamp of the point cloud data, the timestamp of the image data, and the external parameter matrix between the point cloud data and the image data in the current area, where the external parameter matrix includes rotation parameters and translation vectors; The processing module 402 is further configured to align the timestamp of the point cloud data and the timestamp of the image data according to the time interpolation method; The processing module 402 is further configured to perform rotational alignment on the point cloud data and the image data according to the rotation parameters, and perform translational alignment on the point cloud data and the image data according to the translation vectors.
[0116] In some embodiments, the processing module 402 is further configured to segment the fused map according to a preset area segmentation method to obtain a segmented fused map, where the segmented fused map includes multiple different areas; The processing module 402 is further configured to perform semantic recognition on the segmented fused map according to a visual semantic algorithm to obtain the semantic categories of multiple different areas in the fused map, where the semantic categories include dynamic objects and static objects; The processing module 402 is further configured to calculate the motion vectors of the point cloud data in multiple different areas in the segmented fused map, and mark the areas with semantic categories of dynamic objects and motion vectors greater than a preset motion threshold as dynamic areas; The processing module 402 is further configured to perform rejection processing on the point cloud data and the image data corresponding to the dynamic area to obtain a processed fusion map.
[0117] In some embodiments, the recognition module 403 is further configured to calculate multiple target recognition results according to the evidence theory algorithm to obtain multiple result confidence levels of different target recognition results among the multiple target recognition results; The recognition module 403 is further configured to perform probability assignment calculation on the multiple result confidence levels of different target recognition results to obtain the recognition accuracy rate of different target recognition results; The recognition module 403 is further configured to determine the target category of the target recognition result corresponding to the current recognition accuracy rate according to the recognition accuracy rate, and generate a navigation and positioning map according to the target category.
[0118] In some embodiments, the processing module 402 is further configured to perform bilateral filtering denoising processing on the fusion map to obtain a processed fusion map.
[0119] Some modules in the device described in this application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0120] The device or module illustrated in the above application embodiments can be specifically implemented by a computer chip or an entity, or by a product with a certain function. For the convenience of description, the above device is described by dividing it into various modules according to functions. When implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, the module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.
[0121] The methods, devices or modules described in this application can be implemented in the form of computer-readable program code. The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor, and a computer-readable medium that stores computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0122] The embodiments of this application also provide a device, which includes: a processor; a memory for storing instructions executable by the processor; when the processor executes the executable instructions, the method described in the embodiments of this application is implemented.
[0123] The embodiments of this application also provide a non-volatile computer-readable storage medium, on which a computer program or instructions are stored. When the computer program or instructions are executed, the method described in the embodiments of this application is implemented.
[0124] In addition, in each embodiment of the present invention, the various functional modules can be integrated into one processing module, or each module can exist alone, or two or more modules can be integrated into one module.
[0125] The above storage medium includes, but is not limited to, random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD), or memory card. The memory can be used to store computer program instructions.
[0126] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or can also be embodied in the implementation process of data migration. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0127] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. All or part of the present application can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0128] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present application.
Claims
1. A laser vision fusion navigation and positioning method, characterized in that: The method comprises: Acquire point cloud data and image data of the current area, and perform data alignment on the point cloud data and the image data; Performing fusion conversion processing on the aligned point cloud data and the image data to obtain a fused map; According to the fused map and the preset neural network model, a plurality of target recognition results are obtained, the plurality of target recognition results including passable areas, obstacle categories and location information, the preset neural network model is obtained by training an initial neural network model according to sample data collected in a historical navigation and positioning scenario, the sample data including a plurality of historical fused maps collected in the historical navigation and positioning scenario and a plurality of passable areas, a plurality of historical obstacle categories and a plurality of location information in the historical navigation and positioning scenario; The multiple target recognition results are fused according to formula (1) to obtain a navigation positioning map, wherein the navigation positioning map is used for navigation positioning; (1) in, To locate the map for navigation, The number of identification results for the target, is the weight coefficient of the i-th target recognition result, is the i-th target recognition result, and Ci is the confidence of the i-th target recognition result.
2. The method according to claim 1, characterized in that According to the fusion map and the preset neural network model, a plurality of target recognition results are obtained, including: Normalizing the fused map to obtain a normalized fused map; Performing feature extraction on the normalized fusion map to obtain a multi-scale feature map; The multi-scale feature map is processed by multi-task branches to obtain traversable areas, obstacle categories and location information.
3. The method according to claim 2, characterized in that The multi-task branch includes a traversable area segmentation branch, an obstacle category identification branch, and a location information extraction branch; The multi-task branch processing is performed on the multi-scale feature map to obtain the passable area, obstacle category and location information, including: The multi-scale feature map is upsampled by the passable area segmentation branch, and the upsampled multi-scale feature map is input into a preset probability prediction function for calculation to obtain a passable area; Mapping the multi-scale feature map through an obstacle category recognition branch to obtain a possibility score for each obstacle category in the multi-scale feature map, and determining the obstacle category with the highest possibility score as the current obstacle category; The multi-scale feature map is subjected to information extraction processing through a position information extraction branch to obtain position information, wherein the position information includes at least one of spatial coordinates, relative position relationship, and positioning reliability.
4. The method according to claim 1, characterized in that: The acquiring point cloud data and image data of the current area and aligning the point cloud data and the image data comprises: Obtaining a timestamp of the point cloud data of the current region, a timestamp of the image data, and an extrinsic parameter matrix between the point cloud data and the image data, wherein the extrinsic parameter matrix includes a rotation parameter and a translation vector; Aligning the timestamp of the point cloud data and the timestamp of the image data according to a time interpolation method; The point cloud data and the image data are rotationally aligned according to the rotation parameter, and the point cloud data and the image data are translationally aligned according to the translation vector.
5. The method according to claim 1, characterized in that The fused map includes a dynamic area and a static area; After the aligned point cloud data and the image data are fused and converted to obtain a fused map, the method further includes: Segmenting the fused map according to a preset region segmentation method to obtain a segmented fused map, wherein the segmented fused map includes a plurality of different regions; Performing semantic recognition on the segmented fused map according to a visual semantic algorithm to obtain semantic categories of a plurality of different regions in the fused map, wherein the semantic categories include dynamic objects and static objects; Calculating motion vectors of point cloud data in a plurality of different regions in the segmented fusion map, and marking regions whose semantic categories are dynamic objects and whose motion vectors are greater than a preset motion threshold as dynamic regions; The point cloud data and image data corresponding to the dynamic area are eliminated to obtain a processed fused map.
6. The method according to claim 1, characterized in that The fusing the multiple target recognition results to obtain a navigation positioning map includes: Calculating the multiple target recognition results according to the evidence theory algorithm to obtain multiple result confidences of different target recognition results in the multiple target recognition results; Probability distribution calculation is performed on multiple confidence levels of different target recognition results to obtain recognition accuracy rates of different target recognition results; The target category of the target recognition result corresponding to the current recognition accuracy is determined according to the recognition accuracy, and a navigation positioning map is generated according to the target category.
7. The method according to claim 1, characterized in that Before normalizing the fused map to obtain the normalized fused map, the method further includes: The fused map is subjected to bilateral filtering and denoising processing to obtain a processed fused map.
8. A laser vision fusion navigation and positioning device, characterized in that: include: An acquisition module, used to acquire point cloud data and image data of the current area, and perform data alignment on the point cloud data and the image data; A processing module, used for performing fusion conversion processing on the aligned point cloud data and the image data to obtain a fused map; an identification module, for obtaining a plurality of target identification results according to the fused map and a preset neural network model, wherein the plurality of target identification results include passable areas, obstacle categories, and location information, wherein the preset neural network model is obtained by training an initial neural network model according to sample data collected in a historical navigation and positioning scenario, wherein the sample data includes a plurality of historical fused maps collected in the historical navigation and positioning scenario, and a plurality of passable areas, a plurality of historical obstacle categories, and a plurality of location information in the historical navigation and positioning scenario; A fusion module, used for fusing the multiple target recognition results according to formula (1) to obtain a navigation positioning map, wherein the navigation positioning map is used for navigation positioning; (1) in, To locate the map for navigation, The number of identification results for the target, is the weight coefficient of the i-th target recognition result, is the i-th target recognition result, and Ci is the confidence of the i-th target recognition result.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Laser and visual positioning fusion method and equipment
CN111735446A
Map generation method and device, electronic equipment and storage medium
CN114494618A
Point cloud image fusion method and system for vehicle navigation
CN117911829A
Robot positioning method and system based on binocular vision and laser scanning fusion
CN118050734A
Obstacle detection method and device, storage medium and electronic device
CN119625685A
Cited By
Patrol robot navigation positioning system based on multi-modal information fusion
CN120445227A
Navigation method and system for green cutting robot
CN121089710A
A mowing robot navigation method and system
CN121089710B