Positioning Method and System for Complementary Localization Fusion of LiDAR and Stereo Vision

Through the fusion method of complementary positioning of lidar and stereoscopic visual, visual semantic segmentation and lidar point cloud motion vectorization analysis, the problem of positioning deviation in dynamic environment is solved, and the positioning effect with high accuracy and high robustness is achieved.

CN119758363BActive Publication Date: 2025-07-11CUITONG ELECTRONIC TECH (SHANGHAI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510271315.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-08
Publication Date
2025-07-11
Estimated Expiration
2045-03-08

AI Technical Summary

Technical Problem

In the dynamic environment of existing positioning technology, lidar and stereoscopic visual positioning methods have problems with noise point cloud interference and mismatch of visual feature points, resulting in deviations in positioning results. The existing multi-sensor fusion method has failed to effectively eliminate dynamic interference, affecting positioning robustness and accuracy.

Method used

Through the fusion method of complementary positioning of lidar and stereoscopic visual, visual semantic segmentation and lidar point cloud motion vectorization analysis are used to identify dynamic objects and eliminate interfering point clouds, and spatial calibration and IMU motion prediction verification are used to ensure the accuracy of positioning references.

Benefits of technology

High-precision and high-rootability positioning in dynamic environments are achieved, dynamic interference is eliminated, and positioning reference is updated based on static background data, which enhances the adaptability and accuracy of the positioning system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119758363B_ABST
    Figure CN119758363B_ABST
Patent Text Reader

Abstract

The present invention discloses a positioning method and system for complementary positioning fusion of lidar and stereo vision, which relates to the technical field of radio wave positioning. The specific steps of the method are as follows: S100 performs timestamp alignment and spatial calibration on image and point cloud data, S200 identifies dynamic object regions, S300 calculates displacement amounts to identify dynamic interference regions, S400 determines the final true dynamic interference regions, and S500 updates the positioning reference based on the geometric and texture features of static point clouds. Through cross-modal collaboration of visual semantic segmentation and lidar motion vectorization analysis, the present invention realizes precise detection and real-time elimination of dynamic interference. Visual semantic segmentation provides pixel-level dynamic object recognition, while lidar point cloud cluster displacement analysis can quantify the motion trajectory. Through a spatio-temporal alignment mechanism, the two effectively distinguish true dynamic targets from sensor noise, avoiding positioning errors caused by errors in single-sensor data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radio wave positioning, and specifically to a positioning method and system for the complementary positioning fusion of lidar and stereo vision. Background Art

[0002] With the rapid development of technology, intelligent positioning technology has become the core support for many cutting-edge fields. In the field of intelligent logistics, automated guided vehicles need precise positioning to achieve efficient handling and sorting of goods. The accuracy of positioning directly affects the efficiency and cost of logistics operations. In the field of security monitoring, patrol robots rely on accurate positioning to patrol along the preset route and timely detect potential safety hazards. In virtual reality and augmented reality applications, precise positioning can provide users with a more immersive experience, enabling virtual elements to perfectly blend with the real environment. However, the interference of moving objects and the spatio-temporal asynchrony of sensor data in a dynamic and complex environment are important factors leading to positioning drift and map distortion in existing fusion solutions.

[0003] Currently, when existing positioning technologies face a dynamic environment, when there are a large number of dynamic objects in the environment, the point cloud data obtained by lidar will contain a large number of noise point clouds generated by dynamic objects. These noise point clouds will interfere with the extraction of static environmental features by the positioning algorithm, resulting in a large deviation in the positioning result. Traditional stereo vision positioning, although it can obtain rich visual information, in a dynamic scene, due to factors such as light changes and fast object movement, feature points in the visual image are prone to false matching or loss. Moreover, existing multi-sensor fusion positioning methods often simply superimpose the data of lidar and stereo vision, without fully considering the motion characteristics of objects in a dynamic environment, and cannot effectively eliminate dynamic interference, resulting in a significant reduction in the performance of the positioning system in a complex dynamic environment.

[0004] In summary, the limitations of current intelligent positioning technology in a dynamic environment seriously restrict its wide application in actual scenarios. Therefore, there is an urgent need for a new positioning method that can combine visual semantic segmentation and lidar point cloud analysis to achieve real-time elimination of dynamic interference and update the positioning reference, thereby significantly improving the positioning robustness and accuracy in a dynamic scene. Summary of the Invention

[0005] The purpose of the present invention is to make up for the deficiencies of the existing technology, and provide a positioning method and system for the complementary positioning fusion of lidar and stereo vision. It can accurately identify dynamic objects through visual semantic segmentation and lidar point cloud motion vectorization analysis, and real-time eliminate interfering point clouds, thereby providing pure static background data for positioning and effectively reducing positioning deviation caused by dynamic interference.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: On the one hand, a positioning method for the complementary positioning and fusion of lidar and stereo vision. The specific steps of this method are as follows:

[0007] S100. Synchronously acquire lidar point cloud data and stereo vision image data, and perform timestamp alignment and spatial calibration on both;

[0008] S200. Perform real-time semantic segmentation on the stereo vision image, identify the dynamic object area and generate a pixel-level dynamic mask;

[0009] S300. Perform motion vectorization analysis on consecutive frames of lidar point cloud, including:

[0010] Extract point cloud clusters through clustering;

[0011] Based on the displacement vector calculation of the centroids of adjacent frame point cloud clusters, screen out the point cloud clusters with displacement exceeding the dynamic threshold as the dynamic interference area;

[0012] S400. Align the visual dynamic mask and the lidar dynamic point cloud clusters in space-time, and determine the final real dynamic interference area, specifically including:

[0013] Project the visual dynamic mask onto the lidar coordinate system and calculate its overlap rate with the dynamic point cloud clusters;

[0014] For the area with an overlap rate exceeding the preset threshold, it is determined as the real dynamic interference area, and for the area with an overlap rate lower than the preset threshold, it is determined as the uncertain area, and secondary verification is performed in combination with IMU motion prediction;

[0015] S500. Eliminate the point cloud and visual features in the real dynamic interference area, retain the static background point cloud, and update the positioning reference based on the geometric and texture features of the static point cloud.

[0016] Furthermore, the process of spatial calibration of the lidar and the stereo vision camera in S100 is as follows: Collect multiple groups of image and point cloud data within the camera's field of view and the lidar's scanning range. Extract the corner coordinates (x c , y c ) from the images captured by the camera, and simultaneously obtain the three-dimensional coordinates (X l , Y l , Z l ) of the corresponding corners in the lidar point cloud. Establish an overdetermined system of equations for the image coordinates and the lidar coordinate system, that is, for each corner, there is an equation: where (u, v) are the pixel coordinates of the corner in the image, extracted from the images captured by the camera, λ is the scale factor, K is the camera internal parameter matrix and f x 、f yare the focal lengths of the camera in the x and y directions, respectively, and c x and c y are the coordinates of the principal point of the image. R is the rotation matrix used to describe the rotation relationship between the camera coordinate system and the lidar coordinate system, and T is the translation vector that describes the translation offset of the camera coordinate system relative to the lidar coordinate system. After collecting multiple sets of corner coordinates, substitute them into the equation to form an overdetermined system of equations, thereby obtaining the camera internal parameter matrix K, the rotation matrix R, and the translation vector T, and establishing the spatial mapping relationship between the lidar and the stereo vision camera.

[0017] Furthermore, in S200, real-time semantic segmentation is performed on the stereo vision image, and features are continuously extracted from the input image I through the downsampling path. During the downsampling process, the size of the feature map is halved and the number of channels is doubled. That is, the number of downsampling times is n, and the size of the feature map after the i-th downsampling is H i ×W i ×C i , satisfying C i = 2C i-1 , thereby extracting image features of different scales. The upsampling path fuses the high-dimensional features obtained by downsampling with the low-dimensional features of the corresponding layer. The upsampling restores the high-dimensional features to the same size as the corresponding low-dimensional features through deconvolution operations and then performs splicing fusion. Finally, a pixel-level dynamic mask M is output. When splicing and fusing, the feature channels at the same position are merged to accurately identify the pixel range of dynamic objects.

[0018] Furthermore, the steps for the S300 to extract the point cloud clusters are as follows:

[0019] In the lidar point cloud space, set the neighborhood search radius r and the minimum point number threshold N. Taking P0 as the seed point, search for the point set S within its neighborhood with a distance less than r through the data structure;

[0020] Divide the space along the coordinate axes and construct a binary tree data structure;

[0021] During the search, start from the root node and search the left and right subtrees according to the position relationship between the query point and the node hyperplane to narrow the search range;

[0022] When |S|≥N, classify these points into a point cloud cluster, and continue to repeat the search process with the points in S as the seed points until all points are clustered.

[0023] Furthermore, S300 extracts the centroid coordinates (X t , Y t , Z t ) of the t-th frame point cloud cluster C t from the clustered point cloud clusters, and the corresponding point cloud cluster C t+1 centroid coordinates (X t +1, Y t+1 , Z t+1 ), then the displacement vector Its displacement Perform weighted averaging on the coordinates of all points within each point cloud cluster, with a weight of 1, that is where m is the number of points in the point cloud cluster C t X tj represents the X-axis coordinate of the j-th point in the point cloud cluster C t in the t-th frame in the lidar coordinate system, Y tj is the coordinate of this point on the Y-axis, Z tj is its coordinate on the Z-axis. When D > D th , this point cloud cluster is screened as a dynamic interference area.

[0024] Furthermore, in the S400, project the visual dynamic mask onto the lidar coordinate system. Using the camera intrinsic matrix K, rotation matrix R, and translation vector T, for the pixel point (u, v) in the visual dynamic mask, use the inverse matrix K -1 of the camera intrinsic matrix K. Through back-projection λ obtain the three-dimensional coordinates (x c , y c , λ) in the camera coordinate system, where K -1 is the inverse matrix of the camera intrinsic matrix K. Through this inverse matrix, convert the pixel coordinates (u, v) into the ray direction in the camera coordinate system, combine with the depth information λ to obtain the three-dimensional coordinates, and then through coordinate transformation transform to the lidar coordinate system, calculate its overlap rate O with the dynamic point cloud cluster. The overlap rate O is obtained by statistically calculating the ratio of the number of points falling into the lidar dynamic point cloud cluster in the projected visual dynamic mask area to the total number of points in the visual dynamic mask area, that is where P v is the point set in the projected visual dynamic mask area, P l is the point set of the lidar dynamic point cloud cluster. For the area where the overlap rate O exceeds the threshold O th , that is, O ≥ O th , it is determined as a real dynamic interference area.

[0025] Furthermore, in the S400, for the area where the overlap rate O does not exceed the threshold O th , that is, O < O th , it is determined as an uncertain area, and combined with IMU motion prediction for secondary verification. During the verification process, the IMU measures the acceleration and angular velocity to calculate the displacement associated with the attitude change Δθ, i.e., within the time interval Δt, the displacement Displacement attitude change wherein is the initial displacement, is the initial velocity, Δθ0 is the initial attitude, and the attitude change is obtained by integrating the angular velocity Using the predicted displacement and attitude to update the spatial position of the uncertain region, and re-evaluating whether it is a dynamic region. The re-evaluation is performed by transforming the initial coordinates of the uncertain region according to the displacement and attitude changes, i.e., where P old are the initial coordinates of the uncertain region, P new are the updated coordinates, and R(Δθ) is the rotation matrix obtained according to the attitude change Δθ. Then, it is compared with the information of the dynamic point cloud clusters of the lidar again to determine whether it is a real dynamic interference region.

[0026] Furthermore, the specific steps of the S500 are as follows:

[0027] Remove the point cloud and visual features within the real dynamic interference region determined in S400 to eliminate the interference of dynamic factors on positioning;

[0028] Extract plane features and corner features from the remaining static point cloud;

[0029] According to the mapping relationship established by the S100 spatial calibration, match the plane features of the static point cloud with the corner features in the visual image to fuse the static information of different sensors;

[0030] Use the fused static information to predict and update the positioning reference.

[0031] On the other hand, a positioning system for lidar and stereo vision complementary positioning fusion, the components of the system include: a data acquisition module, a dynamic object recognition module, a dynamic interference division module, a dynamic interference determination module, and a positioning fusion module;

[0032] The data acquisition module uses a lidar and a stereo vision camera to respectively collect the point cloud data and visual image data of the environment, align the data according to the time stamp, and perform spatial calibration to construct a spatial mapping relationship;

[0033] The dynamic object recognition module performs semantic segmentation on the images collected by the stereo vision camera, identifies the dynamic object regions in the images, and generates a pixel-level dynamic mask;

[0034] The dynamic interference division module performs motion vectorization analysis on the lidar point cloud data, and screens out the point cloud clusters with a displacement amount exceeding the dynamic threshold as dynamic interference regions;

[0035] The dynamic interference determination module combines the semantic segmentation result and the lidar point cloud data to determine the final real dynamic interference area;

[0036] The positioning fusion module eliminates the dynamic interference point cloud and updates the positioning according to the remaining static background features.

[0037] Compared with the prior art, the positioning method and system for complementary positioning fusion of lidar and stereo vision have the following beneficial effects:

[0038] First, through the cross-modal collaboration of visual semantic segmentation and lidar motion vectorization analysis, the present invention realizes the accurate detection and real-time elimination of dynamic interference. Visual semantic segmentation provides pixel-level dynamic object recognition, while lidar point cloud cluster displacement analysis can quantify the motion trajectory. Through the spatio-temporal alignment mechanism, the two effectively distinguish real dynamic targets from sensor noise, avoiding positioning errors caused by single-sensor data errors.

[0039] Second, the present invention adopts the method of complementary positioning fusion of lidar and stereo vision, giving full play to the advantages of the two sensors and significantly enhancing the adaptability of the positioning system to complex environments. Under different conditions, stereo vision can provide rich texture information to assist positioning, while lidar can stably obtain distance information. By fusing stereo vision images and lidar data and continuously updating the positioning reference based on the static background in dynamic scenarios, the system can quickly adapt to environmental changes.

[0040] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0042] Figure 1 is the operation flowchart of the positioning method for complementary positioning fusion of lidar and stereo vision;

[0043] Figure 2 is the block diagram of the composition of the positioning system for complementary positioning fusion of lidar and stereo vision. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, elaborate in detail on the specific implementation manners, structures, features, and effects of the present invention as follows.

[0045] Embodiment 1

[0046] As Figure 1 shown, this embodiment elaborates in detail the specific steps of the positioning method for the complementary positioning fusion of lidar and stereo vision. By elaborating in detail the entire process from data acquisition to positioning reference update, the working principle of this technology in a complex dynamic environment is demonstrated. Using visual semantic segmentation and lidar point cloud motion vectorization analysis, dynamic interference is effectively identified and eliminated, and the positioning reference is updated based on static background features to achieve high-precision and high-robustness positioning.

[0047] First, enter the data acquisition stage (S100) to obtain the original data of the environment. The lidar and stereo vision camera work simultaneously to collect the point cloud data and visual image data of the environment respectively. The lidar obtains the distance information of surrounding objects by emitting laser beams and measuring the time delay of the reflected light, thereby generating three-dimensional point cloud data. These point cloud data can accurately represent the spatial position of objects but lack texture information. The stereo vision camera obtains rich texture information by taking pictures of the surrounding environment. However, it is difficult to accurately measure the distance of objects only relying on the stereo vision camera. In order to enable the data collected by the two sensors to work together, timestamp alignment and spatial calibration are required. Timestamp alignment ensures that the lidar point cloud data and stereo vision image data are collected at the same moment to avoid data inconsistency problems caused by time differences. Spatial calibration is to establish the conversion relationship between the lidar coordinate system and the camera coordinate system. For the calibration object within the camera's field of view and the lidar scanning range, the camera takes pictures and extracts the corner coordinates (x c , y c ), and at the same time, the lidar obtains the three-dimensional coordinates (X l , Y l , Z l ) of the corresponding corner points. For each corner point, there is an equation: where (u, v) are the pixel coordinates of the corner point in the image, which are extracted from the image taken by the camera, λ is the scale factor, and the camera intrinsic matrix f x , f y are the focal lengths of the camera in the x and y directions respectively, and c x , c yis the principal point coordinates of the image, R is the rotation matrix describing the rotation relationship between the camera coordinate system and the lidar coordinate system, T is the translation vector describing the translation offset of the camera coordinate system relative to the lidar coordinate system. After collecting multiple sets of corner coordinates, substituting them into the equation forms an overdetermined system of equations, thereby obtaining the camera internal parameter matrix K, the rotation matrix R, and the translation vector T, and establishing an accurate spatial mapping relationship between the two. In this way, the information in the camera image can be accurately converted to the lidar coordinate system for unified processing later.

[0048] Then, enter the dynamic object recognition stage (S200). Perform real-time semantic segmentation on the images collected by the stereo vision camera, identify the dynamic object regions in the images and generate a pixel-level dynamic mask. By downsampling the input image I, during this process, continuously extract the features of the image. During downsampling, the size of the feature map is halved and the number of channels is doubled. That is, the number of downsampling times is n, and after the i-th downsampling, the size of the feature map is H i ×W i ×C i , satisfying C i =2C i-1 , in this way, the network can extract image features at different scales and capture the detailed information in the image, such as the edges and textures of objects. After downsampling, the image features are compressed into a small high-dimensional space, and then enter the upsampling path. Upsampling restores the high-dimensional features to the same size as the corresponding low-dimensional features through deconvolution operations and then performs splicing and fusion, expands the channel information of the high-dimensional features, and restores the size of the image. When splicing and fusing, the feature channels at the same position are merged to enrich the feature information. Finally, the network outputs a pixel-level dynamic mask M, which classifies each pixel in the image and clearly indicates whether the pixel belongs to a dynamic object.

[0049] Subsequently, enter the dynamic interference division stage (S300). Perform motion vectorization analysis on the lidar point cloud data, and screen out the point cloud clusters with displacement amounts exceeding the dynamic threshold as dynamic interference regions. By extracting the point cloud clusters, in the lidar point cloud space, set the neighborhood search radius r and the minimum point number threshold N. Taking P0 as the seed point, use the data structure to search for the point set S within its neighborhood with a distance less than r. By continuously dividing the space along the coordinate axes, construct a binary tree structure. When searching, start from the root node and search the left subtree or the right subtree according to the position relationship between the query point and the node hyperplane to quickly narrow the search range and improve the search efficiency. When |S|≥N, classify these points into a point cloud cluster, and then continue to repeat the search process with the points in S as the seed points until all points are clustered, completing the extraction of the point cloud clusters. Calculate the displacement vector and displacement amount for the extracted and clustered point cloud clusters. For the clustered point cloud clusters, extract the centroid coordinates (X t of the t-th frame point cloud cluster Ct , Y t , Z t ), the point cloud cluster C corresponding to the (t + 1)-th frame t+1 The centroid coordinates (X t+1 , Y t+1 , Z t+1 ). The centroid coordinates are obtained by weighted averaging the coordinates of all points within each point cloud cluster, with the weight value being 1, that is where m is the number of points in the point cloud cluster C t , X tj represents the X-axis coordinate of the j-th point in the point cloud cluster C in the lidar coordinate system at the t-th frame, Y t is the Y-axis coordinate of this point, Z tj is its Z-axis coordinate, and the displacement vector tj The displacement amount The displacement amount Preset a dynamic threshold D th , when D > D th , it indicates that the point cloud cluster has undergone displacement between two frames and belongs to a dynamic object. Therefore, this point cloud cluster is screened as a dynamic interference area.

[0050] Secondly, enter the dynamic interference determination stage (S400). Combine the semantic segmentation result (i.e., the pixel-level dynamic mask) and the lidar point cloud data to determine the final real dynamic interference area. Project the visual dynamic mask onto the lidar coordinate system. For the pixel point (u, v) in the visual dynamic mask, use the inverse matrix K -1 of the camera internal parameter matrix K, and through back-projection λ to obtain the three-dimensional coordinates (x c , y c , λ) in the camera coordinate system. Here, K -1 converts the pixel coordinates (u, v) into the ray direction in the camera coordinate system, and combines the depth information λ to obtain the three-dimensional coordinates. After coordinate transformation Convert to the lidar coordinate system, and calculate the overlap rate O between the projected visual dynamic mask area and the lidar dynamic point cloud cluster. The overlap rate O is obtained by statistically calculating the ratio of the number of points falling into the lidar dynamic point cloud cluster within the projected visual dynamic mask area to the total number of points within the visual dynamic mask area, that is where P v is the point set P within the projected visual dynamic mask area l is the point set of the lidar dynamic point cloud cluster. For the area where the overlap rate O exceeds the threshold O th , that is, O ≥ O th , it is determined as the real dynamic interference area. For the area where the overlap rate O does not exceed the threshold O th , that is, O < Oth , it is determined as an uncertain area. For the uncertain area, secondary verification is carried out in combination with IMU motion prediction. The IMU measures acceleration and angular velocity Within the time interval Δt, the displacement and attitude change Δθ are calculated by integration. The displacement where is the initial displacement, is the initial velocity, and the attitude change where Δθ0 is the initial attitude. The predicted displacement and attitude are used to update the spatial position of the uncertain area. The initial coordinates of the uncertain area are transformed according to the displacement and attitude change, that is where P old is the initial coordinate of the uncertain area, and P new is the updated coordinate. R(Δθ) is the rotation matrix obtained according to the attitude change Δθ. It is compared with the lidar dynamic point cloud cluster information again to determine whether it is a real dynamic interference area.

[0051] Finally, it enters the positioning fusion stage (S500). The positioning fusion module removes the point cloud and visual features within the determined real dynamic interference area, excludes the interference of dynamic factors on positioning, extracts the plane features and corner features from the remaining static point cloud. For the static point cloud, the plane features are extracted. According to the mapping relationship established by spatial calibration, the plane features of the static point cloud are matched with the corner features in the visual image, and the static information of different sensors is fused. Through this matching, the geometric information of the lidar can be combined with the texture information of the stereo vision to provide a richer and more accurate environmental description. Using the fused static information, the positioning reference is predicted and updated, and the predicted pose is corrected, so as to obtain a more accurate positioning result and update the positioning reference.

[0052] In summary, through the close cooperation of the above steps S100 - S500, it can be clearly seen the working process of the positioning method of lidar and stereo vision complementary positioning fusion in a complex dynamic environment. From obtaining the original data, through dynamic object recognition, dynamic interference division, and dynamic interference determination, the dynamic interference is gradually and accurately identified and removed. Finally, the positioning reference is updated using the static background features, realizing high-precision and high-robustness positioning.

[0053] Embodiment 2

[0054] Based on Embodiment 1, as Figure 2As shown in the figure, this embodiment provides the system composition of a positioning system that integrates complementary positioning of lidar and stereo vision. By elaborating on the functions and collaborative working methods of each module of the system, it demonstrates the role of the system in achieving high-precision positioning, effectively coping with complex dynamic environments, and providing reliable support for the application of precise positioning.

[0055] The positioning system that integrates complementary positioning of lidar and stereo vision consists of: a data acquisition module, a dynamic object recognition module, a dynamic interference division module, a dynamic interference determination module, and a positioning fusion module. Among them:

[0056] The data acquisition module is responsible for obtaining raw data from the external environment. It integrates a lidar and a stereo vision camera. The lidar emits laser beams and receives reflected light to accurately measure the distance between itself and surrounding objects, and then generates three-dimensional point cloud data. The stereo vision camera captures images of the surrounding environment, obtains rich texture information, and performs timestamp alignment and spatial calibration on the data.

[0057] The dynamic object recognition module works based on the image data collected by the stereo vision camera. This module performs real-time semantic segmentation on the input images to determine whether they belong to dynamic objects and further determine the categories of dynamic objects. After identifying the dynamic object regions, the module generates pixel-level dynamic masks, achieving precise positioning and separation of dynamic objects, and providing an important basis for subsequent interference processing.

[0058] The dynamic interference division module deeply analyzes the point cloud data collected by the lidar. It processes the lidar point cloud data of consecutive frames, divides the point cloud into different point cloud clusters through clustering, tracks the movement of each point cloud cluster between different frames, and determines whether the point cloud cluster is in a moving state by calculating the displacement of the centroid of the point cloud cluster between adjacent frames. If the displacement of the centroid of a point cloud cluster exceeds a pre-set dynamic threshold, then this point cloud cluster is determined to be a possible dynamic interference source.

[0059] The dynamic interference determination module synthesizes the results of the dynamic object recognition module and the dynamic interference division module to accurately determine the real dynamic interference region. It projects the pixel-level dynamic mask generated by the dynamic object recognition module into the lidar coordinate system and compares it with the dynamic point cloud clusters obtained by the dynamic interference division module. By calculating the overlap rate between the two, it preliminarily determines whether a certain region is a dynamic interference region. For uncertain regions, the module uses the data of the IMU for secondary verification to re-determine whether it is a dynamic interference region, thereby improving the accuracy of the judgment.

[0060] The described positioning fusion module is responsible for the final positioning calculation and the update of the positioning reference. It eliminates the point clouds and visual features within the dynamically interfered areas determined by the dynamic interference determination module, extracts key features from the static background data, and uses the mapping relationship established by spatial calibration to match and fuse the planar features of the static point clouds with the corner features in the visual images. Based on the fused static information, it predicts and updates the positioning reference. By continuously updating the positioning reference, the system can adapt to environmental changes in real time and provide high-precision positioning results.

[0061] In summary, this embodiment has introduced in detail each component module of the positioning system for the complementary positioning fusion of lidar and stereo vision and its working principle. The modules cooperate closely with each other, giving full play to the complementary advantages of lidar and stereo vision, enabling the entire positioning system to still work stably and accurately in a complex dynamic environment.

[0062] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to form equivalent embodiments with equivalent changes, but as long as the technical content of the present invention is not departed from, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A positioning method for the complementary positioning fusion of lidar and stereo vision, characterized in that, The specific steps of this method are as follows: S100. Synchronously acquire lidar point cloud data and stereo vision image data, and perform timestamp alignment and spatial calibration on both; S200. Perform real-time semantic segmentation on the stereo vision image, identify the dynamic object area, and generate a pixel-level dynamic mask; S300. Perform motion vectorization analysis on consecutive frames of lidar point cloud, including: Extract point cloud clusters through clustering; Based on the calculation of the displacement vector of the centroid of adjacent frame point cloud clusters, screen out the point cloud clusters with displacement exceeding the dynamic threshold as the dynamic interference area; S400. Align the visual dynamic mask and the lidar dynamic point cloud clusters in space-time, and determine the final real dynamic interference area, specifically including: Project the visual dynamic mask into the lidar coordinate system and calculate its overlap rate with the dynamic point cloud clusters; For the area with an overlap rate exceeding the preset threshold, it is determined as the real dynamic interference area, and for the area with an overlap rate lower than the preset threshold, it is determined as the uncertain area, and secondary verification is performed in combination with IMU motion prediction; S500. Remove the point cloud and visual features in the real dynamic interference area, retain the static background point cloud, and update the positioning reference based on the geometric and texture features of the static point cloud.

2. The positioning method for the complementary positioning fusion of lidar and stereo vision according to claim 1, wherein, The process of spatially calibrating the lidar and the stereo vision camera in S100 is as follows: Collect multiple sets of image and point cloud data within the camera's field of view and the lidar's scanning range. Extract the corner coordinates (x c , t c ) from the images captured by the camera. At the same time, obtain the three-dimensional coordinates (X l , Y l , Z l ) of the corresponding corners in the lidar point cloud. Establish an overdetermined system of equations for the image coordinates and the lidar coordinate system. That is, for each corner, there is an equation: where (u, v) are the pixel coordinates of the corner in the image, extracted from the images captured by the camera, λ is the scale factor, K is the camera intrinsic matrix and f x , f y are the focal lengths of the camera in the x and y directions respectively, c x , c y are the coordinates of the principal point of the image, R is the rotation matrix, used to describe the rotation relationship of the camera coordinate system relative to the lidar coordinate system, and T is the translation vector, describing the translation offset of the camera coordinate system relative to the lidar coordinate system. After collecting multiple sets of corner coordinates, substitute them into the equation to form an overdetermined system of equations, so as to obtain the camera intrinsic matrix K, the rotation matrix R and the translation vector T, and establish the spatial mapping relationship between the lidar and the stereo vision camera.

3. The positioning method for the complementary positioning fusion of lidar and stereo vision according to claim 1, characterized in that, In the S200, real-time semantic segmentation is performed on the stereo vision image. Features are continuously extracted from the input image I through the downsampling path. During the downsampling process, the size of the feature map is halved and the number of channels is doubled. That is, if the number of downsampling times is n, the size of the feature map after the i-th downsampling is H i ×W i ×C i , satisfying C i = 2C i-1 , so as to extract image features of different scales. The upsampling path fuses the high-dimensional features obtained by downsampling with the low-dimensional features of the corresponding layer. The upsampling restores the high-dimensional features to the same size as the corresponding low-dimensional features through deconvolution operations and then performs splicing and fusion. Finally, a pixel-level dynamic mask M is output. During the splicing and fusion, the feature channels at the same position are merged to accurately identify the pixel range of dynamic objects.

4. The positioning method for the complementary positioning fusion of lidar and stereo vision according to claim 1, characterized in that, The steps of extracting point cloud clusters in S300 are as follows: In the lidar point cloud space, set the neighborhood search radius r and the minimum point number threshold N. Taking P0 as the seed point, search for the point set S within its neighborhood with a distance less than r through the data structure; Divide the space along the coordinate axes and construct a binary tree data structure; During the search, start from the root node, and search the left subtree and the right subtree according to the position relationship between the query point and the node hyperplane to narrow the search range; When |S|≥N, these points are grouped into a point cloud cluster, and the search process is repeated with the points in S as the seed points until all points are clustered.

5. The positioning method for the complementary positioning fusion of lidar and stereo vision according to claim 4, characterized in that, The S300 extracts the t-th frame of point cloud cluster C from the clustered point cloud clusters t The centroid coordinates (X t , Y t , Z t ), and the centroid coordinates of the corresponding point cloud cluster C t +1 of the (t + 1)-th frame (X t +1, Y t+1 , Z t+1 ). Then the displacement vector Its displacement Perform weighted averaging on the coordinates of all points within each point cloud cluster, with the weight value being 1, that is where m is the number of points in the point cloud cluster C t , X tj represents the X-axis coordinate of the j-th point in the t-th frame of point cloud cluster C t in the lidar coordinate system, Y tj is the Y-axis coordinate of this point, and Z tj is its Z-axis coordinate. When D > D th , this point cloud cluster is screened as a dynamic interference area.

6. The positioning method for the complementary positioning fusion of lidar and stereo vision according to claim 2, wherein In the S400, the visual dynamic mask is projected into the lidar coordinate system. Using the camera intrinsic matrix K, rotation matrix R, and translation vector T, for the pixel point (u, v) in the visual dynamic mask, the inverse matrix K of the camera intrinsic matrix K is used -1 , and through back-projection , the three-dimensional coordinates (x c , y c , λ) in the camera coordinate system are obtained. Here, K -1 is the inverse matrix of the camera intrinsic matrix K. Through this inverse matrix, the pixel coordinates (u, v) are converted into the ray direction in the camera coordinate system. Combining the depth information λ, the three-dimensional coordinates are obtained, and then through coordinate transformation , they are transformed into the lidar coordinate system, and the overlap rate O with the dynamic point cloud cluster is calculated. The overlap rate O is obtained by statistically calculating the ratio of the number of points falling into the lidar dynamic point cloud cluster in the projected visual dynamic mask area to the total number of points in the visual dynamic mask area, that is where P v is the point set in the projected visual dynamic mask area, and P l is the point set of the lidar dynamic point cloud cluster. For the area where the overlap rate O exceeds the threshold O th , that is, O ≥ O th , it is determined as the real dynamic interference area.

7. The positioning method for the complementary positioning fusion of lidar and stereo vision according to claim 1, wherein In the S400, for the region where the overlap rate O does not exceed the threshold O th That is, O < O th , it is determined as an uncertain region, and secondary verification is carried out in combination with IMU motion prediction. During the verification process, the IMU measures the acceleration and the angular velocity to obtain the displacement and the attitude change Δθ through integral calculation. That is, within the time interval Δt, the displacement Displacement attitude change where is the initial displacement, is the initial velocity, Δθ0 is the initial attitude, and the attitude change is obtained by integrating the angular velocity . The spatial position of the uncertain region is updated using the predicted displacement and attitude, and it is re-evaluated whether it is a dynamic region. The re-evaluation is carried out by transforming the initial coordinates of the uncertain region according to the displacement and attitude changes, that is where P old is the initial coordinate of the uncertain region, P new is the updated coordinate, and R(Δθ) is the rotation matrix obtained according to the attitude change Δθ. It is compared with the lidar dynamic point cloud cluster information again to determine whether it is a real dynamic interference region.

8. The positioning method for the complementary positioning fusion of lidar and stereo vision according to claim 1, characterized in that, The specific steps of S500 are as follows: Remove the point cloud and visual features in the real dynamic interference area determined in S400 to exclude the interference of dynamic factors on positioning; Extract plane features and corner features from the remaining static point cloud; According to the mapping relationship established by the S100 spatial calibration, match the plane features of the static point cloud with the corner features in the visual image to fuse the static information of different sensors; Use the fused static information to predict and update the positioning reference.

9. A positioning system for the complementary positioning fusion of lidar and stereo vision, applicable to the method for the complementary positioning fusion of lidar and stereo vision according to any one of claims 1-8, characterized in that, The components of this system include: a data acquisition module, a dynamic object recognition module, a dynamic interference division module, a dynamic interference determination module, and a positioning fusion module; The data acquisition module uses a lidar and a stereo vision camera to collect the point cloud data and visual image data of the environment respectively, align the data according to the timestamp, and perform spatial calibration to construct a spatial mapping relationship; The dynamic object recognition module performs semantic segmentation on the images collected by the stereo vision camera, identifies the dynamic object areas in the images, and generates a pixel-level dynamic mask; The dynamic interference division module performs motion vectorization analysis on the lidar point cloud data, and screens out the point cloud clusters with displacement exceeding the dynamic threshold as the dynamic interference area; The dynamic interference determination module combines the semantic segmentation results and the lidar point cloud data to determine the final real dynamic interference area; The positioning and fusion module eliminates dynamic interference point clouds and updates the positioning according to the remaining static background features.

Citation Information

Patent Citations

  • Target detection and motion state estimation method based on vision and laser radar

    CN111951305A

  • Vehicle positioning method and system considering communication delay in vehicle-road cooperation environment

    CN114877883A