Device based on vision and 2D laser fusion positioning
The fusion of 2D laser and vertical vision modules with advanced data processing addresses the limitations of single-sensor systems, providing high-precision and robust positioning in complex industrial environments.
Patent Information
- Application Number
- CN202510808472.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
现有技术中,激光定位在重复特征和高动态环境下易失效,视觉定位受光照变化和遮挡干扰显著,单一模态定位方案难以实现工业场景全工况覆盖。
Using a device based on the fusion positioning of vision and 2D lasers, through the complementary layout of the horizontal ranging module and the vertical vision module, combined with the calculation processing module, a laser profile map and visual feature map are built with coordinate system aligned, dual-mode observation filtering update is realized, geometric constraints of laser profile matching and semantic constraints of visual reprojection are fused, composite maps are constructed and global consistency optimization is performed.
While ensuring the millimeter-level positioning accuracy, it significantly improves the robustness of industrial interference factors such as light changes, dynamic obstacles and structural repetition, providing all-weather and reliable positioning guarantee.
Smart Images

Figure CN120313613A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of positioning, and in particular to a device for visual and 2D laser fusion positioning. Background Art
[0002] Currently, in the field of mobile robots, similar to the debate between lidar routes and pure vision routes in driverless, there are also differences in technical routes between vision and laser in the natural navigation and positioning technology of AMR / AGV in industrial scenarios (i.e., positioning technology without using artificial markers such as QR codes and reflectors). Due to having accurate distance measurement information, the difficulty of the laser positioning technology route is less than that of vision positioning, and the current technology maturity is also significantly higher than that of vision positioning. Laser positioning has stability and operation accuracy to meet the application requirements of industrial scenarios in 80%-90% indoor / semi-indoor environments. Therefore, domestic AMR manufacturers basically focus on laser SLAM technology. However, laser positioning has a risk of loss in environments with repetitive features (such as long corridors) and high dynamics. In contrast, vision positioning technology is an emerging positioning technology. Although compared with laser positioning, a camera can capture rich environmental information, such as colors, textures, etc., which helps the robot to more accurately identify and understand the environment, vision positioning faces challenges in various complex environments such as light changes and obstacle occlusions, and its engineering application still faces huge challenges.
[0003] Therefore, how to effectively integrate the advantageous features of the two perception modalities and break through the performance limitations of a single-modal sensor under specific working conditions has become a key technical challenge for improving the full-scenario positioning ability of industrial mobile robots. Summary of the Invention
[0004] (I) Technical Problems to be Solved In view of the above-mentioned disadvantages and deficiencies of the prior art, the present invention provides a device for visual and 2D laser fusion positioning, which solves the technical problems that existing laser positioning is prone to failure in environments with repetitive features and high dynamics, vision positioning is significantly affected by light changes and occlusion interference, and a single-modal positioning scheme is difficult to cover all working conditions in industrial scenarios.
[0005] (II) Technical Solutions To achieve the above object, the main technical solutions adopted by the present invention include: In a first aspect, an embodiment of the present invention provides a device for visual and 2D laser fusion positioning, including: A horizontal ranging module, horizontally installed on the robot body, for collecting contour ranging data of the motion plane; A vertical vision module, top-view installed on the robot body, for collecting visual feature data of the vertical space; A calculation and processing module, configured to: construct a laser contour map and a visual feature map with coordinate system alignment based on contour ranging data and visual feature data, determine the initial pose of the robot according to the input given position or the optimal position obtained by globally matching the laser contour map and the visual feature map, generate a predicted pose in combination with motion measurement values, and iteratively perform dual-mode observation filtering updates of laser contour matching optimization and visual reprojection matching optimization centered on the predicted pose.
[0006] Optionally, a horizontal ranging module, including a 2D lidar, is adjusted by a pre-configured dynamic pitch compensation mechanism, and the parallelism error between the laser scanning plane of the 2D lidar and the traveling direction of the robot is less than a preset parallelism error. A vertical vision module, including an infrared sensor, is adjusted by a pre-configured tunable gimbal. The optical axis of the lens of the infrared sensor forms an included angle of 88° - 92° with the 2D lidar, so that the vertical field of view coverage range of the infrared sensor and the horizontal scanning sector of the 2D lidar form a non-overlapping complementary observation domain in the space coordinate system.
[0007] Optionally, the complementary observation domain satisfies the following spatial constraint conditions: The 2D lidar forms a scanning sector θ ∈ [0°, 270°] with adaptive opening angle adjustment in the X-Y plane of the robot's motion coordinate system. The infrared sensor constructs a three-dimensional pyramid-shaped detection area in the Z-axis direction of the Cartesian coordinate system. The pitch angle φ of the infrared sensor ∈ [55°, 95°] and forms a buffer isolation zone of not less than 15% with the scanning sector of the lidar in the spatial projection plane.
[0008] Optionally, an active infrared supplementary light source device is configured, and the active infrared supplementary light source device includes: An annular distribution unit, which arranges N independently addressable infrared emission sources circumferentially around the infrared sensor. A spectral modulation unit, which dynamically selects a working wavelength of 850 nm or 940 nm according to the ambient light entropy value and drives a tunable bandpass filter for matching spectral filtering. A light emission synchronization controller, which aligns the pulse light emission periods of each infrared emission source with the global shutter exposure window of the infrared sensor through a hardware trigger signal. A light source selection module, including a switchable LED light source group and a VCSEL light source group. Among them, the LED light source group includes multiple LEDs arranged in an annular array, and the light emitting surfaces of all LEDs are installed obliquely outward with respect to the central normal line, and the inclination angle range is 5° - 15°, and the inclination directions of the LEDs in the same layer are symmetrically distributed; the VCSEL light source group includes multiple vertical cavity surface emitting lasers.
[0009] Optionally, the calculation and processing module includes: A registration and mapping unit, configured to construct a laser contour map and a visual feature map with aligned coordinate systems based on contour ranging data and visual feature data; A pose prediction unit, configured to determine an initial pose of the robot according to a given input position or the optimal position obtained by globally matching the laser contour map and the visual feature map, and generate a predicted pose by combining motion measurement values; An iterative execution unit, configured to iteratively execute with the predicted pose as the center: perform an optimization calculation of the contour coincidence degree between the current contour ranging data and the laser contour map centered on the current predicted pose, and implement a first observation filtering update; extract semantic features of the current visual feature data, and perform an optimization of the feature point reprojection matching of the visual feature map using the semantic features centered on the current predicted pose, and implement a second observation filtering update; perform a positioning correction on the results of the first observation filtering update and the second observation filtering update, and use the corrected pose as the predicted input for the next cycle.
[0010] Optionally, the registration and mapping unit includes: A dynamic environment compensation sub-unit, configured to construct a motion distortion compensation field based on the integration of the angular velocity / linear velocity of the IMU and the odometer, perform point-by-point motion distortion removal processing on the contour ranging data, perform dynamic target detection based on density clustering on the distortion-removed contour ranging data, remove data with a moving speed exceeding a preset speed, and dynamically adjust the extraction sensitivity of the visual feature data according to the environmental illumination intensity, and enable adaptive histogram equalization preprocessing for the visual feature data when the illumination is less than a preset illumination threshold; A dynamic joint calibration sub-unit, configured to calculate an external parameter residual matrix between the lidar coordinate system and the visual sensor coordinate system according to the synchronously acquired contour ranging data and visual feature data, and trigger calibration when the external parameter residual exceeds a preset residual value or the pose estimation covariance value of a continuous preset number of frames is greater than a preset covariance value. Perform online calibration work by performing bimodal feature extraction on the reflectivity mutation edge of the contour ranging data and the temperature difference feature corner of the visual feature data; A multi-modal map construction sub-unit, configured to adopt a tightly coupled SLAM framework to jointly optimize the edge features of the contour ranging data and the SIFT features of the visual feature data, construct a laser contour map as the underlying information and a visual feature map as the top-level information respectively, and then perform the fusion of the laser contour map and the visual feature map, and establish a cross-modal feature association between the laser and the vision through bidirectional projection to obtain a composite map; A map optimization subunit, which is used to obtain a laser constraint edge representing the connection of continuous robot pose nodes by calculating the geometric contour matching degree between the current contour ranging data and the laser contour map on the composite map, and obtain a visual constraint edge representing the connection of visual feature observation nodes by calculating the three-dimensional feature point reprojection error between the current visual feature data and the visual feature map. An adaptive weight is set between the laser constraint edge and the visual constraint edge, and the weighted sum of squared residuals of the laser constraint edge and the visual constraint edge is synchronously minimized to achieve global consistency optimization of the composite map.
[0011] Optionally, a composite calibration target is used for on-line calibration work. The composite calibration target includes: Laser high-reflection stripes, which are grid patterns composed of alternating 20% reflectivity and 80% reflectivity; Infrared feature markers, which are AprilTag tags embedded in the temperature difference module and present controllable temperature difference features under thermal imaging.
[0012] Optionally, the pose prediction unit includes: A manual repositioning subunit, which is used to parse the initial position coordinates and their confidence radii input by the user, construct a three-degree-of-freedom search space, adopt a multi-resolution rasterization strategy, refine the matching granularity layer by layer in the search space to find the coordinate with the best matching degree as the initial pose, and generate a confidence report containing the position covariance matrix at the same time; An automatic repositioning subunit, which is used to synchronously perform laser contour matching and visual feature retrieval, generate a joint score by dynamically weighted fusion of laser geometric similarity and visual semantic matching degree, and obtain the best coordinate as the initial pose according to the comprehensive matching degree; A motion prediction subunit, which is used to fuse wheel odometer, IMU angular velocity and laser odometer data, and deduce the pose change amount by using a kinematic model; A prediction optimization subunit, which is used to generate multiple candidate pose particles based on Monte Carlo sampling, perform a composite likelihood evaluation of laser projection coverage and visual reprojection error for each particle, update the candidate pose particles by importance resampling, and output the weighted average pose as the prediction result.
[0013] Optionally, the iterative execution unit includes: A first observation filtering and updating subunit, which is used to establish an adaptive search area in the laser contour map with the current predicted pose as the center, adopt a multi-scale iterative closest point matching strategy, optimize the contour coincidence degree layer by layer between coarse-grained and fine-grained resolutions, apply an exponentially decaying weight to the contour ranging data points whose moving speed exceeds the preset speed threshold, and calculate the laser matching residual based on the contour matching result; The second observation filtering and updating subunit is used to extract semantic features in the current visual data in real time through a lightweight semantic network running on an embedded NPU, construct a three-dimensional projection space within the predicted pose neighborhood, reproject the three-dimensional feature points in the visual feature map to the current image coordinate system and perform two-way matching verification, and calculate the visual reprojection error based on the valid matching pairs obtained from the verification. The observation result correction subunit is used to construct a multi-modal Kalman filter, receive the laser matching residual and the visual reprojection error respectively, dynamically allocate sensor weights according to the environmental degradation factors, where the environmental degradation factors include the proportion of dynamic obstacles in the laser and the visual illumination uniformity. When the bimodal positioning deviation exceeds the safety threshold, an arbitration mechanism is activated, and the sensor data stream with a confidence level continuously higher than the preset confidence threshold in the past preset period is preferentially adopted, and the corrected pose is output as the prediction input for the next cycle.
[0014] (3) Beneficial effects The beneficial effects of the present invention are: through the deep integration of the heterogeneous sensor architecture and the computing platform, the present invention constructs an all-weather positioning solution for industrial scenarios.
[0015] Specifically, the complementary spatial layout of the horizontal ranging module and the vertical vision module can capture both the precise geometric information of the moving plane and the rich semantic features in the vertical direction at the same time, effectively overcoming the perception limitations of a single horizontal laser scan in complex three-dimensional scenarios. Secondly, the deep integration architecture of software and hardware significantly enhances the spatio-temporal consistency of multi-source data, not only eliminating the fusion error caused by communication delay in the traditional split solution, but also achieving sub-millimeter alignment of laser point clouds and visual features through a unified spatio-temporal reference, laying a physical foundation for multi-modal data fusion.
[0016] Furthermore, the bimodal observation filtering and updating mechanism adopted by the present invention innovatively integrates the geometric constraint of laser contour matching and the semantic constraint of visual reprojection to form a dual anti-interference ability in a dynamic interference environment: when there are temporary occlusions or sudden changes in illumination in the environment, the laser observation provides a stable geometric reference; while in scenarios with repetitive structures such as long corridors, visual features inject discriminative semantic information, and the two form a dynamic complementary fault tolerance mechanism. In addition, the composite map construction method with strictly aligned coordinate systems enables the system to concurrently utilize the precise geometric prior of the laser and the topological recognition ability of the vision during the global positioning stage. Even in the case of cold start of the device or severe loss of pose, it can still quickly restore a stable pose through cross-modal joint search.
[0017] Thus, through the optimization of hardware layout and the innovation of algorithm architecture, the present invention has successfully solved the core pain points of modal characteristic conflicts and single environmental adaptability in traditional solutions. While ensuring millimeter-level positioning accuracy, it has significantly improved the robustness against typical industrial interference factors such as light changes, dynamic obstacles, and structural repetitions, providing all-weather and highly reliable positioning guarantee for scenarios such as unmanned handling. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 FIG. is a schematic diagram of the composition of a device for visual and 2D laser fusion positioning provided by an embodiment of the present invention; Figure 2 FIG. is a schematic diagram of the internal process of the calculation and processing module of a device for visual and 2D laser fusion positioning provided by an embodiment of the present invention.
[0019]
DESCRIPTION OF REFERENCE NUMERALS
[0020] In order to better explain the present invention for easy understanding, the present invention will be described in detail below with reference to the accompanying drawings through specific embodiments.
[0021] As Figure 1 shown, a device for visual and 2D laser fusion positioning proposed by an embodiment of the present invention includes: a horizontal ranging module, horizontally installed on the robot body for collecting contour ranging data of the motion plane; a vertical vision module, top-view installed on the robot body for collecting visual feature data of the vertical space; a calculation and processing module configured to: construct a laser contour map and a visual feature map with coordinate system alignment based on the contour ranging data and visual feature data, determine the initial pose of the robot according to the input given position or the optimal position of self-global matching of the laser contour map and the visual feature map, generate a predicted pose in combination with the motion measurement value, and iteratively perform dual-modal observation filtering update of laser contour matching optimization and visual reprojection matching optimization centered on the predicted pose.
[0022] The present invention constructs an all-weather positioning solution for industrial scenarios through the deep integration of heterogeneous sensor architectures and computing platforms.
[0023] Specifically, the complementary spatial layout of the horizontal ranging module and the vertical vision module can capture both the precise geometric information of the moving plane and the rich semantic features in the vertical direction, effectively overcoming the perception limitations of a single horizontal laser scan in complex three-dimensional scenes. Secondly, the deep integration architecture of software and hardware significantly enhances the spatio-temporal consistency of multi-source data. It not only eliminates the fusion errors caused by communication delays in traditional split-type solutions but also achieves sub-millimeter alignment of laser point clouds and visual features through a unified spatio-temporal reference, laying a physical foundation for multi-modal data fusion.
[0024] Furthermore, the dual-modal observation filtering and updating mechanism adopted in the present invention innovatively integrates the geometric constraints of laser contour matching and the semantic constraints of visual reprojection to form a dual anti-interference ability in a dynamic interference environment: when there are temporary occlusions or sudden changes in illumination in the environment, the laser observation provides a stable geometric reference; while in scenes with repetitive structures such as long corridors, visual features inject discriminative semantic information, and the two form a dynamically complementary fault-tolerant mechanism. In addition, the composite map construction method with strictly aligned coordinate systems enables the system to concurrently utilize the precise geometric prior of the laser and the topological recognition ability of the vision during the global positioning stage. Even in the case of cold start of the device or severe loss of pose, it can still quickly restore a stable pose through cross-modal joint search.
[0025] Thus, through the optimization of the hardware layout and the innovation of the algorithm architecture, the present invention successfully solves the core pain points of modal characteristic conflicts and single environmental adaptability in traditional solutions. While ensuring millimeter-level positioning accuracy, it significantly improves the robustness against industrial typical interference factors such as illumination changes, dynamic obstacles, and repetitive structures, providing all-weather and highly reliable positioning guarantee for scenarios such as unmanned handling.
[0026] To better understand the above technical solutions, the exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more clear and thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0027] Specifically, the horizontal ranging module 1 includes a 2D lidar, which is adjusted by a pre-configured dynamic pitch compensation mechanism so that the parallelism error between the laser scanning plane of the 2D lidar and the robot's traveling direction is less than the set parallelism error. Referring again to Figure 1 , the 2D lidar is horizontally installed. The dynamic pitch compensation mechanism is composed of a linkage of a gyroscope sensor and a servo motor, which real-time detects the change in the robot's traveling posture and drives the rotation axis of the 2D lidar, so that the parallelism error between the laser scanning plane and the X-axis (traveling direction) of the robot's motion coordinate system is controlled within the range of ±0.5°.
[0028] The vertical vision module 2 includes an infrared sensor, which is adjusted by a pre-configured tunable gimbal. The tunable gimbal is built with an angle encoder and a locking mechanism, and supports the configuration that the optical axis of the lens of the infrared sensor forms an included angle of 88°-92° with the 2D lidar, so that the coverage range of the vertical field of view angle of the infrared sensor and the horizontal scanning sector of the 2D lidar form a non-overlapping complementary observation domain in the space coordinate system.
[0029] In particular, the complementary observation domain satisfies the following spatial constraint conditions: The 2D lidar forms a scanning sector with an adaptive opening angle adjustment in the X-Y plane of the robot motion coordinate system θ ∈[0°, 270°], and at the same time, the horizontal scanning sector θ of the 2D lidar is dynamically adjusted through an adaptive opening angle algorithm: when a narrow environment such as a long corridor is detected, it automatically shrinks to θ =120° to reduce redundant data; in the open area, it expands to θ =270° to achieve large-range detection. The infrared sensor constructs a three-dimensional pyramid-shaped detection area in the Z-axis direction of the Cartesian coordinate system. The pitch angle φ of the infrared sensor ∈[55°, 95°] and forms a buffer isolation zone of not less than 15% with the scanning sector of the lidar in the spatial projection plane. It should be emphasized that the parameter design of the buffer isolation zone of the positive interactive complementary detection domain mimics the distribution law of the rod cells of the compound eyes of insects, and realizes the optimal spatial coverage efficiency under limited sensor resources.
[0030] Next, the device of the present invention is configured with an active infrared supplementary light source device 3, and the active infrared supplementary light source device includes: An annular distribution unit is arranged circumferentially around the infrared sensor N independently addressable infrared emission sources; where N ∈[2, 8] and satisfies N = ceil ( H max / ΔH ), ceil is rounding up, H max is the extreme value of the vertical height of the scene, ΔH is the preset height resolution threshold.
[0031] A spectral modulation unit dynamically selects a working wavelength of 850 nm or 940 nm according to the environmental light entropy value, and drives a tunable band-pass filter for matching spectral filtering; when the short-wave infrared detector is activated, a wavelength of 940 nm is preferentially used to avoid visible light interference; when the long-wave infrared detector works, it automatically switches to a wavelength of 850 nm to enhance the detection sensitivity of the thermal contrast.
[0032] The light-emitting synchronization controller aligns the pulsed light-emitting periods of each infrared emission source with the global shutter exposure window of the infrared sensor through a hardware trigger signal.
[0033] The light source selection module includes a switchable LED light source group and a VCSEL light source group. Among them, the LED light source group includes multiple LEDs arranged in a circular array, and the light-emitting surfaces of all LEDs are installed obliquely outward with the central normal as the reference, and the inclination angle ranges from 5° to 15°, and the inclination directions of the LEDs on the same layer are symmetrically distributed; the VCSEL light source group includes multiple vertical cavity surface-emitting lasers. The LED light source group is suitable for scenes dominated by diffuse reflection, and the divergence angle ≥ 120°; the VCSEL light source group is suitable for scenes with specular reflection suppression, the quasi-vertical angle ≤ 10° and the power density > 200mW / cm².
[0034] It should be noted that the device of the present invention is also provided with a communication interface module 5 and a power supply module 6. The device of the present invention defaults to use a gigabit Ethernet communication interface. Through this interface, the positioning pose is published to the robot body. At the same time, according to the usage requirements, CAN communication or other communication methods can also be used to interact with the navigation system, etc. The power supply module 6 is responsible for stabilizing the voltage and supplying power to the calculation and processing unit, lidar module, infrared active supplementary lighting device, image sensor, etc. Since the infrared automatic supplementary lighting device is instantaneously supplemented with light and the current changes greatly instantaneously, it is necessary to ensure the isolation between modules.
[0035] Furthermore, as Figure 2 shown, the calculation and processing module 4 includes: The registration and mapping unit is used to construct a laser contour map and a visual feature map with coordinate system alignment based on the contour ranging data and visual feature data.
[0036] In one embodiment, the mapping mode is started, and the mobile robot carrying the device is controlled to walk along the path in the running environment. When controlling the mobile robot to travel along the preset path, the horizontal contour ranging data of the 2D lidar and the vertical visual feature data of the infrared sensor are synchronously collected at a frequency of 10Hz, and the spatio-temporal alignment of the dual-modal data is realized through the hardware timestamp. After that, the obtained laser contour map is the plane contour map scanned by the lidar, and the visual feature map is the upper visual feature map in the environment. The visual feature point cloud is projected onto the laser contour map coordinate system through the extrinsic calibration matrix, and the origin (the initial pose point of the mobile robot) and the X / Y axis directions (consistent with the lidar scanning plane) of the two types of maps are forced to be unified to achieve consistent spatial mapping. After the mapping work is completed, the map is loaded into the device, and the positioning mode can be entered.
[0037] The pose prediction unit is used to determine the initial pose of the robot according to the input given position or the optimal position obtained by globally matching the laser contour map and the visual feature map, and generate a predicted pose in combination with the motion measurement value.
[0038] The pose prediction unit innovatively integrates dual modes of manual positioning (i.e., manual prior guidance) and automatic positioning (i.e., multi-sensor data-driven): (1) In the manual mode, the restricted neighborhood search algorithm is adopted. The roughly specified position manually is used as the center of the Gaussian distribution. Through laser profile matching and visual feature correlation calculation, sub-meter-level fast convergence is achieved, solving the problem that traditional manual calibration is easily affected by subjective errors, and finding the coordinates with the best matching degree as the initial pose; (2) In the automatic mode, according to the obtained lidar data and image data, they are simultaneously matched with the laser map and the visual map, and the coordinates with the best matching degree in the global map are obtained as the initial pose.
[0039] And, the iterative execution unit is used to iteratively execute centered on the predicted pose: optimize the contour coincidence degree calculation between the current contour ranging data and the laser contour map centered on the current predicted pose, and implement the first observation filter update; extract the semantic features of the current visual feature data, and use the semantic features to optimize the feature point reprojection matching of the visual feature map centered on the current predicted pose, and implement the second observation filter update; perform positioning correction on the results of the first observation filter update and the second observation filter update, and use the corrected pose as the prediction input for the next cycle.
[0040] Furthermore, the registration and mapping unit includes: The dynamic environment compensation sub-unit is used to construct a motion distortion compensation field based on the integration of the angular velocity / linear velocity of the IMU and the odometer, perform point-by-point motion distortion removal processing on the contour ranging data, perform dynamic target detection based on density clustering on the distortion-removed contour ranging data, remove data with a moving speed exceeding the set speed (such as 0.2 m / s), and at the same time dynamically adjust the extraction sensitivity of the visual feature data according to the environmental illumination intensity, and enable adaptive histogram equalization preprocessing for the visual feature data when the illuminance < 50 lux.
[0041] Specifically, based on the IMU angular velocity ω and the odometer linear velocity v , the motion compensation basic matrix is constructed as: ; In the formula, ω ( t ) is the instantaneous angular velocity, is the cumulative rotation angle from the start time t0 of the scanning period to the current moment t , v x ( t ) and v y ( t ) are the X / Y axis linear velocities provided by the odometer respectively, Construct the motion distortion compensation field based on the motion compensation fundamental matrix as follows: ; In the formula, λ k is the contribution weight of the k th laser point to the current compensation field, ( x , y ) is the coordinate of any point in the compensation field, ([[]] x k , y k ) is the original measurement value of the k th point actually collected by the 2D lidar, σ is the Gaussian kernel bandwidth parameter, σ = R • tan ( θ / 2), R is the measured distance, θ is the angular resolution, N is the total number of data collected by the 2D lidar.
[0042] It should be clear that the online calibration work uses a composite calibration target, which includes: laser high-reflection stripes, a grid pattern composed of alternating 20% reflectivity and 80% reflectivity; infrared feature markers, AprilTag tags embedded in the temperature difference module, which present controllable temperature difference characteristics under thermal imaging.
[0043] The dynamic joint calibration sub-unit is used to calculate the external parameter residual matrix between the lidar coordinate system and the vision sensor coordinate system according to the synchronously collected contour ranging data and visual feature data. When the external parameter residual exceeds the set residual value of 0.15 or the pose estimation covariance determinant for 5 consecutive frames triggers calibration, online calibration work is carried out by performing dual-modal feature extraction on the reflectivity mutation edge of the contour ranging data and the temperature difference feature corner of the visual feature data.
[0044] Specifically, the online calibration work includes: extracting metal edge features with a reflectivity > 25% / m and an incident angle > 60° (such as door frames, shelf edges) to generate a laser edge point set P L , and detecting infrared high-contrast corner points with a temperature difference > 2℃ / pixel (such as lamp heat sources, ventilation openings) to generate a visual corner point set P V , projecting P L onto the image plane and performing two-way RANSAC matching with P V to establish matching pairs for solving the optimal external parameters.
[0045] The multi-modal map construction subunit is used to adopt a tightly coupled SLAM framework to jointly optimize the edge features of contour ranging data and the SIFT features of visual feature data, construct a laser contour map as the underlying information and a visual feature map as the top-level information respectively, then perform the fusion of the laser contour map and the visual feature map, establish cross-modal feature associations between laser and vision through bidirectional projection, and obtain a composite map.
[0046] In this embodiment, for the laser edge point set P L and the visual corner point set P V , the following bidirectional projection constraints are imposed: (1) Laser → vision forward projection: Project the laser edge point set P L to the visual feature map through calibration parameters, calculate the pixel distance between it and the nearest neighbor visual feature point, and establish a geometric consistency constraint; (2) Vision → laser back projection: Back project the visual corner point set P V to the laser contour map plane, perform spatial alignment verification with adjacent laser points, and ensure the physical position consistency of cross-modal features in three-dimensional space.
[0047] The pixel deviation of the forward projection and the spatial offset of the back projection are weighted and fused, and for the matching points with residuals exceeding the threshold (such as pixel deviation > 15 pixels or spatial offset > 20 cm), their weights are reduced to nearly zero. For point pairs determined to be abnormal for more than 3 consecutive frames, they are temporarily removed from the matching pool to avoid cumulative errors. That is, it realizes the adaptive adjustment of the weights of the two types of residuals according to the environmental dynamics, such as increasing the laser constraint weight in scenes with drastic lighting changes and enhancing the visual constraint effect in areas where lasers fail such as glass curtain walls.
[0048] The map optimization subunit is used to obtain the laser constraint edges representing the connection of continuous robot pose nodes by calculating the geometric contour matching degree between the current contour ranging data and the laser contour map on the composite map, and obtain the visual constraint edges representing the connection of visual feature observation nodes by calculating the three-dimensional feature point reprojection error between the current visual feature data and the visual feature map, set an adaptive weight between the laser constraint edges and the visual constraint edges, and synchronously minimize the weighted sum of squared residuals of the laser constraint edges and the visual constraint edges to achieve the global consistency optimization of the composite map.
[0049] It should be understood that the laser constraint edge is a geometric consistency constraint that connects the continuous pose nodes of the robot and is generated by the matching degree between real-time laser ranging data and a pre-built laser profile map. Specifically, when the robot moves, the currently scanned contour point cloud (such as the edge of a wall or a column) is highly accurately aligned with the geometric structure in the laser map, and the overlap degree and shape similarity of the two contours are calculated. During the matching process, the relative motion constraints between pose nodes are extracted to form an edge connecting the poses at the previous and current moments. This constraint enforces that the movement amount between adjacent poses during the optimization process must conform to the geometric laws of laser observations, effectively suppressing the cumulative error of the odometer and ensuring the local geometric accuracy of the map.
[0050] The visual constraint edge is a cross-modal constraint that connects the robot pose node and the visual feature observation node and is constructed by the reprojection error of three-dimensional visual feature points. During the positioning process, the semantic features extracted from the current image (such as ceiling texture, fixed signs) are matched with the three-dimensional points in the visual feature map. The map points are projected onto the current camera view, and the pixel-level deviation between the projected position and the actual image detection position is calculated. This deviation is quantified as the reprojection error to form a constraint edge connecting the pose node and the feature node. This constraint enforces that the optimized pose must align the map feature points with the real-time observations in the projection space, thereby using visual semantic information to correct pose drift and enhancing the robustness to dynamic disturbances.
[0051] Furthermore, the pose prediction unit includes: A manual repositioning sub-unit, which is used to parse the initial position coordinates and their confidence radius input by the user, construct a three-degree-of-freedom search space, adopt a multi-resolution rasterization strategy, and refine the matching granularity layer by layer within the search space to find the coordinate with the best matching degree as the initial pose, and at the same time generate a confidence report containing the position covariance matrix. Specifically, an adaptive multi-resolution rasterization strategy is adopted, an initial grid size is set, the search space is divided into several coarse-grained grid cells, the matching scores between the center points of each grid and the laser map are calculated through a contour matching algorithm (such as NDT), the high-probability candidate areas with scores higher than the threshold are screened out, and the screened high-probability candidate areas are refined layer by layer to find the coordinate with the best matching degree.
[0052] An automatic repositioning sub-unit, which is used to synchronously perform laser contour matching and visual feature retrieval, generate a joint score by dynamically weighted fusion of the laser geometric similarity and the visual semantic matching degree, and obtain the best coordinate as the initial pose according to the comprehensive matching degree.
[0053] In a specific embodiment, in the automatic repositioning sub-unit, the NDT algorithm is used to perform laser contour matching to obtain the laser geometric similarity: ; where, Nm is the number of current matching points, N t is the total number of points, and RMSE is the root mean square error of registration. k 1 = 10 is the attenuation coefficient.
[0054] and retrieving similar scenes in the visual map through the BoW (Bag of Words) model
[0055] In the formula, M is the number of matching features, T f ( f i ) is the feature weight, cos ( d c , d m ) is the current feature descriptor d c and the cosine similarity with the map feature descriptor d m .
[0056] The combined score is: , where are the variances of the laser matching error and the visual retrieval error respectively.
[0057] The motion prediction subunit is used to fuse the data of the wheel odometer, the IMU angular velocity, and the laser odometer, and estimate the pose change amount using a kinematic model.
[0058] The kinematic model is: ; Among them, v L , v R are the left and right wheel speeds, L is the wheelbase, ω IMU is the IMU angular velocity, ωΔt is the angle change amount within the time period Δt .
[0059] The prediction optimization subunit is used to generate N candidate pose particles based on Monte Carlo sampling, perform a composite likelihood evaluation of the laser projection coverage rate and the visual reprojection error for each particle, update the candidate pose particles through importance resampling, and output the weighted average pose as the prediction result.
[0060] In the prediction optimization subunit, based on the motion prediction result, generate within the 3 σ confidence ellipseN = 200 candidate particles are then selected, followed by calculating the laser likelihood: the matching coverage rate of projecting the laser point cloud onto the laser map, and the visual likelihood: the average pixel error of reprojecting the visual feature points. Subsequently, the normalized weights are calculated, and the systematic resampling method is used to retain the particles with high weights and eliminate the particles with weights lower than the threshold (such as 1 / 2N). Finally, the weighted average pose is output as the prediction result.
[0061] Furthermore, the iterative execution unit includes: The first observation filtering and updating subunit is used to establish an adaptive search area in the laser contour map centered on the current predicted pose. Its search range dynamically adjusts the horizontal radius and the heading angle coverage interval according to the robot's movement speed. The multi-scale iterative closest point matching strategy is adopted to optimize the contour coincidence degree layer by layer between the coarse-grained and fine-grained resolutions. An exponential decay weight is applied to the contour ranging data points whose movement speed exceeds the preset speed threshold, and the laser matching residual is calculated based on the contour matching result.
[0062] Specifically, in the first observation filtering and updating subunit, the coarse-grained and fine-grained resolutions follow the following logic: Coarse-grained layer (resolution 0.2m): Quickly screen the candidate matching areas and eliminate the areas that deviate significantly; Medium-grained layer (resolution 0.1m): Optimize the point cloud alignment through the normal vector consistency constraint; Fine-grained layer (resolution 0.05m): Achieve sub-centimeter-level registration based on the curvature feature. It should be noted that an exponential decay weight is applied to the point cloud of dynamic obstacles whose movement speed exceeds the preset speed threshold to suppress the influence of dynamic interference on the matching result. The exponential decay weight is: , where t 1 is the duration of the obstacle's persistence, and the weight of the obstacle that persists for more than 2 seconds is reset to zero.
[0063] After that, the laser matching residual is calculated based on the contour matching result. The residual includes the root mean square error of the position residual and the heading angle deviation, and is output to the observation result correction subunit.
[0064] The second observation filtering and updating subunit is used to real-time extract the semantic features in the current visual data through the lightweight semantic network running on the embedded NPU, construct a three-dimensional projection space in the neighborhood of the predicted pose, reproject the three-dimensional feature points in the visual feature map to the current image coordinate system and perform two-way matching verification, and calculate the visual reprojection error based on the valid matching pairs obtained from the verification.
[0065] In the second observation filtering update sub-unit, the lightweight YOLO-Mini network (model size <8MB) is run through the embedded NPU to extract semantic features in real time. It mainly extracts preset categories, including navigation markers such as safety exit signs, QR code tags, and reflective positioning balls, as well as environmental structure features such as ceiling beams, wall distribution boxes, and floor guiding lines. Subsequently, 2D-3D feature point pairs with semantic labels are output, and there is a transmission delay. A three-dimensional projection space is constructed within the predicted pose neighborhood. Specifically, a three-dimensional projection pyramid is constructed, where the number of projection layers is 3, the heights correspond to 0.5m / 1.5m / 2.5m respectively, and the resolution is 0.1m / layer to cover the robot's motion space.
[0066] In addition, a bidirectional matching verification mechanism is adopted. On the one hand, forward projection projects the three-dimensional points in the visual feature map onto the current image plane and calculates the pixel-level deviation; on the other hand, reverse matching back-projects the 2D features extracted from the current image into the map space to verify the three-dimensional consistency. The valid matching pairs need to meet the following conditions: the bidirectional projection errors are both <3 pixels, the semantic label consistency reaches 100%, and the feature scale change rate <20% (resistant to view changes).
[0067] Furthermore, the following formula is used for the visual reprojection error; ; In the formula, m is the total number of three-dimensional feature points for matching the current frame with the map, δ i is the value of the binary indicator function taking {0,1}, which is 1 for valid matching pairs δ i =1, O i is the semantic confidence weighting, the classification probability of the feature points output by the lightweight semantic network, is the projection coordinate, T pred is the predicted pose matrix for the current cycle, that is, the transformation relationship from the map coordinate system to the current camera coordinate system, the coordinate of the i th feature point in the visual feature map, is the two-dimensional pixel coordinate of the i th feature point actually detected in the current image.
[0068] The observation result correction subunit is used to construct a multi-modal Kalman filter, which respectively receives the laser matching residual and the visual reprojection error, dynamically allocates sensor weights according to the environmental degradation factor, and the environmental degradation factor includes the proportion of dynamic obstacles in the laser and the visual illumination uniformity. When the bimodal positioning deviation exceeds the safety threshold, the arbitration mechanism is started, the historical confidence data is traced back, and the sensor data stream with a confidence level continuously higher than 80% in the past N seconds is preferentially adopted. The confidence level is calculated through the condition number of the residual distribution covariance matrix, and the multi-modal Kalman filter update is executed to output the corrected pose and covariance matrix as the prediction input for the next cycle.
[0069] In summary, the present invention provides a device for visual and 2D laser fusion positioning, which constructs a high-precision and high-robustness positioning solution through the deep integration of a heterogeneous sensor architecture and a computing platform. Specifically, the device includes three core modules: a horizontal ranging module, a vertical vision module, and a computing and processing module 4, and the three form a collaborative innovation in terms of spatial layout and data processing.
[0070] First, in the hardware architecture design, the horizontal ranging module uses a 2D lidar, maintains the strict parallelism between the laser scanning plane and the robot's motion trajectory through a dynamic pitch compensation mechanism, and obtains millimeter-level horizontal contour data in real time; the vertical vision module is equipped with an infrared sensor and an active lighting array, installed at a right angle, and combines an infrared band-pass filter to eliminate ambient light interference, focusing on collecting vertical space semantic features such as ceiling structures and hanging markers. This orthogonal layout not only physically isolates the signal crosstalk between sensors, but also lays a geometric foundation for multi-modal data fusion through coordinate system mapping.
[0071] On this basis, the computing and processing module 4 realizes the collaborative innovation of software and hardware. The multi-core heterogeneous computing platform (CPU+NPU / GPU) improves the positioning performance through the following technical breakthroughs: Based on the multi-core heterogeneous computing platform (CPU+NPU / GPU), execute: construct a laser contour map and a visual feature map with coordinate system alignment based on the contour ranging data and the visual feature data, determine the initial pose of the robot according to the given position input or the optimal position of the self-global matching laser contour map and the visual feature map, generate a predicted pose in combination with the motion measurement value, and iteratively execute the dual-modal observation filtering update of laser contour matching optimization and visual reprojection matching optimization centered on the predicted pose. Thus, the present invention constructs an all-weather positioning solution for industrial scenarios through the deep integration of a heterogeneous sensor architecture and a computing platform.
[0072] Particularly critically, this device solves the industry pain points through deep integration of software and hardware: reducing the communication delay from sensor data acquisition to the output of the final result; ensuring strict spatio-temporal alignment of laser scanning and image exposure with the help of a hardware trigger signal to eliminate the matching error caused by motion blur; and the NPU reuse design enables the sharing of computing resources for visual feature extraction and laser point cloud processing, reducing the overall hardware cost.
[0073] Thus, through orthogonal sensor layout, multimodal data fusion algorithm, and software and hardware collaborative optimization, the present invention constructs the first set of "all-weather - all-terrain - all-dynamic" positioning system in the field of industrial mobile robots, providing a reliable autonomous navigation infrastructure for intelligent manufacturing scenarios.
[0074] Since the system / device described in the above embodiments of the present invention is the system / device adopted for implementing the method in the above embodiments of the present invention, based on the method described in the above embodiments of the present invention, those skilled in the art can understand the specific structure and variations of the system / device, and thus will not be elaborated herein. Any system / device adopted by the method in the above embodiments of the present invention falls within the scope of protection of the present invention.
[0075] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0076] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions.
[0077] It should be noted that the words "a" or "an" preceding a component do not exclude the existence of a plurality of such components. The present invention can be implemented by means of hardware including several different components and by means of a properly programmed computer. Among the several devices listed, several of these devices can be embodied by the same hardware. The use of the words first, second, third, etc. is only for convenience of expression and does not represent any order. These words can be understood as part of the component name.
[0078] In addition, it should be noted that in the description of this specification, the descriptions of terms such as "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0079] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications after learning the basic creative concepts. Therefore, the technical solutions should be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0080] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the technical solutions of the present invention and their equivalent technologies, the present invention should also include these modifications and variations.
Claims
1. A device based on visual and 2D laser fusion positioning, characterized in that, Comprising: A horizontal ranging module, horizontally installed on the robot body, for collecting contour ranging data of the motion plane; A vertical vision module, top-view installed on the robot body, for collecting visual feature data of the vertical space; A calculation and processing module, configured to: construct a laser contour map and a visual feature map with coordinate system alignment based on the contour ranging data and the visual feature data, determine the initial pose of the robot according to the given position input or the optimal position of globally matching the laser contour map and the visual feature map, generate a predicted pose in combination with the motion measurement value, and iteratively execute dual-modal observation filtering update of laser contour matching optimization and visual reprojection matching optimization centered on the predicted pose.
2. The device for visual and 2D laser fusion positioning according to claim 1, wherein The horizontal ranging module includes a 2D lidar, which is adjusted by a pre-configured dynamic pitch compensation mechanism, and the parallelism error between the laser scanning plane of the 2D lidar and the traveling direction of the robot is less than the preset parallelism error; The vertical vision module includes an infrared sensor, which is adjusted by a pre-configured tunable gimbal, and the optical axis of the lens of the infrared sensor forms an included angle of 88° - 92° with the 2D lidar, so that the vertical field of view coverage range of the infrared sensor and the horizontal scanning sector of the 2D lidar form a non-overlapping complementary observation domain in the space coordinate system.
3. The device for visual and 2D laser fusion positioning according to claim 2, wherein The complementary observation domain satisfies the following spatial constraint conditions: The 2D lidar forms a scanning sector with adaptive opening angle adjustment θ ∈ [0°, 270°] in the X-Y plane of the robot motion coordinate system; The infrared sensor constructs a three-dimensional pyramid-shaped detection area in the Z-axis direction of the Cartesian coordinate system, and the pitch angle φ of the infrared sensor ∈ [55°, 95°] and forms a buffer isolation zone of not less than 15% with the scanning sector of the lidar on the spatial projection plane.
4. The device for visual and 2D laser fusion positioning according to claim 1, wherein, An active infrared supplementary light source device is configured, and the active infrared supplementary light source device includes: An annular distribution unit, arranging N independently addressable infrared emission sources circumferentially around the infrared sensor; A spectral modulation unit, dynamically selecting a working wavelength of 850nm or 940nm according to the ambient light entropy value, and driving a tunable band-pass filter for matching spectral filtering; A light emission synchronization controller, aligning the pulse emission period of each infrared emission source with the global shutter exposure window of the infrared sensor through a hardware trigger signal; A light source selection module, including a switchable LED light source group and a VCSEL light source group. Among them, the LED light source group includes multiple LEDs arranged in an annular array, and the light emitting surfaces of all LEDs are installed obliquely outward with the central normal as the reference, and the inclination angle range is 5° - 15°, and the inclination directions of the LEDs in the same layer are symmetrically distributed; the VCSEL light source group includes multiple vertical cavity surface emitting lasers.
5. The device for visual and 2D laser fusion positioning according to any one of claims 1-4, characterized in that, The calculation and processing module includes: A registration and mapping unit, for constructing a laser contour map and a visual feature map with coordinate system alignment based on the contour ranging data and the visual feature data; A pose prediction unit, for determining the initial pose of the robot according to the given position input or the optimal position of globally matching the laser contour map and the visual feature map, and generating a predicted pose in combination with the motion measurement value; An iterative execution unit for iteratively executing with the predicted pose as the center: performing optimization calculation of the contour coincidence degree between the current contour ranging data and the laser contour map with the current predicted pose as the center, and implementing the first observation filtering update; extracting the semantic features of the current visual feature data, and using the semantic features to perform feature point reprojection matching optimization of the visual feature map with the current predicted pose as the center, and implementing the second observation filtering update; performing positioning correction on the results of the first observation filtering update and the second observation filtering update, and using the corrected pose as the predicted input for the next cycle.
6. The device for visual and 2D laser fusion positioning according to claim 5, characterized in that The registration and mapping unit includes: A dynamic environment compensation sub-unit for constructing a motion distortion compensation field based on the integration of the angular velocity / linear velocity of the IMU and the odometer, performing point-by-point motion distortion removal processing on the contour ranging data, performing dynamic target detection based on density clustering on the de-distorted contour ranging data, removing data with a moving speed exceeding the preset speed, and dynamically adjusting the extraction sensitivity of the visual feature data according to the environmental light intensity, and enabling adaptive histogram equalization preprocessing for the visual feature data when the illuminance is less than the preset illuminance threshold; A dynamic joint calibration sub-unit for calculating the external parameter residual matrix between the lidar coordinate system and the visual sensor coordinate system according to the synchronously collected contour ranging data and visual feature data, and triggering calibration when the external parameter residual exceeds the preset residual value or the pose estimation covariance value of a continuous preset number of frames is greater than the preset covariance value. Performing online calibration work by implementing bimodal feature extraction on the reflectivity mutation edge of the contour ranging data and the temperature difference feature corner of the visual feature data; A multi-modal map construction sub-unit for adopting a tightly coupled SLAM framework to jointly optimize the edge features of the contour ranging data and the SIFT features of the visual feature data, respectively constructing a laser contour map as the underlying information and a visual feature map as the top-level information, and then performing fusion of the laser contour map and the visual feature map, establishing cross-modal feature associations between the laser and the vision through bidirectional projection, and obtaining a composite map; A map optimization sub-unit for obtaining a laser constraint edge representing the connection of continuous robot pose nodes by calculating the geometric contour matching degree between the current contour ranging data and the laser contour map on the composite map, and obtaining a visual constraint edge representing the connection of visual feature observation nodes by calculating the three-dimensional feature point reprojection error between the current visual feature data and the visual feature map, setting an adaptive weight between the laser constraint edge and the visual constraint edge, and synchronously minimizing the weighted residual sum of squares of the laser constraint edge and the visual constraint edge to achieve global consistency optimization of the composite map.
7. The device for visual and 2D laser fusion positioning according to claim 6, wherein The online calibration work uses a composite calibration target, and the composite calibration target includes: Laser high-reflection stripes, a grid pattern composed of alternating 20% reflectivity and 80% reflectivity; Infrared feature markers, AprilTag tags embedded in the temperature difference module, presenting controllable temperature difference features under thermal imaging.
8. The device for visual and 2D laser fusion positioning according to claim 5, characterized in that, The pose prediction unit includes: A manual repositioning subunit, which is used to parse the initial position coordinates and their confidence radii input by the user, construct a three-degree-of-freedom search space, adopt a multi-resolution rasterization strategy, refine the matching granularity layer by layer within the search space to find the coordinates with the best matching degree as the initial pose, and generate a confidence report including a position covariance matrix at the same time; An automatic repositioning subunit, which is used to synchronously execute laser contour matching and visual feature retrieval, generate a joint score by dynamically weighted fusion of laser geometric similarity and visual semantic matching degree, and obtain the best coordinates as the initial pose according to the comprehensive matching degree; A motion prediction subunit, which is used to fuse wheel odometer, IMU angular velocity and laser odometer data, and calculate the pose change amount by using a kinematic model; A prediction optimization subunit, which is used to generate multiple candidate pose particles based on Monte Carlo sampling, perform a composite likelihood evaluation of laser projection coverage and visual reprojection error for each particle, update the candidate pose particles by importance resampling, and output the weighted average pose as the prediction result.
9. The device for visual and 2D laser fusion positioning according to claim 5, wherein The iterative execution unit includes: A first observation filtering and updating subunit, which is used to establish an adaptive search area in the laser contour map with the current predicted pose as the center, adopt a multi-scale iterative closest point matching strategy, optimize the contour coincidence degree layer by layer between coarse-grained and fine-grained resolutions, apply an exponential decay weight to the contour ranging data points whose moving speed exceeds the preset speed threshold, and calculate the laser matching residual based on the contour matching result; A second observation filtering and updating subunit, which is used to real-time extract the semantic features in the current visual data through a lightweight semantic network running on an embedded NPU, construct a three-dimensional projection space in the neighborhood of the predicted pose, reproject the three-dimensional feature points in the visual feature map to the current image coordinate system and perform bidirectional matching verification, and calculate the visual reprojection error based on the valid matching pairs obtained by the verification; An observation result correction subunit, which is used to construct a multi-modal Kalman filter, receive the laser matching residual and the visual reprojection error respectively, dynamically allocate sensor weights according to the environmental degradation factor, the environmental degradation factor includes the proportion of laser dynamic obstacles and the visual illumination uniformity, when the bimodal positioning deviation exceeds the safety threshold, start an arbitration mechanism, preferentially adopt the sensor data stream whose confidence has been continuously higher than the preset confidence threshold in the past preset time period, and output the corrected pose as the prediction input for the next cycle.
Citation Information
Patent Citations
Repositioning method for mobile robot
CN105652871A
Robot indoor mapping method and system based on vision and laser slam
CN111076733A
Laser and vision fused integrated global positioning method for indoor robot
CN112596064A
SLAM method based on tight coupling of 2D laser radar and binocular camera
CN112785702A
Positioning method and system
CN114184193A
Cited By
Vehicle-mounted laser radar point cloud real-time target detection method and system
CN120630151A
Large-scale intelligent driving beam transporting vehicle obstacle avoidance control method and system based on accurate positioning
CN120972981A
Unmanned aerial vehicle automatic homeward voyage method and system based on inertial navigation and visual system
CN121115818A
Construction site-oriented multi-modal adaptive weight SLAM method and system
CN121383989A
Indoor narrow space large-scale equipment self-adaptive installation method based on multi-source information fusion
CN121594843A