A device based on vision and 2D laser fusion positioning

Through the visual and 2D laser fusion positioning device, the complementary layout of the horizontal ranging module and the vertical vision module and the dual-mode observation filter update of the calculation and processing module are solved, and the ease of failure of laser positioning in repeated characteristics and high dynamic environments is achieved, achieving all-weather and highly reliable industrial scene positioning.

CN120313613BActive Publication Date: 2025-08-29HANGZHOU LANXIN TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510808472.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-29
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

In the prior art, laser positioning is prone to failure under repeated features and high dynamic environments, visual positioning is significantly disturbed by light changes and occlusion, and a single modal positioning scheme is difficult to achieve full operating conditions coverage in industrial scenarios.

Method used

Using a device based on vision and 2D laser fusion positioning, through the complementary layout of the horizontal ranging module and the vertical vision module, the dual-modal observation filtering update is performed in combination with the calculation processing module, and a coordinate system-aligned laser profile map and visual feature map are constructed to realize multimodal data fusion and enhance the robustness of the system.

Benefits of technology

While ensuring the millimeter-level positioning accuracy, it significantly improves the robustness of industrial interference factors such as light changes, dynamic obstacles, and structural repetition, providing all-weather and high-reliable positioning guarantee.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120313613B_ABST
    Figure CN120313613B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of positioning technology, and in particular to a device based on vision and 2D laser fusion positioning, the device comprising: a horizontal ranging module, mounted horizontally on the robot body, for collecting contour ranging data of the motion plane; a vertical vision module, mounted top-view on the robot body, for collecting visual feature data in the vertical space; and a computing and processing module, configured to: construct a laser contour map and a visual feature map aligned with the coordinate system based on the contour ranging data and the visual feature data; determine the robot's initial pose based on a given position or the optimal position of the laser contour map and the visual feature map from a global match; generate a predicted pose based on motion measurement values; and iteratively perform dual-modal observation filtering updates of laser contour matching optimization and visual reprojection matching optimization centered on the predicted pose. The present invention constructs an all-weather positioning solution for industrial scenarios through the deep integration of heterogeneous sensor architecture and computing platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of positioning technology, and in particular to a device based on vision and 2D laser fusion positioning. Background Art

[0002] Currently, in the field of mobile robotics, similar to the debate between lidar and pure vision in autonomous driving, natural navigation and positioning technology for AMR / AGVs in industrial scenarios (i.e., positioning technology that does not use artificial markers such as QR codes and reflectors) also faces a technical distinction between vision and laser. Laser positioning, due to its precise distance measurement information, is less challenging and significantly more mature than visual positioning. Laser positioning offers the stability and accuracy required for industrial applications in 80%-90% of indoor and semi-indoor environments. Consequently, domestic AMR manufacturers primarily focus on laser SLAM technology. However, laser positioning carries the risk of loss in environments with repetitive features (such as long silhouettes) and high dynamic range. In contrast, visual positioning is an emerging positioning technology. While cameras can capture richer environmental information, such as color and texture, than laser positioning, which helps robots more accurately identify and understand their environment, visual positioning faces challenges in complex environments, such as varying lighting and obstructions, and its engineering application remains a significant challenge.

[0003] Therefore, how to effectively integrate the advantages of the two perception modalities and break through the performance limitations of single-modality sensors under specific working conditions has become a key technical challenge to improve the full-scene positioning capabilities of industrial mobile robots. Summary of the Invention

[0004] (1) Technical issues to be resolved

[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a device based on vision and 2D laser fusion positioning, which solves the technical problems that the existing laser positioning is prone to failure in repetitive features and high-dynamic environments, visual positioning is significantly affected by lighting changes and occlusions, and a single modality positioning solution is difficult to achieve full working condition coverage of industrial scenarios.

[0006] (2) Technical solution

[0007] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:

[0008] In a first aspect, an embodiment of the present invention provides a device based on vision and 2D laser fusion positioning, comprising:

[0009] The horizontal ranging module is installed horizontally on the robot body and is used to collect the contour ranging data of the motion plane;

[0010] Vertical vision module, top view installed on the robot body, used to collect visual feature data in the vertical space;

[0011] The computing and processing module is configured to: construct a laser contour map and a visual feature map with aligned coordinate systems based on the contour ranging data and the visual feature data; determine the initial position of the robot according to the input given position or the optimal position of the laser contour map and the visual feature map from the global matching; generate a predicted position in combination with the motion measurement value; and iteratively perform a dual-modal observation filtering update of the laser contour matching optimization and the visual reprojection matching optimization with the predicted position as the center.

[0012] Optionally, the horizontal ranging module includes a 2D laser radar, which is adjusted by a pre-configured dynamic pitch compensation mechanism so that the parallelism error between the laser scanning plane of the 2D laser radar and the direction of travel of the robot is maintained less than a preset parallelism error;

[0013] The vertical vision module includes an infrared sensor. Through the pre-configured tunable universal joint adjustment, the optical axis of the infrared sensor lens forms an 88°-92° angle configuration with the 2D lidar, so that the vertical field of view angle coverage of the infrared sensor and the horizontal scanning sector of the 2D lidar form a non-overlapping complementary observation domain in the spatial coordinate system.

[0014] Optionally, the complementary observation domain satisfies the following spatial constraints:

[0015] The 2D laser radar forms a scanning sector θ∈[0°,270°] with adaptive opening angle adjustment in the XY plane of the robot motion coordinate system;

[0016] The infrared sensor constructs a three-dimensional pyramid-shaped detection area in the Z-axis direction of the Cartesian coordinate system. The pitch angle φ∈[55°,95°] of the infrared sensor forms a buffer isolation zone of no less than 15% with the scanning sector of the lidar on the spatial projection surface.

[0017] Optionally, an active infrared fill light source device is configured, and the active infrared fill light source device includes:

[0018] A ring-shaped distribution unit, which arranges N independently addressable infrared emission sources around the infrared sensor;

[0019] The spectrum modulation unit dynamically selects the operating wavelength of 850nm or 940nm according to the ambient light entropy value, and drives the tunable bandpass filter to perform matched spectrum filtering;

[0020] The light emitting synchronization controller uses hardware trigger signals to align the pulse emission periods of each infrared emission source with the global shutter exposure window of the infrared sensor;

[0021] The light source selection module includes a switchable LED light source group and a VCSEL light source group. The LED light source group includes multiple LEDs arranged in a circular array, and the light-emitting surfaces of all LEDs are installed tilted outward with respect to the center normal. The tilt angle range is 5°-15°, and the tilt direction of the LEDs on the same layer is symmetrically distributed; the VCSEL light source group includes multiple vertical cavity surface emitting lasers.

[0022] Optionally, the calculation processing module includes:

[0023] A registration and mapping unit is used to construct a laser contour map and a visual feature map with aligned coordinate systems based on the contour ranging data and the visual feature data;

[0024] The pose prediction unit is used to determine the initial pose of the robot based on the input given position or the optimal position from the global matching laser contour map and visual feature map, and generate the predicted pose in combination with the motion measurement value;

[0025] The iterative execution unit is used to iteratively execute with the predicted posture as the center: optimize the calculation of the contour coincidence between the current contour ranging data and the laser contour map with the current predicted posture as the center, and implement the first observation filter update; extract the semantic features of the current visual feature data, use the semantic features to optimize the feature point reprojection matching of the visual feature map with the current predicted posture as the center, and implement the second observation filter update; perform positioning correction on the results of the first observation filter update and the second observation filter update, and use the corrected posture as the prediction input for the next cycle.

[0026] Optionally, the registration and mapping unit includes:

[0027] The dynamic environment compensation subunit is used to construct a motion distortion compensation field based on the angular velocity / linear velocity integration of the IMU and the odometer, perform point-by-point motion dedistortion processing on the contour ranging data, perform density clustering-based dynamic target detection on the dedistorted contour ranging data, remove data moving faster than a preset speed, and dynamically adjust the extraction sensitivity of the visual feature data according to the ambient light intensity. When the illumination is less than the preset illumination threshold, adaptive histogram equalization preprocessing is enabled for the visual feature data;

[0028] The dynamic joint calibration subunit is used to calculate the extrinsic parameter residual matrix of the lidar coordinate system and the visual sensor coordinate system based on the synchronously collected contour ranging data and visual feature data. When the extrinsic parameter residual exceeds the preset residual value or the covariance value of the pose estimation of a preset number of consecutive frames is greater than the preset covariance value, the calibration is triggered. The dual-modal feature extraction is performed on the reflectivity mutation edge of the contour ranging data and the temperature difference feature corner point of the visual feature data to perform online calibration.

[0029] The multimodal map construction subunit is used to jointly optimize the edge features of the contour ranging data and the SIFT features of the visual feature data using a tightly coupled SLAM framework, respectively constructing a laser contour map as the bottom layer of information and a visual feature map as the top layer of information. The laser contour map and the visual feature map are then fused, and a cross-modal feature association between laser and vision is established through bidirectional projection to obtain a composite map.

[0030] The map optimization subunit is used to obtain the laser constraint edges that represent the connection between continuous robot pose nodes on the composite map by calculating the geometric contour matching degree between the current contour ranging data and the laser contour map, and to obtain the visual constraint edges that represent the connection between visual feature observation nodes by calculating the three-dimensional feature point reprojection error between the current visual feature data and the visual feature map. Adaptive weights are set between the laser constraint edges and the visual constraint edges, and the weighted residual sum of squares of the laser constraint edges and the visual constraint edges are simultaneously minimized to achieve global consistency optimization of the composite map.

[0031] Optionally, online calibration uses a composite calibration target, which includes:

[0032] Laser high-reflectivity stripes, a grid pattern consisting of alternating 20% ​​and 80% reflectivity;

[0033] Infrared feature marker, AprilTag tag embedded with temperature differential module, presents controllable temperature difference characteristics under thermal imaging.

[0034] Optionally, the pose prediction unit includes:

[0035] The manual repositioning subunit is used to parse the initial position coordinates and their confidence radius input by the user, construct a three-degree-of-freedom search space, and adopt a multi-resolution rasterization strategy to gradually refine the matching granularity within the search space to find the coordinates with the best matching degree as the initial pose. At the same time, it generates a confidence report including the position covariance matrix;

[0036] The automatic repositioning subunit is used to synchronously perform laser contour matching and visual feature retrieval. It generates a joint score by dynamically weighting the fusion of laser geometric similarity and visual semantic matching, and obtains the optimal coordinates as the initial pose based on the comprehensive matching degree.

[0037] The motion prediction subunit is used to integrate the wheel odometer, IMU angular velocity and laser odometer data, and use the kinematic model to calculate the pose change;

[0038] The prediction optimization subunit is used to generate multiple candidate pose particles based on Monte Carlo sampling, perform a composite likelihood evaluation of laser projection coverage and visual reprojection error on each particle, update the candidate pose particles through importance resampling, and output the weighted average pose as the prediction result.

[0039] Optionally, the iterative execution unit includes:

[0040] The first observation filter update subunit is used to establish an adaptive search area in the laser contour map with the current predicted pose as the center, adopt a multi-scale iterative closest point matching strategy, optimize the contour coincidence layer by layer between coarse-grained and fine-grained resolutions, apply exponential decay weights to contour ranging data points whose moving speed exceeds a preset speed threshold, and calculate the laser matching residual based on the contour matching results;

[0041] The second observation filter update subunit is used to extract semantic features from the current visual data in real time through a lightweight semantic network running on the embedded NPU, construct a three-dimensional projection space within the predicted pose neighborhood, reproject the three-dimensional feature points in the visual feature map to the current image coordinate system and perform two-way matching verification, and calculate the visual reprojection error based on the verified valid matching pairs;

[0042] The observation result correction subunit is used to construct a multimodal Kalman filter, which receives the laser matching residual and visual reprojection error respectively, and dynamically allocates sensor weights according to the environmental degradation factors. The environmental degradation factors include the proportion of laser dynamic obstacles and the uniformity of visual illumination. When the dual-modal positioning deviation exceeds the safety threshold, the arbitration mechanism is activated, and the sensor data stream whose built-in confidence has been continuously higher than the preset confidence threshold in the past preset period is given priority, and the corrected posture is output as the prediction input for the next cycle.

[0043] (3) Beneficial effects

[0044] The beneficial effect of the present invention is that the present invention constructs an all-weather positioning solution for industrial scenarios through the deep integration of heterogeneous sensor architecture and computing platform.

[0045] Specifically, the complementary spatial layout of the horizontal ranging module and the vertical vision module can simultaneously capture precise geometric information about the moving plane and rich semantic features in the vertical direction, effectively overcoming the perception limitations of single horizontal laser scanning in complex stereoscopic scenes. Secondly, the deeply integrated hardware and software architecture significantly enhances the spatiotemporal consistency of multi-source data. This not only eliminates fusion errors caused by communication delays in traditional split-type solutions, but also achieves submillimeter alignment of laser point clouds and visual features through a unified spatiotemporal reference, laying the physical foundation for multimodal data fusion.

[0046] Furthermore, the dual-modal observation filter update mechanism employed by this invention innovatively integrates the geometric constraints of laser contour matching with the semantic constraints of visual reprojection, creating dual anti-interference capabilities in dynamic interference environments: when the environment experiences temporary occlusion or sudden changes in illumination, laser observations provide a stable geometric baseline; while in repetitive structural scenes such as long corridors, visual features inject discriminative semantic information, forming a dynamic and complementary fault-tolerant mechanism. Furthermore, the composite map construction method with strictly aligned coordinate systems enables the system to concurrently utilize the precise geometric priors of lasers and the topological recognition capabilities of vision during the global positioning phase. Even in the event of a device cold start or severe position loss, a cross-modal joint search can rapidly restore a stable position.

[0047] Therefore, the present invention successfully solves the core pain points of modal characteristic conflicts and single environmental adaptability in traditional solutions through hardware layout optimization and algorithm architecture innovation. While ensuring millimeter-level positioning accuracy, it significantly improves the robustness to typical industrial interference factors such as lighting changes, dynamic obstacles, and structural repetition, providing all-weather high-reliability positioning guarantees for scenarios such as unmanned transportation. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A schematic diagram of the composition of a device based on vision and 2D laser fusion positioning provided by an embodiment of the present invention;

[0049] Figure 2 A schematic diagram of the internal flow of a computing and processing module of a device based on vision and 2D laser fusion positioning provided in an embodiment of the present invention.

[0050] [Description of Reference Numerals]

[0051] 1: Horizontal ranging module; 2: Vertical vision module; 3: Active infrared fill light source device; 4: Computation and processing module; 5: Communication interface module; 6: Power supply module. DETAILED DESCRIPTION

[0052] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.

[0053] like Figure 1As shown, an embodiment of the present invention proposes a device based on vision and 2D laser fusion positioning, including: a horizontal ranging module, horizontally installed on the robot body, for collecting contour ranging data of the motion plane; a vertical vision module, top-view installed on the robot body, for collecting visual feature data in the vertical space; a calculation and processing module, configured to: construct a laser contour map and a visual feature map aligned with the coordinate system based on the contour ranging data and the visual feature data, determine the initial posture of the robot according to an input given position or the optimal position of the laser contour map and the visual feature map from the global matching, generate a predicted posture in combination with the motion measurement value, and iteratively perform dual-modal observation filtering update of laser contour matching optimization and visual reprojection matching optimization with the predicted posture as the center.

[0054] This invention builds an all-weather positioning solution for industrial scenarios through deep integration of heterogeneous sensor architecture and computing platform.

[0055] Specifically, the complementary spatial layout of the horizontal ranging module and the vertical vision module can simultaneously capture precise geometric information about the moving plane and rich semantic features in the vertical direction, effectively overcoming the perception limitations of single horizontal laser scanning in complex stereoscopic scenes. Secondly, the deeply integrated hardware and software architecture significantly enhances the spatiotemporal consistency of multi-source data. This not only eliminates fusion errors caused by communication delays in traditional split-type solutions, but also achieves submillimeter alignment of laser point clouds and visual features through a unified spatiotemporal reference, laying the physical foundation for multimodal data fusion.

[0056] Furthermore, the dual-modal observation filter update mechanism employed by this invention innovatively integrates the geometric constraints of laser contour matching with the semantic constraints of visual reprojection, creating dual anti-interference capabilities in dynamic interference environments: when the environment experiences temporary occlusion or sudden changes in illumination, laser observations provide a stable geometric baseline; while in repetitive structural scenes such as long corridors, visual features inject discriminative semantic information, forming a dynamic and complementary fault-tolerant mechanism. Furthermore, the composite map construction method with strictly aligned coordinate systems enables the system to concurrently utilize the precise geometric priors of lasers and the topological recognition capabilities of vision during the global positioning phase. Even in the event of a device cold start or severe position loss, a cross-modal joint search can rapidly restore a stable position.

[0057] Therefore, the present invention successfully solves the core pain points of modal characteristic conflicts and single environmental adaptability in traditional solutions through hardware layout optimization and algorithm architecture innovation. While ensuring millimeter-level positioning accuracy, it significantly improves the robustness to typical industrial interference factors such as lighting changes, dynamic obstacles, and structural repetition, providing all-weather high-reliability positioning guarantees for scenarios such as unmanned transportation.

[0058] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0059] Specifically, the horizontal ranging module 1 includes a 2D laser radar, which is adjusted by a pre-configured dynamic pitch compensation mechanism so that the parallelism error between the 2D laser radar laser scanning plane and the robot's travel direction is less than the set parallelism error. Figure 1 The 2D laser radar is installed horizontally, and the dynamic pitch compensation mechanism is composed of a gyroscope sensor and a servo motor. It detects the changes in the robot's movement posture in real time and drives the 2D laser radar rotation axis to control the parallelism error between the laser scanning plane and the robot's motion coordinate system X-axis (travel direction) within the range of ±0.5°.

[0060] The vertical vision module 2 includes an infrared sensor, which is adjusted through a pre-configured tunable gimbal. The tunable gimbal has a built-in angle encoder and locking mechanism, which supports the infrared sensor's lens optical axis to form an 88°-92° angle configuration with the 2D lidar. This ensures that the infrared sensor's vertical field of view angle coverage and the 2D lidar's horizontal scanning sector form a non-overlapping complementary observation domain in the spatial coordinate system.

[0061] In particular, the complementary observation domain satisfies the following spatial constraints: the 2D lidar forms a scanning sector with adaptive opening angle adjustment in the XY plane of the robot motion coordinate system i ∈[0°,270°], and the horizontal scanning sector θ∈[0°,270°] of the 2D laser radar is dynamically adjusted by the adaptive opening angle algorithm: when a narrow corridor-like environment is detected, it automatically shrinks to i =120° to reduce redundant data; in open areas, it extends to i =270°, achieving wide-area detection. The infrared sensor constructs a three-dimensional pyramidal detection zone along the Z-axis of the Cartesian coordinate system. The infrared sensor's elevation angle φ∈[55°,95°] forms a buffer zone of at least 15% with the lidar's scanning sector on the spatial projection plane. It is important to emphasize that the buffer zone parameters of the orthogonal complementary detection domain are designed to mimic the distribution of rods in insect compound eyes, achieving optimal spatial coverage efficiency with limited sensor resources.

[0062] Next, the device of the present invention is equipped with an active infrared fill light source device 3, which includes:

[0063] Ring distribution unit, arranged circumferentially around the infrared sensor N independently addressable infrared emission sources; among which, N ∈[2,8] and satisfies N = ceiling ( H max / ΔH ), ceil is rounded up, H max is the extreme vertical height of the scene, ΔH is the preset height resolution threshold.

[0064] The spectral modulation unit dynamically selects an operating wavelength of 850nm or 940nm according to the ambient light entropy value, and drives a tunable bandpass filter for matched spectral filtering. When the short-wave infrared detector is activated, the 940nm wavelength is preferentially used to avoid visible light interference. When the long-wave infrared detector is working, it automatically switches to an 850nm wavelength to enhance the thermal contrast detection sensitivity.

[0065] The light emitting synchronization controller uses hardware trigger signals to align the pulse light emitting periods of each infrared emission source with the global shutter exposure window of the infrared sensor.

[0066] The light source selection module includes switchable LED and VCSEL light source groups. The LED light source group consists of multiple LEDs arranged in a circular array, with the light-emitting surfaces of all LEDs tilted outward from the center normal. The tilt angle range is 5°-15°, and the LEDs on the same layer are symmetrically distributed. The VCSEL light source group includes multiple vertical cavity surface emitting lasers. The LED light source group is suitable for scenes with diffuse reflection, with a divergence angle of ≥120°. The VCSEL light source group is suitable for scenes with specular reflection suppression, with a collimation angle of ≤10° and a power density of >200mW / cm².

[0067] It should be noted that the device of the present invention is also equipped with a communication interface module 5 and a power supply module 6. By default, the device uses a Gigabit Ethernet communication interface, through which the positioning posture is transmitted to the robot body. CAN communication or other communication methods can also be used to interact with navigation systems, etc., depending on the usage requirements. The power supply module 6 is responsible for stabilizing the voltage and supplying power to the processing unit, lidar module, infrared active fill light device, image sensor, etc. Because the infrared automatic fill light device is instantaneous, the current changes greatly instantly, so isolation between the modules is required.

[0068] Furthermore, if Figure 2 As shown, the calculation processing module 4 includes:

[0069] The registration and mapping unit is used to construct a laser contour map and a visual feature map with aligned coordinate systems based on the contour ranging data and the visual feature data.

[0070] In one embodiment, mapping mode is activated, and a mobile robot equipped with the device is controlled to move along a path within the operating environment. While the mobile robot is controlled to move along the preset path, horizontal profile ranging data from a 2D lidar and vertical visual feature data from an infrared sensor are synchronously collected at a 10Hz frequency. Hardware timestamping is used to achieve spatiotemporal alignment of the dual-modal data. The resulting laser profile map represents the contour map of the plane scanned by the lidar, while the visual feature map represents the visual features of the upper portion of the environment. The visual feature point cloud is projected onto the laser profile map coordinate system using an extrinsic calibration matrix. The origin (the initial pose of the mobile robot) and X / Y axis directions (aligned with the lidar scanning plane) of the two maps are forcibly aligned to achieve consistent spatial mapping. Once mapping is complete, the map is loaded into the device, and positioning mode is entered.

[0071] The pose prediction unit is used to determine the robot's initial pose based on the input given position or the optimal position from the global matching laser contour map and visual feature map, and generate the predicted pose in combination with the motion measurement value.

[0072] The pose prediction unit innovatively integrates the dual modes of manual positioning (i.e., manual prior guidance) and automatic positioning (i.e., multi-sensor data driven): (1) The manual mode adopts a restricted neighborhood search algorithm, takes the manually specified approximate position as the center of the Gaussian distribution, and achieves sub-meter fast convergence through laser contour matching and visual feature correlation calculation, solving the problem that traditional manual calibration is easily affected by subjective errors, and finding the coordinates with the best matching degree as the initial pose; (2) The automatic mode matches the obtained lidar data and image data with the laser map and visual map at the same time, and obtains the coordinates with the best matching degree in the global map as the initial pose.

[0073] And, an iterative execution unit is used to iteratively execute with the predicted posture as the center: optimize the calculation of the contour coincidence between the current contour ranging data and the laser contour map with the current predicted posture as the center, and implement the first observation filter update; extract the semantic features of the current visual feature data, use the semantic features to perform feature point reprojection matching optimization of the visual feature map with the current predicted posture as the center, and implement the second observation filter update; perform positioning correction on the results of the first observation filter update and the second observation filter update, and use the corrected posture as the prediction input for the next cycle.

[0074] Furthermore, the registration and mapping unit includes:

[0075] The dynamic environment compensation subunit is used to construct a motion distortion compensation field based on the angular velocity / linear velocity integration of the IMU and odometer, perform point-by-point motion dedistortion processing on the profile ranging data, perform density clustering-based dynamic target detection on the dedistorted profile ranging data, remove data moving faster than a set speed (such as 0.2m / s), and dynamically adjust the extraction sensitivity of the visual feature data according to the ambient light intensity. When the illumination is <50lux, adaptive histogram equalization preprocessing of the visual feature data is enabled.

[0076] Specifically, based on the IMU angular velocity oh Linear speed with odometer v , construct the motion compensation basic matrix as:

[0077] ;

[0078] Where, oh ( t ) is the instantaneous angular velocity, From the start time t0 of the scan cycle to the current time t The cumulative rotation angle of v x ( t ), v y ( t ) are the X / Y axis linear speeds provided by the odometer,

[0079] The motion distortion compensation field is constructed based on the motion compensation basic matrix:

[0080] ;

[0081] Where, l k For the k The contribution weight of each laser point to the current compensation field, ( x , y ) is the coordinate of any point in the compensation field, ( x k , y k ) is the first k The original measurement value of the point, s is the Gaussian kernel bandwidth parameter, s = R • tan ( i / 2), R To measure distance, i is the angular resolution, N The total number of 2D lidar data collected.

[0082] It should be clarified that online calibration uses a composite calibration target, which includes: laser high-reflectivity stripes, a grid pattern consisting of alternating 20% ​​and 80% reflectivity; infrared feature markers, and AprilTag tags embedded in the temperature differential module, which present controllable temperature difference characteristics under thermal imaging.

[0083] The dynamic joint calibration subunit is used to calculate the external parameter residual matrix of the lidar coordinate system and the visual sensor coordinate system based on the synchronously collected contour ranging data and visual feature data. When the external parameter residual exceeds the set residual value of 0.15 or the covariance determinant of the pose estimation for 5 consecutive frames When the calibration is triggered, the online calibration is performed by extracting dual-modal features from the reflectivity mutation edge of the contour ranging data and the temperature difference feature corner points of the visual feature data.

[0084] Specifically, the online calibration work includes: extracting metal edge features (such as door frames and shelf edges) with reflectivity > 25% / m and incident angle > 60°, generating a laser edge point set P L , and detect infrared high-contrast corner points (such as lamp heat sources and vents) with temperature differences > 2°C / pixel to generate a visual corner point set P V ,Will P L Projected onto the image plane, P V Perform bidirectional RANSAC matching and establish matching pairs to solve the optimal external parameters.

[0085] The multimodal map construction subunit is used to jointly optimize the edge features of the contour ranging data and the SIFT features of the visual feature data using a tightly coupled SLAM framework, respectively constructing a laser contour map as the bottom-level information and a visual feature map as the top-level information. It then performs the fusion of the laser contour map and the visual feature map, establishes cross-modal feature association between laser and vision through bidirectional projection, and obtains a composite map.

[0086] In this embodiment, the laser edge point set P L and visual corner point set P V , impose the following bidirectional projection constraints:

[0087] (1) Laser → Visual Forward Projection: Transform the laser edge point set P L By projecting the calibration parameters onto the visual feature map, the pixel distance between it and the nearest neighboring visual feature point is calculated to establish geometric consistency constraints;

[0088] (2) Vision → Laser Back Projection: Set the visual corner points PV Back-projection is performed onto the laser contour map plane and spatial alignment verification is performed with adjacent laser points to ensure the consistency of the physical position of cross-modal features in three-dimensional space.

[0089] The forward-projected pixel deviation and the back-projected spatial offset are weighted and fused. Matching points with residuals exceeding a threshold (e.g., pixel deviation > 15 pixels or spatial offset > 20 cm) have their weights reduced to near zero. Point pairs identified as anomalous for more than three consecutive frames are temporarily removed from the matching pool to prevent error accumulation. This allows for adaptive adjustment of the weights of the two residuals based on environmental dynamics. For example, this increases the laser constraint weight in scenes with drastic lighting changes, and enhances visual constraints in areas where lasers fail, such as glass curtain walls.

[0090] The map optimization subunit is used to obtain the laser constraint edges that represent the connection between continuous robot pose nodes on the composite map by calculating the geometric contour matching degree between the current contour ranging data and the laser contour map, and to obtain the visual constraint edges that represent the connection between visual feature observation nodes by calculating the three-dimensional feature point reprojection error between the current visual feature data and the visual feature map. Adaptive weights are set between the laser constraint edges and the visual constraint edges, and the weighted residual sum of squares of the laser constraint edges and the visual constraint edges are simultaneously minimized to achieve global consistency optimization of the composite map.

[0091] It's important to understand that laser constraint edges are geometric consistency constraints connecting the robot's consecutive pose nodes. They are generated by matching real-time laser ranging data with a pre-built laser contour map. Specifically, as the robot moves, the currently scanned contour point cloud (such as the edges of walls and columns) is aligned with the geometric structure in the laser map with high precision, and the overlap and shape similarity of the two contours are calculated. During the matching process, relative motion constraints between pose nodes are extracted to form edges connecting the poses at previous and subsequent moments. This constraint forces the movement of adjacent poses during the optimization process to conform to the geometric laws of laser observation, effectively suppressing the accumulated error of the odometry and ensuring the local geometric accuracy of the map.

[0092] The visual constraint edge is a cross-modal constraint connecting the robot's pose node with the visual feature observation node, constructed through the reprojection error of three-dimensional visual feature points. During the localization process, the semantic features extracted from the current image (such as ceiling texture and fixed signs) are matched with the three-dimensional points in the visual feature map. The map points are projected to the current camera's perspective, and the pixel-level deviation between their projected position and the actual image detection position is calculated. This deviation is quantified as a reprojection error, forming a constraint edge connecting the pose node and the feature node. This constraint forces the optimized pose to align the map feature points with the real-time observation in the projected space, thereby using visual semantic information to correct pose drift and enhance robustness to dynamic interference.

[0093] Furthermore, the pose prediction unit includes:

[0094] The manual repositioning subunit is used to parse the initial position coordinates and their confidence radius entered by the user, construct a three-degree-of-freedom search space, and adopt a multi-resolution rasterization strategy to gradually refine the matching granularity within the search space to find the coordinates with the best matching degree as the initial pose. At the same time, a confidence report containing the position covariance matrix is ​​generated. Specifically, an adaptive multi-resolution rasterization strategy is adopted, the initial grid size is set, and the search space is divided into several coarse-grained grid cells. The matching score between each grid center point and the laser map is calculated using a contour matching algorithm (such as NDT). High-probability candidate areas with scores above the threshold are screened out, and these high-probability candidate areas are gradually refined to find the coordinates with the best matching degree.

[0095] The automatic repositioning subunit is used to synchronously perform laser contour matching and visual feature retrieval. It generates a joint score by dynamically weighted fusion of laser geometric similarity and visual semantic matching, and obtains the best coordinates as the initial pose based on the comprehensive matching degree.

[0096] In a specific embodiment, in the automatic repositioning subunit, laser contour matching is performed using an NDT algorithm to obtain the laser geometric similarity:

[0097] ;

[0098] Where, N m is the current number of matching points, N t is the total number of points, RMSE is the root mean square error of registration, k 1=10 is the attenuation coefficient.

[0099] And retrieve similar scenes in the visual map through the BoW bag of words model

[0100]

[0101] Where, M is the number of matching features, T f ( f i ) is the feature weight, cos ( d c , d m ) is the current feature descriptor d c Map feature descriptor d m The cosine similarity of .

[0102] The combined score is: , where are the variance of laser matching error and visual retrieval error, respectively.

[0103] The motion prediction subunit is used to fuse the wheel odometry, IMU angular velocity and laser odometry data, and use the kinematic model to calculate the change in posture.

[0104] The kinematic model is:

[0105] ;

[0106] in, v L , v R is the left and right wheel speed, L is the wheelbase, oh IMU is the IMU angular velocity, ωΔt For the time period Δt The amount of change in the inner angle.

[0107] Prediction optimization subunit, used to generate N candidate pose particles, performs composite likelihood evaluation of laser projection coverage and visual reprojection error on each particle, updates the candidate pose particles through importance resampling, and outputs the weighted average pose as the prediction result.

[0108] In the prediction optimization subunit, based on the motion prediction results, s Confidence ellipse generated N = 200 candidate particles. Next, the laser likelihood is calculated: the matching coverage between the projected laser point cloud and the laser map. The visual likelihood is calculated: the average pixel error of the reprojected visual feature points. Next, normalized weights are calculated, and a systematic resampling method is used to retain high-weight particles and eliminate particles with weights below a threshold (e.g., 1 / 2N). Finally, the weighted average pose is output as the prediction result.

[0109] Furthermore, the iterative execution unit includes:

[0110] The first observation filter update subunit is used to establish an adaptive search area in the laser contour map with the current predicted posture as the center. The search range dynamically adjusts the horizontal radius and heading angle coverage interval according to the robot's movement speed. A multi-scale iterative nearest point matching strategy is adopted to optimize the contour overlap layer by layer between coarse-grained and fine-grained resolutions. An exponential decay weight is applied to contour ranging data points whose moving speed exceeds a preset speed threshold, and the laser matching residual is calculated based on the contour matching results.

[0111] Specifically, in the first observation filter update subunit, the coarse-grained and fine-grained resolutions follow the following logic: Coarse-grained layer (resolution 0.2m): Rapidly screen candidate matching areas and eliminate areas with obvious deviations; Medium-grained layer (resolution 0.1m): Optimize point cloud alignment through normal vector consistency constraints; Fine-grained layer (resolution 0.05m): Achieve sub-centimeter-level registration based on curvature features. It is worth noting that an exponential decay weight is applied to the point clouds of dynamic obstacles whose movement speed exceeds a preset speed threshold to suppress the impact of dynamic interference on the matching results. The exponential decay weight is: ,in, t 1 is the duration of the obstacle. The weight of obstacles exceeding 2 seconds is reset to zero.

[0112] Afterwards, the laser matching residual is calculated based on the contour matching result. The residual includes the position residual root mean square error and heading angle deviation, and is output to the observation result correction subunit.

[0113] The second observation filter update subunit is used to extract semantic features from the current visual data in real time through a lightweight semantic network running on the embedded NPU, construct a three-dimensional projection space within the predicted pose neighborhood, reproject the three-dimensional feature points in the visual feature map to the current image coordinate system and perform two-way matching verification, and calculate the visual reprojection error based on the valid matching pairs obtained through verification.

[0114] In the second observation filter update subunit, a lightweight YOLO-Mini network (model size <8MB) is run through the embedded NPU to extract semantic features in real time. The main extraction is preset categories, including navigation markers such as safety exit signs, QR code labels, reflective positioning balls, and environmental structural features such as ceiling beams, wall distribution boxes, and ground guide wires. Subsequently, 2D-3D feature point pairs with semantic labels are output, and the transmission delay is small. Within the predicted pose neighborhood, a three-dimensional projection space is constructed. Specifically, a three-dimensional projection pyramid is constructed with three projection layers, with heights corresponding to 0.5m / 1.5m / 2.5m respectively, and a resolution of 0.1m / layer to cover the robot's motion space.

[0115] In addition, a bidirectional matching verification mechanism is employed. Forward projection projects 3D points from the visual feature map onto the current image plane and calculates pixel-level deviations. Backward matching back-projects 2D features extracted from the current image back into the map space to verify 3D consistency. A valid matching pair must meet the following criteria: bidirectional projection error is less than 3 pixels, semantic label consistency reaches 100%, and feature scale variation is less than 20% (resistance to viewport variations).

[0116] Furthermore, the following formula is used for visual reprojection error;

[0117] ;

[0118] Where, m is the total number of 3D feature points matched between the current frame and the map, d i For a binary indicator function, the value is {0,1}, and the effective matching pair is d i =1, O i is the semantic confidence weighted, the classification probability of the feature points output by the lightweight semantic network, is the projection coordinate, T pred is the predicted pose matrix of the current cycle, that is, the transformation relationship from the map coordinate system to the current camera coordinate system, The visual feature map i The coordinates of the feature points, is the first detected i The two-dimensional pixel coordinates of the feature points.

[0119] The observation result correction subunit is used to construct a multimodal Kalman filter, which receives the laser matching residual and visual reprojection error respectively, and dynamically allocates sensor weights according to the environmental degradation factor. The environmental degradation factor includes the proportion of laser dynamic obstacles and the uniformity of visual illumination. When the dual-modal positioning deviation exceeds the safety threshold, the arbitration mechanism is activated, and the historical confidence data is reviewed. The sensor data stream with a built-in confidence continuously higher than 80% in the past N seconds is given priority. The confidence is calculated through the condition number of the residual distribution covariance matrix, and the multimodal Kalman filter update is executed to output the corrected pose and covariance matrix as the prediction input for the next cycle.

[0120] In summary, the present invention provides a device for fusion positioning based on vision and 2D laser technology. This device, through the deep integration of a heterogeneous sensor architecture and a computing platform, creates a high-precision, robust positioning solution. Specifically, the device comprises three core modules: a horizontal ranging module, a vertical vision module, and a computational processing module 4. These three modules form a collaborative innovation in spatial layout and data processing.

[0121] First, in terms of hardware architecture design, the horizontal ranging module uses a 2D lidar. A dynamic pitch compensation mechanism maintains strict parallelism between the laser scanning plane and the robot's motion trajectory, acquiring millimeter-level horizontal profile data in real time. The vertical vision module, equipped with an infrared sensor and an active fill light array, is mounted at orthogonal angles. Combined with an infrared bandpass filter to eliminate ambient light interference, it focuses on capturing vertical spatial semantic features such as ceiling structures and hanging signs. This orthogonal layout not only physically isolates signal crosstalk between sensors but also, through coordinate system mapping, lays the geometric foundation for multimodal data fusion.

[0122] On this basis, the computing and processing module 4 achieves collaborative innovation in software and hardware. The multi-core heterogeneous computing platform (CPU + NPU / GPU) improves positioning performance through the following technological breakthroughs:

[0123] Based on a multi-core heterogeneous computing platform (CPU + NPU / GPU), this system performs the following: constructing a coordinate-aligned laser profile map and visual feature map based on profile ranging data and visual feature data; determining the robot's initial pose based on a given input position or the optimal position from a global match between the laser profile map and the visual feature map; generating a predicted pose based on motion measurements; and iteratively performing a dual-modal observation filter update centered on the predicted pose, optimizing laser profile matching and visual reprojection matching. This system, through the deep integration of a heterogeneous sensor architecture and computing platform, creates an all-weather positioning solution for industrial scenarios.

[0124] Crucially, this device solves industry pain points through deep integration of software and hardware: it reduces the communication delay from sensor data acquisition to final result output; uses hardware trigger signals to ensure strict spatiotemporal alignment of laser scanning and image exposure, eliminating matching errors caused by motion blur; and the NPU multiplexing design enables visual feature extraction and laser point cloud processing to share computing resources, reducing overall hardware costs.

[0125] Therefore, the present invention constructs the first "all-weather, all-terrain, and full-dynamic" positioning system in the field of industrial mobile robots through orthogonal sensor layout, multimodal data fusion algorithm, and coordinated optimization of software and hardware, providing a reliable autonomous navigation infrastructure for intelligent manufacturing scenarios.

[0126] Since the systems / devices described in the above embodiments of the present invention are systems / devices used to implement the methods of the above embodiments of the present invention, those skilled in the art will be able to understand the specific structures and variations of these systems / devices based on the methods described in the above embodiments of the present invention, and thus will not be described in detail here. All systems / devices used in the methods of the above embodiments of the present invention are within the scope of protection of the present invention.

[0127] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0128] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.

[0129] It should be noted that the word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The present invention may be implemented by means of hardware comprising several distinct components and by means of a suitably programmed computer. Among the several devices listed, several of these devices may be embodied by the same hardware. The use of the words first, second, third, etc., is for convenience only and does not imply any order. These words should be understood as part of the component name.

[0130] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0131] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concepts. Therefore, the technical solutions should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0132] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the technical solution of the present invention and its equivalents, the present invention shall also include such modifications and variations.

Claims

1. A device based on vision and 2D laser fusion positioning, characterized in that: include: The horizontal ranging module is installed horizontally on the robot body and is used to collect the contour ranging data of the motion plane; Vertical vision module, top view installed on the robot body, used to collect visual feature data in the vertical space; The computing and processing module is configured to: construct a laser contour map and a visual feature map with coordinate system alignment based on the contour ranging data and the visual feature data; determine the initial position of the robot according to a given input position or the optimal position of the laser contour map and the visual feature map from the global matching; generate a predicted position in combination with the motion measurement value; and iteratively perform a dual-modal observation filter update of the laser contour matching optimization and the visual reprojection matching optimization centered on the predicted position; Among them, the calculation processing module includes an iterative execution unit, which includes: a first observation filter update subunit, which is used to establish an adaptive search area in the laser contour map with the current predicted posture as the center, adopt a multi-scale iterative nearest point matching strategy, optimize the contour overlap layer by layer between coarse-grained and fine-grained resolutions, apply exponential decay weights to contour ranging data points whose moving speed exceeds a preset speed threshold, and calculate the laser matching residual based on the contour matching result; a second observation filter update subunit, which is used to extract semantic features in the current visual feature data in real time through a lightweight semantic network running on an embedded NPU, and construct a three-dimensional projection space within the predicted posture neighborhood. , reproject the three-dimensional feature points in the visual feature map to the current image coordinate system and perform two-way matching verification, and calculate the visual reprojection error based on the valid matching pairs obtained by verification; the observation result correction subunit is used to construct a multimodal Kalman filter, which receives the laser matching residual and the visual reprojection error respectively, and dynamically allocates sensor weights according to the environmental degradation factor. The environmental degradation factor includes the proportion of laser dynamic obstacles and the uniformity of visual illumination. When the dual-modal positioning deviation exceeds the safety threshold, the arbitration mechanism is activated, and the sensor data stream whose built-in confidence has been continuously higher than the preset confidence threshold in the past preset period is given priority, and the corrected posture is output as the prediction input for the next cycle.

2. The device based on vision and 2D laser fusion positioning according to claim 1, characterized in that: The horizontal ranging module includes a 2D laser radar. Through the pre-configured dynamic pitch compensation mechanism, the 2D laser radar laser scanning plane and the robot's direction of travel are kept parallel to each other with an error less than the preset parallelism error. The vertical vision module includes an infrared sensor. Through the pre-configured tunable universal joint adjustment, the optical axis of the infrared sensor lens forms an 88°-92° angle configuration with the 2D lidar, so that the vertical field of view angle coverage of the infrared sensor and the horizontal scanning sector of the 2D lidar form a non-overlapping complementary observation domain in the spatial coordinate system.

3. The device based on vision and 2D laser fusion positioning according to claim 2, characterized in that: The complementary observation domain satisfies the following spatial constraints: The 2D laser radar forms a scanning sector θ∈[0°,270°] with adaptive opening angle adjustment in the XY plane of the robot motion coordinate system; The infrared sensor constructs a three-dimensional pyramid-shaped detection area in the Z-axis direction of the Cartesian coordinate system. The pitch angle φ∈[55°,95°] of the infrared sensor forms a buffer isolation zone of no less than 15% with the scanning sector of the lidar on the spatial projection surface.

4. The device based on vision and 2D laser fusion positioning according to claim 1, characterized in that: An active infrared fill light source device is configured, and the active infrared fill light source device includes: A ring-shaped distribution unit, which arranges N independently addressable infrared emission sources around the infrared sensor; The spectrum modulation unit dynamically selects the 850nm or 940nm operating wavelength according to the ambient light entropy value and drives the tunable bandpass filter to perform matched spectrum filtering; The light emitting synchronization controller uses hardware trigger signals to align the pulse emission periods of each infrared emission source with the global shutter exposure window of the infrared sensor; The light source selection module includes a switchable LED light source group and a VCSEL light source group. The LED light source group includes multiple LEDs arranged in a circular array, and the light-emitting surfaces of all LEDs are installed tilted outward with respect to the center normal. The tilt angle range is 5°-15°, and the tilt direction of the LEDs on the same layer is symmetrically distributed; the VCSEL light source group includes multiple vertical cavity surface emitting lasers.

5. The device based on vision and 2D laser fusion positioning according to any one of claims 1 to 4, characterized in that: The calculation processing module includes: A registration and mapping unit is used to construct a laser contour map and a visual feature map with aligned coordinate systems based on the contour ranging data and the visual feature data; The pose prediction unit is used to determine the robot's initial pose based on the input given position or the optimal position from the global matching laser contour map and visual feature map, and generate the predicted pose in combination with the motion measurement value.

6. The device based on vision and 2D laser fusion positioning according to claim 5, characterized in that: The registration and mapping unit includes: The dynamic environment compensation subunit is used to construct a motion distortion compensation field based on the angular velocity / linear velocity integration of the IMU and the odometer, perform point-by-point motion dedistortion processing on the contour ranging data, perform density clustering-based dynamic target detection on the dedistorted contour ranging data, remove data moving faster than a preset speed, and dynamically adjust the extraction sensitivity of the visual feature data according to the ambient light intensity. When the illumination is less than the preset illumination threshold, adaptive histogram equalization preprocessing is enabled for the visual feature data; The dynamic joint calibration subunit is used to calculate the extrinsic parameter residual matrix of the lidar coordinate system and the visual sensor coordinate system based on the synchronously collected contour ranging data and visual feature data. When the extrinsic parameter residual exceeds the preset residual value or the covariance value of the pose estimation of a preset number of consecutive frames is greater than the preset covariance value, the calibration is triggered. The dual-modal feature extraction is performed on the reflectivity mutation edge of the contour ranging data and the temperature difference feature corner point of the visual feature data to perform online calibration. The multimodal map construction subunit is used to jointly optimize the edge features of the contour ranging data and the SIFT features of the visual feature data using a tightly coupled SLAM framework, respectively constructing a laser contour map as the bottom layer of information and a visual feature map as the top layer of information. The laser contour map and the visual feature map are then fused, and a cross-modal feature association between laser and vision is established through bidirectional projection to obtain a composite map. The map optimization subunit is used to obtain the laser constraint edges that represent the connection between continuous robot pose nodes on the composite map by calculating the geometric contour matching degree between the current contour ranging data and the laser contour map, and to obtain the visual constraint edges that represent the connection between visual feature observation nodes by calculating the three-dimensional feature point reprojection error between the current visual feature data and the visual feature map. Adaptive weights are set between the laser constraint edges and the visual constraint edges, and the weighted residual sum of squares of the laser constraint edges and the visual constraint edges are simultaneously minimized to achieve global consistency optimization of the composite map.

7. The device based on vision and 2D laser fusion positioning according to claim 6, characterized in that: Online calibration uses a composite calibration target, which includes: Laser high-reflectivity stripes, a grid pattern consisting of alternating 20% ​​and 80% reflectivity; Infrared feature marker, AprilTag tag embedded with temperature differential module, presents controllable temperature difference characteristics under thermal imaging.

8. The device based on vision and 2D laser fusion positioning according to claim 5, characterized in that: The pose prediction unit includes: The manual repositioning subunit is used to parse the initial position coordinates and their confidence radius input by the user, construct a three-degree-of-freedom search space, and adopt a multi-resolution rasterization strategy to gradually refine the matching granularity within the search space to find the coordinates with the best matching degree as the initial pose. At the same time, it generates a confidence report including the position covariance matrix; The automatic repositioning subunit is used to synchronously perform laser contour matching and visual feature retrieval. It generates a joint score by dynamically weighting the fusion of laser geometric similarity and visual semantic matching, and obtains the optimal coordinates as the initial pose based on the comprehensive matching degree. The motion prediction subunit is used to integrate the wheel odometer, IMU angular velocity and laser odometer data, and use the kinematic model to calculate the pose change; The prediction optimization subunit is used to generate multiple candidate pose particles based on Monte Carlo sampling, perform a composite likelihood evaluation of laser projection coverage and visual reprojection error on each particle, update the candidate pose particles through importance resampling, and output the weighted average pose as the prediction result.

Citation Information

Patent Citations

  • Laser and vision fused integrated global positioning method for indoor robot

    CN112596064A

  • SLAM method based on tight coupling of 2D laser radar and binocular camera

    CN112785702A

  • Positioning method and system

    CN114184193A

  • Instant positioning and mapping method, device and system and storage medium

    CN115597582A