Vision sensor extrinsic calibration method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202611302640.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0002]传统相机-雷达外参标定高度依赖人工靶标(如棋盘格、KSAC球靶),部署时需人工放置、精密测量靶标尺寸,单次准备耗时超15分钟,且靶标易磨损、受环境光照影响,难以适应机器人频繁换场、长期运行的需求
本实施例通过滑动窗口动态维护最近N帧同步采集的双模态的传感数据,在避免历史数据无限累积导致算力爆炸的同时,保证了优化所需的信息冗余度;通过提取各帧双模态的几何基元并结合语义SLAM构建带唯一语义标识的全局环境地图,利用场景中自然存在的稳定结构作为标定参考,摆脱了对人工标定靶的依赖;通过语义标识自动完成跨模态几何基元关联,省去了人工标注对应关系的繁琐流程;仅在标定状态不稳定时才以最近一次的外参为初值,联合优化外参与滑动窗口内各几何基元对的单帧几何参数,既避免了稳定状态下的无效算力消耗,又通过多帧观测的统计特性抑制了单帧噪声干扰,兼顾了在线标定的高精度、低延迟与低功耗,解决了传统一次性离线标定无法感知运行中外参漂移、依赖人工靶标部署成本高的问题。
Smart Images

Figure CN122820864A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and more specifically, to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for calibrating the extrinsic parameters of a vision sensor. Background Technology
[0002] Traditional camera-radar extrinsic calibration relies heavily on manual targets (such as checkerboard or KSAC ball targets). Deployment requires manual placement and precise measurement of target dimensions, with each preparation taking more than 15 minutes. Furthermore, the targets are prone to wear and tear and are affected by ambient light, making it difficult to adapt to the needs of robots that frequently change locations and operate for extended periods.
[0003] While existing online calibration schemes have freed themselves from target dependence, they generally have limitations: they may only use edge features, which are sensitive to lighting and texture, resulting in poor stability; or they may not have built a joint optimization framework, making it difficult to integrate multimodal observation information; and they may lack observability quantification mechanisms, which cannot actively compensate when the scene degrades, making calibration failures likely. Summary of the Invention
[0004] This application provides an external parameter calibration method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can solve the above-mentioned problems of the prior art. The technical solution is as follows: According to one aspect of the embodiments of this application, a method for calibrating the extrinsic parameters of a visual sensor is provided, the method comprising: The robot's vision sensor synchronously collects sensor data from two modalities at a preset frame rate. The sensor data from the two modalities collected in the same frame are combined into a data frame, and a sliding window is constructed. The sliding window includes the latest N data frames. For each data frame within the sliding window, extract the geometric primitives in each sensing data in the data frame and the single-frame geometric parameters of the geometric primitives in the coordinate system of the corresponding mode; An environment map is constructed based on the data frames within the sliding window. The environment map includes each semantic geometric primitive and the global geometric parameters of the semantic geometric primitive in the robot's base coordinate system. The semantic geometric primitive is a geometric primitive with a semantic identifier. The semantic identifier is used to uniquely indicate the physical structure corresponding to the corresponding geometric primitive. For each data frame within the sliding window, the geometric primitives of the two modalities in the data frame are associated according to the environment map to obtain geometric primitive pairs of the same semantic geometric primitive in the two modalities; If it is determined that the calibration state of the extrinsic parameters of the vision sensor is unstable, then the base coordinate system is used as the world coordinate system, the most recently updated extrinsic parameters are used as the initial values for the current frame iterative optimization, the extrinsic parameters and the single-frame geometric parameters of each geometric primitive pair within the sliding window are used as variables to be optimized, and iterative optimization is performed by minimizing the observation constraint set constructed based on geometric consistency to determine the calibration result of the extrinsic parameters in the current frame.
[0005] According to another aspect of the embodiments of this application, a visual sensor extrinsic parameter calibration device is provided, the device comprising: The data acquisition module is used to acquire sensor data from two modalities synchronously collected by the robot's vision sensor at a preset frame rate, combine the sensor data from the two modalities collected in the same frame into a data frame, and construct a sliding window, wherein the sliding window includes the latest N data frames. The parameter extraction module is used to extract the geometric primitives in each sensing data in each data frame within the sliding window, as well as the single-frame geometric parameters of the geometric primitives in the coordinate system of the corresponding mode. The map building module is used to build an environment map based on the data frames in the sliding window. The environment map includes each semantic geometric primitive and the global geometric parameters of the semantic geometric primitive in the robot's base coordinate system. The semantic geometric primitive is a geometric primitive with a semantic identifier. The semantic identifier is used to uniquely indicate the physical structure corresponding to the corresponding geometric primitive. The association module is used to associate the geometric primitives of two modalities in each data frame within the sliding window according to the environment map, so as to obtain the geometric primitive pairs of the same semantic geometric primitive in the two modalities. The iterative optimization module is used to determine the calibration result of the extrinsic parameters in the current frame if the calibration state of the extrinsic parameters of the visual sensor is determined to be unstable. The module uses the base coordinate system as the world coordinate system, the most recently updated extrinsic parameters as the initial values for iterative optimization in the current frame, and the extrinsic parameters and the single-frame geometric parameters of each geometric primitive pair within the sliding window as variables to be optimized. Iterative optimization is performed by minimizing the observation constraint set constructed based on geometric consistency to determine the calibration result of the extrinsic parameters in the current frame.
[0006] According to another aspect of the embodiments of this application, an electronic device is provided, the electronic device including a memory, a processor and a computer program stored in the memory, the processor executing the computer program to implement the above-described method.
[0007] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the above-described method.
[0008] According to one aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described method.
[0009] The beneficial effects of the technical solutions provided in this application are: This embodiment dynamically maintains the bimodal sensor data synchronously acquired in the most recent N frames through a sliding window. This avoids the computational power explosion caused by the infinite accumulation of historical data while ensuring the information redundancy required for optimization. By extracting the geometric primitives of each frame's bimodality and combining them with semantic SLAM to construct a global environment map with unique semantic identifiers, it utilizes naturally existing stable structures in the scene as calibration references, eliminating the dependence on manual calibration targets. Semantic identifiers automatically complete the cross-modal geometric primitive association, saving the tedious process of manually labeling correspondences. Only when the calibration state is unstable is the most recent extrinsic parameter used as the initial value to jointly optimize the single-frame geometric parameters of the extrinsic parameter and each geometric primitive pair within the sliding window. This avoids the ineffective computational power consumption in the stable state and suppresses single-frame noise interference through the statistical characteristics of multi-frame observations. It balances the high accuracy, low latency, and low power consumption of online calibration, solving the problems of traditional one-time offline calibration's inability to detect extrinsic parameter drift during operation and the high cost of relying on manual target deployment. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0011] Figure 1 A schematic diagram of a robot system for implementing a method for calibrating the extrinsic parameters of a vision sensor, as provided in an embodiment of this application; Figure 2 This application provides a method for calibrating the extrinsic parameters of a vision sensor. Figure 3 A schematic diagram of the physical space model provided in the embodiments of this application; Figure 4 A schematic diagram of an online joint calibration system for vision-radar extrinsic parameters of a robot platform, provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an external parameter calibration device for a vision sensor provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0013] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0015] First, let's introduce and explain several terms used in this application: Extrinsic parameters: Rigid body transformation parameters of the sensor coordinate system relative to the robot's base coordinate system, including the rotation matrix R∈SO(3) and the translation vector t ∈ R. 3 It has a total of 6 degrees of freedom. This application calibrates head→camera (T HC ) and base→LiDAR(T BL Two sets of external parameters, totaling 12DOF.
[0016] Semantic geometric primitives: Semantic structures in a scene with explicit geometric parameters, including planes. (normal vector n∈S) 2 +intercept d∈R), edge line (Plücker coordinates), cylinder / sphere (Axis + Radius). Remains globally invariant in the base coordinate system, naturally serving as the reference frame for extrinsic parameter calibration.
[0017] Sliding Window Factor Graph: An online optimization framework that constructs a factor graph by retaining the most recent W frames of sensor data, and removing the oldest frame (FIFO) when new data arrives. A fixed window size W ensures a constant computational cost of O(W²·M² + W·M·N), making it suitable for long-term online operation.
[0018] Fisher Information Matrix: Defined as F = J^TΣ - ¹J, where J is the Jacobian matrix of the residual vector with respect to the parameter vector, and Σ is the observation noise covariance matrix. F -1 Give the Cramér-Rao lower bound (CRLB) and Cov( for any unbiased estimator) )≥F -1 .
[0019] Cramér-Rao Lower Bound (CRLB): A theoretical lower bound on the covariance of any unbiased estimator. This application demonstrates, through a CRLB comparative analysis, that the semantic plane method has a tighter theoretical accuracy bound than the sphere-object method (KSAC).
[0020] Cross-modal plane consistency: Observations of the same physical plane from the camera and LiDAR should be aligned after transformation to the base frame. This is achieved through the plane transformation operator T(R,t)×(n,d) =(Rn, d). t T Rn) unifies the constraints between the two sets of observations. The normal component depends only on the rotational extrinsic parameter, while the intercept component carries translation information.
[0021] Random Sample Consensus (RANSAC) False Association Removal: A Plane Pair Selection Process Based on the RANSAC Algorithm. Three pairs of candidate pairs with non-coplanar normals are randomly selected. Temporary extrinsic parameters are solved using a closed-form algorithm. Consistency voting is used to mark interior points (threshold τ). inlier =0.05m+0.1rad), iteration K ransac =50 times.
[0022] The AX=YB problem is a mathematical form of the classic hand-eye calibration problem. This application extends it to a semantic version: replacing the sphere center coordinates with planar parameters (n,d), coupling the two extrinsic parameters through shared planar parameters, and solving for the initial rotation values using the dual quaternion method.
[0023] SE(3): A three-dimensional special Euclidean group, containing the rotation group SO(3) and translation R.3 There are a total of 6 DOF. All extrinsic parameters and poses in this application are represented by SE(3) left perturbation, consistent with the Sophus library.
[0024] Kinematics-anchored Sphere Alignment Calibration (KSAC): A multi-sensor extrinsic parameter calibration method based on a spherical target, which establishes constraints by using the distance between the elliptical tangent and the LiDAR sphere.
[0025] Huber Robust Kernel: An M-estimation loss function that uses L2 cost (efficient for normal observations) when residuals are less than a threshold δ, and L1 cost (robust to outlying points) when residuals exceed the threshold. The LiDAR factor δ in this application... L =0.02m, camera factor δ C =1.0px.
[0026] KSAC is the calibration method most similar to this invention, which establishes camera-LiDAR extrinsic parameter constraints using a spherical target. Its core process includes: (1) Place foamed balls in the robot's workspace; (2) Accurately measure the radius of the sphere; (3) Acquire multiple frames of camera images and LiDAR point clouds; (4) Establish constraints using the distance between the ellipse tangent and the sphere; (5) AX=YB Closed-loop solution of initial values followed by nonlinear optimization.
[0027] The advantages of KSAC include a clear geometric model of a single target, stable closed-form initial value solutions, and the ability to resist occlusion when the elliptical portion is visible. However, its essential limitation lies in: (a) The ball still needs to be placed manually and its radius measured (with mm-level accuracy requirements), and calibration data cannot be supplemented after deployment; (b) The observation of the sphere center is a single-point constraint, and the estimation error of the point cloud normal vector is directly transmitted to the sphere center fitting. (c) Once calibration is complete, external parameters are frozen, and drift caused by collisions / temperature changes / mechanical fatigue during operation cannot be detected; (d) The ball does not exist in the natural working scenario, so lifelong online calibration cannot be achieved.
[0028] MARSCalib extends spherical target calibration to multi-robot scenarios, enhancing robustness in harsh environments (outdoors, off-ground). However, it still relies on spherical targets, lacks online drift detection and proactive adaptation mechanisms, and the multi-robot assumption is not applicable to long-term single-robot deployment scenarios.
[0029] Related technology one also uses RANSAC for plane detection and fitting in LiDAR point clouds, combined with edge detection to achieve camera-LiDAR calibration. However, it essentially still uses traditional calibration boards such as checkerboards as reference targets; it lacks online optimization of sliding window factor maps; it lacks Fisher information theory analysis; and it lacks an active exploration mechanism. Essentially, it uses RANSAC to improve the robustness of traditional target methods, and it has not broken away from the target-dependent framework.
[0030] Related technology two uses semantic features to achieve online calibration through quality index optimization. Its limitations include: using only a single quality index and not constructing a factor graph joint optimization framework; lacking cross-modal plane consistency constraints; lacking Fisher information-driven active exploration; and lacking CRLB theoretical accuracy analysis.
[0031] Related technology three utilizes cascaded optimization of motion trajectories and edge features to achieve online calibration of LiDAR cameras. However, it uses edge features (rather than the semantic plane) as constraints, and edge detection is sensitive to illumination and texture; it lacks multi-layer residual factors (especially the normal-intercept decomposition of cross-modal consistency factors); and it lacks an active exploration mechanism to compensate for unobservable dimensions.
[0032] The extrinsic parameter calibration method, apparatus, electronic device, computer-readable storage medium, and computer program product for visual sensors provided in this application aim to solve the above-mentioned technical problems of the prior art. This application transfers the calibration reference from artificial targets to the semantic geometric structure that objectively exists in the scene, and realizes the non-interventional online extrinsic parameter calibration of the robot throughout its entire life cycle through three integrated technical modules.
[0033] This application includes the following features: Feature 1: Target-independent semantic geometric primitive labeling framework Traditional calibration methods (KSAC, checkerboard, MARSCalib) rely on manual calibration objects—requiring pre-placement, radius measurement (with mm-level accuracy), and surface quality maintenance. This application completely abandons manual targets, instead using naturally occurring semantic geometric structures in the scene (walls, floors, columns, traffic signs, etc.), through a unified planar transformation operator:
[0034] Establish the mathematical foundation for cross-modal semantic alignment. Planar observations provide 6-DOF constraints (vs. 4-DOF for spheres), significantly improving information gain: the plane normal vector is obtained from statistical estimation of multiple points, which greatly reduces the sensitivity to LiDAR ranging noise compared to single-point sphere center constraints.
[0035] Feature 2: Online lifetime calibration of sliding window factor plots Breaking away from the one-step process of acquisition → offline optimization → freeze, a sliding window factor graph (W=30 frames) is constructed to achieve calibration while performing tasks. Four types of residual factors—LiDAR plane point distance factor, camera plane reprojection factor, cross-modal plane consistency factor, and optional plane parameter prior factor—are all jointly optimized using Huber robust kernels. Incremental updates are performed when new frames arrive, and the oldest frame is deleted to maintain a constant computational load (50-100ms / iteration, Intel i5). The cross-modal consistency residual is innovatively decomposed into normal and intercept components: the normal component depends only on rotation extrinsic parameters, while the intercept component... The Rn scalar projection carries translation information—this decomposition directly reveals the degradation mechanism of translation 2DOF when all plane normals are parallel.
[0036] Feature 3: Fisher's Information-Driven Proactive Exploration When online monitoring detects that some extrinsic parameters are not observable (e.g., information missing ratio > 0.2), the system automatically plans robot motion to maximize Fisher information gain. Missing constraint dimensions are identified through Fisher matrix eigenvalue decomposition, and compensating motion trajectories—circular scan (for missing plane diversity) and 3-axis sweep (for missing rotation diversity)—are planned based on the eigenvector directions. Execution is performed incrementally at 5° / s, with one frame sampled per second for real-time closed-loop feedback. The system stops immediately when the information missing ratio drops below a threshold, resuming task execution. Simultaneously, residual mean monitoring (if it rises above 30%) detects long-term extrinsic parameter drift, automatically triggering proactive exploration and reinforcement calibration.
[0037] The table below shows the optional technical parameters for embodiments of this application.
[0038] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0039] Figure 1This is a schematic diagram of a robot system implementing the extrinsic parameter calibration method for a vision sensor according to an embodiment of this application. The vision sensor 110 is installed in the robot 100 and can communicate with the calibration device 120 of the robot 100 via wired or wireless means. The calibration device 120 acquires sensing data from two modalities synchronously collected by the vision sensor 110 at a preset frame rate. The sensing data from the two modalities collected in the same frame are combined into a data frame, and a sliding window is constructed, including the latest N data frames. For each data frame within the sliding window, geometric primitives in each sensing data element and the single-frame geometric parameters of the geometric primitives in the coordinate system of the corresponding modality are extracted. An environment map is constructed based on the data frames within the sliding window. The environment map includes various semantic geometric primitives and the global geometric parameters of the semantic geometric primitives in the robot's base coordinate system. The semantic geometric primitives are geometric primitives with semantic identifiers. Semantic identifiers are used to uniquely indicate the physical structure corresponding to the corresponding geometric primitives. For each data frame within the sliding window, the geometric primitives of the two modalities in the data frame are associated according to the environment map to obtain the geometric primitive pairs of the same semantic geometric primitives in the two modalities. If it is determined that the calibration state of the extrinsic parameters of the vision sensor is unstable, the base coordinate system is used as the world coordinate system, the most recently updated extrinsic parameters are used as the initial values for iterative optimization of the current frame, and the extrinsic parameters and the single-frame geometric parameters of each geometric primitive pair within the sliding window are used as variables to be optimized. Iterative optimization is performed by minimizing the observation constraint set constructed based on geometric consistency to determine the calibration result of the extrinsic parameters in the current frame.
[0040] In one embodiment, the calibration device 120 may be, but is not limited to, a mobile phone, computer, smart voice interaction device, smart home appliance, vehicle terminal, or a processor built into a robot.
[0041] This application provides a method for calibrating the extrinsic parameters of a visual sensor, such as... Figure 2 As shown, the method includes: S101. Acquire sensor data of two modalities synchronously collected by the robot's vision sensor at a preset frame rate, combine the sensor data of the two modalities collected in the same frame into a data frame, and construct a sliding window.
[0042] In some embodiments, the robot's vision sensors include an RGB-D camera and a LiDAR (Light Detection and Ranging) sensor, which are synchronized via hardware triggering. Optionally, the synchronization acquisition frequency is 10Hz to 30Hz. In some embodiments, the RGB-D camera outputs a color image with a resolution of 640×480 and the corresponding depth point cloud, while the LiDAR outputs the XYZ coordinates and reflection intensity information of a 32-line mechanically rotated point cloud.
[0043] In some embodiments, the size N of the sliding window can be dynamically adjusted according to computing power, preferably 30 frames. The window is maintained using a First-In-First-Out (FIFO) mechanism: whenever a new frame of synchronization data arrives, the frame is added to the end of the window, while the oldest frame at the beginning of the window is deleted, ensuring that the window always retains the sensor data of the most recent N frames. This guarantees the amount of historical information required for optimization while controlling the computational complexity of a single iteration to be stable at 50ms~100ms (Intel i5 processor platform), meeting the requirements for online operation.
[0044] It should be noted that this application is not limited to the combination of RGB-D camera and LiDAR. In other embodiments, the observation device can also be replaced by a stereo camera and solid-state LiDAR, or a multi-sensor combination including IMU, as long as the observation data of the two modalities can be independently detected to obtain semantic geometric primitives.
[0045] S102. For each data frame within the sliding window, extract the geometric primitives in each sensing data in the data frame and the single-frame geometric parameters of the geometric primitives in the coordinate system of the corresponding mode.
[0046] In this embodiment, for each data frame within the sliding window, the geometric primitives and corresponding single-frame geometric parameters of the two modal sensing data are extracted respectively.
[0047] In some embodiments, for a specific camera mode, a 3D point cloud in the camera coordinate system is extracted from the RGB-D depth map. Candidate planes are detected using the Plane R-CNN plane detection network, and the least squares method is applied to each candidate plane. The fitting yields the set parameters of a single frame in the camera coordinate system—the planar parameters. ,in ∈S 2 It is a plane normal vector. ∈R is the plane intercept.
[0048] In some embodiments, for radar modes, the RANSAC algorithm is used to detect planes in the radar point cloud accumulated in the current frame, and the single-frame set parameters—plane parameters—are fitted to obtain them in the radar coordinate system. The parameter definitions are consistent with those of the camera modes.
[0049] In some embodiments, in addition to the planar primitives mentioned above, other geometric primitives such as edge lines (represented by Plücker coordinates) and cylinders / spheres (represented by axes and radii) can be extracted according to scene requirements. All primitives are parameterized in the coordinate system of the corresponding modality.
[0050] S103. Construct an environment map based on the data frames within the sliding window.
[0051] The environment map in this embodiment includes each semantic geometric primitive and the global geometric parameters of the semantic geometric primitive in the robot's base coordinate system. The semantic geometric primitive is a geometric primitive with a semantic identifier; the semantic identifier is used to uniquely indicate the physical structure corresponding to the corresponding geometric primitive.
[0052] This application embodiment constructs an environment map using the ORB-SLAM3 semantic SLAM system. The environment map contains all observed semantic geometric primitives and their global geometric parameters in the robot's base coordinate system: the global geometric parameters of the planar primitives are ( , The global geometric parameters of the edge line primitives are Plücker coordinates. , Each primitive carries a unique semantic identifier, which is used to uniquely indicate the corresponding physical structure.
[0053] Specifically, in this embodiment, the planar primitives detected in each frame within the sliding window are back-projected onto the base coordinate system through the current temporary extrinsic parameters. Through loop closure detection and pose graph optimization in semantic SLAM, planar primitives of the same physical structure are associated as the same semantic geometric primitives and assigned unique semantic identifiers, such as wall 001, cylinder 002, etc.
[0054] Each semantic geometric primitive maintains a set of global geometric parameters in the base coordinate system. The global geometric parameters are continuously updated through statistical averaging of multiple frames of observations, and are not affected by the observation noise of a single frame, thus naturally serving as a stable reference frame for extrinsic parameter calibration.
[0055] It should be noted that the environment map is dynamically updated with the addition of new frames: when a newly detected planar primitive fails to match an existing semantic identifier, a new semantic geometric primitive is created; when a semantic geometric primitive is not observed for several consecutive frames, it is aged out and deleted from the map.
[0056] S104. For each data frame within the sliding window, the geometric primitives of the two modalities in the data frame are associated according to the environment map to obtain geometric primitive pairs of the same semantic geometric primitive in the two modalities.
[0057] This step enables automatic cross-modal plane matching without the need for manual labeling of correspondences.
[0058] For the camera end plane detected in the current frame ( , First, the current temporary extrinsic parameters are transformed to the base coordinate system and compared one by one with the global plane parameters in the environment map. The matching conditions are: normal deviation is less than 0.1 rad and intercept deviation is less than 0.05 m. After a successful match, the corresponding semantic label is added to the camera plane.
[0059] Similarly, for the LiDAR end plane detected in the current frame ( , The same matching process is then performed. Ultimately, the camera-side plane and LiDAR-side plane corresponding to the same semantic identifier are combined into geometric primitive pairs with the same semantic geometric primitive {( , ), ( , )}.
[0060] To further eliminate erroneous associations, the geometric primitive pairs can be further filtered using the RANSAC process to retain the set of interior points with high consistency, thus avoiding the pollution of optimization results by erroneous constraints.
[0061] In this embodiment, for each data frame within the sliding window, the geometric primitives of the two modalities are associated based on the aforementioned environment map: the geometric primitives extracted from the camera in the current frame are transformed to the base coordinate system through temporary extrinsic parameters and matched with the global set parameters of existing primitives in the map; the same matching process is performed on the geometric primitives extracted from the radar; the geometric primitives of the two modalities that have been matched and have the same semantic identifier are formed into a pair of geometric primitives with the same semantic geometric primitive, such as the single-frame geometric parameters of the camera on the left wall and the single-frame geometric parameters of the radar forming a pair of geometric primitives.
[0062] S105. If it is determined that the calibration state of the extrinsic parameters of the visual sensor is unstable, then the base coordinate system is used as the world coordinate system, the most recently updated extrinsic parameters are used as the initial values for iterative optimization in the current frame, the extrinsic parameters and the single-frame geometric parameters of each geometric primitive pair in the sliding window are used as variables to be optimized, and iterative optimization is performed by minimizing the observation constraint set constructed based on geometric consistency to determine the calibration result of the extrinsic parameters in the current frame.
[0063] In some embodiments, the criteria for determining an unstable calibration state include, but are not limited to: (1) The number of valid geometric primitive pairs within the sliding window is less than 5; (2) The maximum included angle of all plane normals is less than 10°, which poses a risk of degradation; (3) The increment of the previous optimization is greater than the convergence threshold (translation 0.1mm, rotation 0.01°). (4) The mean residual value increased by more than 30% compared to the benchmark value.
[0064] When any of the above conditions are met, the calibration state is determined to be unstable, triggering iterative optimization. Iterative optimization uses the base coordinate system as the world coordinate system and constructs an observation constraint set containing multiple types of residual factors.
[0065] In this embodiment, the extrinsic parameters of the vision sensor include two sets of rigid body transformation parameters: extrinsic parameters from the head coordinate system to the camera coordinate system. T HC ∈SE(3) (rotation matrix) R HC Translation vector t HC ) and the extrinsic parameters from the base coordinate system to the radar coordinate system T BL ∈SE(3) (rotation matrix) R BL Translation vector t BL ( ), totaling 12 degrees of freedom.
[0066] When the calibration state is determined to be unstable, the base coordinate system is used as the unified world coordinate system, and the most recently updated extrinsic parameters are used as the initial values for the current frame iteration optimization. The aforementioned extrinsic parameters, along with the global geometric parameters of all valid geometric primitive pairs within the sliding window in the base coordinate system (e.g., the values of planar primitives), are then used. , ).
[0067] This embodiment constructs an observation constraint set containing multiple types of residual factors, with the optimization objective being to minimize the sum of weighted Huber costs of all residual factors. Nonlinear iterative optimization is performed using Ceres Solver. After the optimization converges, the calibration results of the extrinsic parameters in the current frame are output and updated to the robot's perception system for use in downstream navigation, grasping, and other tasks.
[0068] In some embodiments, the observation constraint set includes at least one of the following: LiDAR Planar Point Distance Constraint: Constrains the distance from points in the radar point cloud belonging to a certain plane to the fitted plane of that plane in the radar coordinate system. Huber threshold. =0.02 m ; Camera plane reprojection constraint: Constrains the distance from the image boundary pixels to the contour line obtained by backprojecting the plane in the camera coordinate system, Huber threshold. =1.0m px ; Cross-modal plane consistency constraint: The core constraint term aligns observations of the same plane from the camera and radar ends to the base coordinate system via extrinsic parameter transformation. The residuals are decomposed into normal components (dependent only on rotational extrinsic parameters) and intercept components (carrying translation information through scalar projection). The Huber threshold is set to 0.02 for each modality.m / 1.0 px ; Prior constraints on planar parameters: If a plane has accumulated multiple frames of observations in the semantic map, there are coarsely estimated global geometric parameters ( , and covariance Adding prior constraints accelerates convergence.
[0069] After optimization and convergence (increment less than 0.1mm translation and 0.01° rotation), the extrinsic parameter calibration results of the current frame are output and updated to the robot's perception system for use in downstream navigation, grasping and other tasks.
[0070] Taking an indoor inspection robot as an example, no calibration target needs to be placed after the robot starts up: In an office setting, the system automatically detects planar structures such as the floor, walls, and sides of desks to build a global semantic map; a sliding window maintains the camera images and LiDAR point cloud of the most recent 30 frames in real time; when the robot moves and causes a change in the planar observation perspective, it automatically triggers optimization and updates the extrinsic parameters of the visual sensor; if the robot accidentally collides and causes the extrinsic parameters to shift, and the residual mean increases by more than 30%, the system automatically determines it as drift and restarts optimization; when all planar normals are nearly parallel, resulting in an unobservable translation dimension, the system automatically plans a circular scan motion to compensate for missing information. The entire process requires no manual intervention, achieving maintenance-free online calibration throughout its entire lifecycle.
[0071] This embodiment dynamically maintains the bimodal sensor data synchronously acquired in the most recent N frames through a sliding window. This avoids the computational power explosion caused by the infinite accumulation of historical data while ensuring the information redundancy required for optimization. By extracting the geometric primitives of each frame's bimodality and combining them with semantic SLAM to construct a global environment map with unique semantic identifiers, it utilizes naturally existing stable structures in the scene as calibration references, eliminating the dependence on manual calibration targets. Semantic identifiers automatically complete the cross-modal geometric primitive association, saving the tedious process of manually labeling correspondences. Only when the calibration state is unstable is the most recent extrinsic parameter used as the initial value to jointly optimize the single-frame geometric parameters of the extrinsic parameter and each geometric primitive pair within the sliding window. This avoids the ineffective computational power consumption in the stable state and suppresses single-frame noise interference through the statistical characteristics of multi-frame observations. It balances the high accuracy, low latency, and low power consumption of online calibration, solving the problems of traditional one-time offline calibration's inability to detect extrinsic parameter drift during operation and the high cost of relying on manual target deployment.
[0072] Please see Figure 3 The figure illustrates a schematic diagram of the physical space model provided in this embodiment, as shown in the figure. This embodiment involves four core coordinate systems: Base coordinate system (B): Fixed at the base center of the robot body, serving as the global reference benchmark for the entire system.
[0073] Head coordinate system (H): Fixed to the robot's head, its posture changes with the movement of the head joints.
[0074] Camera coordinate system (C): Fixed inside the RealSense D435i camera and connected to the camera's optical center.
[0075] The LiDAR coordinate system (L) is fixed inside the Ouster OS1-32 lidar and is permanently connected to the lidar's optical center.
[0076] During the system initialization phase, the robot body's high-precision extrinsic parameters relative to the base are known through high-precision joint encoders or factory calibration. However, the extrinsic parameters of the camera relative to the head (including rotation and translation) and the extrinsic parameters of the lidar relative to the base are in a state of uncalibrated condition due to mechanical installation errors, temperature drift caused by long-term operation, or vibration.
[0077] This embodiment uses two sensors installed on the robot's head to simultaneously perceive the surrounding environment: Camera-side observation: The RealSense D435i extracts planar features from the environment by fusing depth and RGB images and using deep learning models such as Plane R-CNN to obtain planar parameters in the camera coordinate system. n C , d C ),in n C It is a plane normal vector. d C The distance from the plane to the optical center of the camera. LiDAR observation: Ouster OS1-32 uses the RANSAC algorithm to fit planar features from the point cloud, obtaining planar parameters in the LiDAR coordinate system. n L , d L ).
[0078] The extracted planar parameters, under the semantic SLAM mapping, are assigned globally invariant semantic identifiers (such as ground and walls) and transformed to the Base coordinate system to form scene semantic primitives (such as planes P1, P2, and P3), whose parameters are ( n B , d B ).
[0079] like Figure 3As shown, the semantic geometric primitives in the scene (such as walls P1 and P2, and the ground P3) are fixed in space, forming a naturally invariant global calibration reference. This means that: The actual parameters of the same physical plane (e.g., P1) in the Base coordinate system ( n B1 , d B1 () is a constant that does not change with time.
[0080] Although the camera and lidar have unknown external parameters, resulting in variations in the plane parameters they observe ( n C , d C )and( n L , d L There are differences, but these differences can be solved. T HC and T BL This is done by utilizing the observation differences of the same physical plane under different sensors to construct cross-modal consistent residuals, thereby driving the optimization of extrinsic parameters.
[0081] During system operation, the robot body (serving as a calibration anchor point) first acquires its own posture through joint encoders, and then head movements drive the cameras and LiDAR to collect environmental data. The system correlates the bimodal planar observation data with a semantic map in the Base coordinate system, and then, within the factor graph optimization framework, [the system further optimizes the data]. T HC and T BL As the variable to be optimized, the global parameters of the semantic primitives are used as priors. By minimizing the cross-modal residuals, the high-precision extrinsic parameters are finally solved.
[0082] In some embodiments, as an optional implementation, it further includes: S201. Determine the Jacobian matrix of each residual pair of the variable to be optimized corresponding to the observation constraint set during the iterative optimization, and determine the Fisher information matrix based on the Jacobian matrix.
[0083] In the embodiments of this application, during the residual calculation stage of each iteration optimization, the Jacobian matrix of all residuals in the observation constraint set for the variables to be optimized is calculated simultaneously. J The variables to be optimized include: two sets of extrinsic parameters ( T HC The 6 degrees of freedom T BLThe model has 6 degrees of freedom (12 dimensions in total) and global geometric parameters of M semantic geometric primitives within a sliding window (each planar primitive contains a 3-dimensional normal vector + a 1-dimensional intercept, for a total of 4M dimensions), resulting in a total dimension of 12 + 4M. It should be noted that the size N of the sliding window represents the number of data frames, while M represents the number of independent semantic geometric primitives observed during the sliding window. The same semantic geometric primitive can be observed repeatedly in multiple frames, but it corresponds to only one set of global geometric parameters in the factor graph. Therefore, the dimension of the variable to be optimized is determined by the number of primitives M, not the number of frames N.
[0084] Jacobian matrix J Each row corresponds to a residual term, each column corresponds to a dimension of the variable to be optimized, and the element value is the partial derivative of the residual with respect to the corresponding variable: For point cloud planar distance residual Calculate its pair T BL The rotation and translation components, and the partial derivatives of the corresponding planar radar coordinate system parameters; For camera plane reprojection residual Calculate its pair T HC The rotation and translation components, and the partial derivatives of the corresponding planar camera coordinate system parameters; For the cross-modal plane consistency residual, calculate the partial derivatives of the normal component with respect to the rotational extrinsic parameter and the partial derivatives of the intercept component with respect to the translational extrinsic parameter, respectively. For the a priori residuals of the plane parameters, calculate their partial derivatives with respect to the global geometric parameters of the plane.
[0085] Observation noise covariance matrix This is a diagonal matrix, with diagonal elements set according to sensor characteristics: LiDAR single-point noise. =0.01m, corresponding to the covariance of the residuals Camera pixel noise =1.0 px The covariance of the corresponding residuals .
[0086] According to the Jacobian matrix J and noise covariance matrix According to the formula F = J T J Calculate the Fisher information matrix F . F Let be a real symmetric positive definite matrix, and its inverse matrix be... F -1 The lower bound of the covariance of the extrinsic parameter estimation (CRLB) represents the theoretical optimal accuracy of the extrinsic parameter estimation under the current observation data.
[0087] Perform eigenvalue decomposition on the target: ,in: =diag( , ,…, ) is an eigenvalue diagonal matrix. K The total dimension of the variable to be optimized, and each feature value >0, The larger the value, the higher the observable information of the parameters corresponding to that dimension; the smaller the value, the less information there is in that dimension and the greater the uncertainty in parameter estimation. This is an eigenvector matrix, where each column is an eigenvector corresponding to an eigenvalue, representing the physical direction of that dimension in the robot's base coordinate system. For example, if an eigenvector corresponds to a translational degree of freedom, its direction is the direction of translational non-observable in the base coordinate system; if it corresponds to a rotational degree of freedom, its direction is the direction of the rotational axis, which is not observable in rotation.
[0088] S202. Count the number of feature values less than the preset feature value threshold and calculate the information missing ratio.
[0089] This application embodiment addresses the feature value threshold. The value is not limited; for example, it can be 1.0, corresponding to the lower bound of the covariance. F -1 The diagonal elements do not exceed 1.0, and the statistics are as follows. F All less than The number of feature values is used to calculate the information missing ratio using the formula:
[0090] This indicator directly reflects the proportion of unobservable dimensions in the current external parameter estimation: Info-Deficit=0 indicates that all dimensions have sufficient observations and the parameter estimation has reached the theoretical optimal accuracy; Info-Deficit>0 indicates that some dimensions have insufficient information and the calibration results have uncertainty.
[0091] In this embodiment, the preset information missing ratio threshold is 0.2. That is, when more than 20% of the degrees of freedom are missing, the reliability of the current calibration result is determined to be insufficient, and the active exploration process needs to be triggered.
[0092] S203. If the information missing ratio is greater than a preset information missing ratio threshold, then the feature vector corresponding to the feature value that is less than the preset threshold is determined as the direction of unobservable degrees of freedom.
[0093] If Info-Deficit > 0.2, then it will be less than The eigenvectors corresponding to the eigenvalues are extracted and determined as the set of directions of unobservable degrees of freedom. }. Each This corresponds to a considerable degree of freedom: like If the magnitude of the first three dimensions (corresponding to translational degrees of freedom) is greater than the magnitude of the last three dimensions (corresponding to rotational degrees of freedom), then the unobservable degree of freedom is determined to be a translational degree of freedom, and its physical direction is determined by... The first 3 dimensions of the vector indicate; Conversely, it is determined to be a rotational degree of freedom, and its rotation axis direction is from The last 3-dimensional vector indication.
[0094] For example, when there is only a single plane parallel to the ground in the scene, the eigenvalue decomposition of the Fisher matrix will show that the eigenvalues corresponding to the two translational degrees of freedom along the direction of gravity are less than [value missing]. That is, vertical translation is not observable. S204. Control the robot to perform compensating motion according to the direction of the unobservable degrees of freedom, move the data frame collected during the execution of the compensating motion into the sliding window to replace the old data frame, and re-execute the iterative optimization according to the updated sliding window until the information missing ratio is not greater than the preset threshold, stop the compensating motion, and obtain the new calibration result of the external parameters.
[0095] This embodiment upgrades the empirical judgment degradation in traditional calibration methods to quantitative index evaluation through eigenvalue decomposition of the Fisher information matrix. It can accurately identify which degrees of freedom (DOFs) have sufficient information and which are missing, solving the problem of determining when to supplement data in online calibration. The correspondence between eigenvectors and physical degrees of freedom can intuitively explain the reasons for degradation: for example, if an eigenvector points vertically, it indicates that the planes in the scene are mostly horizontal, lacking vertical translation constraints; if an eigenvector corresponds to pitch axis rotation, it indicates that the head lacks pitch direction movement, making it impossible to constrain the external parameters of rotation along that axis. The directions of the unobservable degrees of freedom identified in this embodiment directly guide the planning of subsequent compensation motions, avoiding the inefficiency of blind motion sampling in traditional methods. This reduces the number of effective frames required for calibration from 12-15 frames in KSAC to 8-12 frames, improving sample efficiency by 25%-33%.
[0096] Assuming LiDAR single-point noise σ = 0.01m, camera noise 1.0 px, calibration range of camera 0.5-1.5m / LiDAR 1.5-3.0m, and high planar diversity across 20 frames. CRLB theoretical estimation results: - Rotation RMSE: KSAC 0.080° → S 3 -Calib 0.045° (decreased by 44%) - Translation RMSE: KSAC 3.5mm → S 3 -Calib 1.8mm (decreased by 49%) - Rotation RMSE: KSAC 0.15° → S 3 -Calib 0.08° (decreased by 47%) - Translation RMSE: KSAC 8.0mm → S 3 -Calib 4.2mm (decreased by 48%) For example, when the robot operates in a pure corridor environment, the walls and floor on both sides are vertical planes, and their normals are parallel to the horizontal plane. In this case, the eigenvalue decomposition of the Fisher matrix shows that the eigenvalues corresponding to the two translational degrees of freedom along the vertical direction (z-axis) are only 0.2 < The information missing ratio reached 2 / 12 ≈ 0.17, close to the threshold of 0.2. The system identified the unobservable direction as z-axis translation and automatically planned a circular scan of the head around the base (no vertical movement; the circular scan can introduce planar observations from different perspectives to supplement the z-axis translation constraint). After performing 3 scans, 9 frames of new data were collected, and the Fisher matrix was recalculated, showing that all eigenvalues were greater than 0.2%. The information missing rate dropped to 0, and the uncertainty of external parameter estimation was significantly reduced.
[0097] Based on the above embodiments, as an optional embodiment, the unobservable degrees of freedom include at least one of rotational degrees of freedom and translational degrees of freedom; The step of controlling the robot to perform compensatory motion according to the direction of the unobservable degrees of freedom includes: If the unobservable degrees of freedom include translational degrees of freedom, then control the robot to perform a circular scanning action; If the unobservable degrees of freedom include rotational degrees of freedom, then the robot is controlled to sequentially perform scanning actions involving yaw, pitch, and roll rotations along the three axes.
[0098] This embodiment addresses the set of directions of unobservable degrees of freedom { }, one by one for each Perform type determination: Determination of translational degrees of freedom: If the eigenvectors If the magnitude of the first three elements (corresponding to the translational components of the SE(3) perturbation) exceeds 70% (the remaining 30% is coupled by rotational components), then the unobservable degree of freedom is determined to be a translational degree of freedom, and its physical direction is the vector formed by the first three elements. Instructions, for example =[0,0,1] T This indicates that the translation along the z-axis of the base coordinate system is not observable.
[0099] Determination of rotational degrees of freedom: If the magnitude of the last three elements of the eigenvector (corresponding to the rotational components of the SE(3) perturbation, i.e., the spinor part of the Lie algebra) exceeds 70%, then the unobservable degree of freedom is determined to be a rotational degree of freedom, and its rotation axis direction is the vector formed by the last three elements. Instructions, for example =[0,0,1] T This indicates that the rotation (roll motion) around the x-axis of the base coordinate system is not observable.
[0100] If there are multiple unobservable degrees of freedom, they are arranged in ascending order of eigenvalues, and the motion corresponding to the degree of freedom with the smallest eigenvalue (least information) is compensated first.
[0101] It should be noted that the core reason why the translational degrees of freedom are not observable is that all plane normals are parallel to the unobservable translational directions, resulting in the intercept component of the cross-modal consistency residual. r d =( - )- ( - )middle, t T Rn The term is insensitive to translation in unobservable directions. Circular scanning, by changing the head position, makes the changes in different frames... Rn The direction changes, thus allowing t T Rn The term carries information about the unobservable translation direction, supplementing the constraints of the intercept component. For example, when the vertical translation is unobservable, a circular scan allows the head to observe the same horizontal ground at different heights, and the ground's normal vector... n =[0,0,1] T In different frames Rn Changes in the horizontal component will introduce t z (Vertical translation) constraints.
[0102] If the unobservable degrees of freedom include translational degrees of freedom, then the robot head is controlled to perform a circular scanning motion, and the specific process is as follows: Plan a circular trajectory in a plane perpendicular to the direction of unobservable translation, with the center of the robot base as the origin. For example, if the direction of unobservable translation is the z-axis ( =[0,0,1] T If a circular trajectory with a radius of 0.3m is planned in the xy plane (which can be adjusted according to the robot's workspace, and is not specifically limited in this embodiment), the optical axis of the head camera is always oriented toward the center of the trajectory circle to ensure that the planar primitives in the scene are always within the camera's field of view.
[0103] Low-speed, stable motion with an angular velocity of 5° / s is employed to avoid image blurring or point cloud distortion caused by high-speed motion. The scanning range is a full 360° circle, ensuring that the same plane primitives are observed from different perspectives, and introducing constraint information from the translation dimension.
[0104] During the scanning process, one frame of synchronized RGB-D data and radar point cloud data is acquired every second. New data frames are added to a sliding window, while the oldest frame in the window is deleted, keeping the window size N=30 constant. After each new data frame is added, the Fisher information matrix calculation and information loss ratio assessment are immediately re-executed, and the compensation effect is monitored in real time.
[0105] When the information loss ratio drops below 0.2, or when the information loss ratio does not decrease significantly after completing a 360° circular scan (the decrease is less than 0.05), the circular scan is stopped and the robot resumes its original task execution.
[0106] The core reason for the unimaginable rotational degrees of freedom is that the head motion does not cover the missing rotational axis, resulting in the normal component of the cross-modal consistency residual. r n = [ i ] - In the middle, rotation matrix or The corresponding axial parameters cannot be observed. Triaxial scanning, by actively covering all possible rotational axes, makes the normal component... r n It responds to rotational changes along different axes, thus supplementing the constraints of the rotational dimension. For example, when rotation around the y-axis is not observable, the pitch scan changes. The pitch angle makes different frames The change in the y-component, and thus the pitch angle parameter constrained by the deviation.
[0107] If the unobservable degrees of freedom include rotational degrees of freedom, then the robot head is controlled to sequentially perform rotational scanning actions along the three axes of yaw, pitch, and roll. The specific process is as follows: The head is controlled to rotate around the z-axis of the base coordinate system, with a scanning range of -90° to +90° and an angular velocity of 5° / s, while keeping the optical axis horizontal. Yaw scanning is mainly used to supplement the constraint information of rotation around the z-axis, and is especially suitable for cases where the planar normal is concentrated in the xz or yz plane.
[0108] After the yaw scan is completed, the camera head is controlled to rotate around the y-axis of the base coordinate system, with a scanning range of -45° to +45° (to avoid excessive camera tilting up / down causing planar primitives to leave the field of view), and an angular velocity of 5° / s. The pitch scan is mainly used to supplement the constraint information of rotation around the y-axis and is suitable for cases where the planar normal is concentrated in the xy plane.
[0109] After the pitch scan is completed, the head is controlled to rotate around its own optical axis (x-axis), with a scanning range of -30° to +30° and an angular velocity of 5° / s. The roll scan is mainly used to supplement the constraint information of rotation around the x-axis and is suitable for situations where the planar normal is uniformly distributed but lacks roll direction movement.
[0110] Similar to circular scanning, each axial movement of the triaxial scan samples one frame of data per second to assess the information loss ratio in real time. If the information loss ratio drops below 0.2 after any axial scan is completed, subsequent scans are immediately stopped and task execution resumes.
[0111] It should be noted that if there are multiple unobservable degrees of freedom, including both translation and rotation, a circular scan of the translational degrees of freedom is performed first, followed by a three-axis scan of the rotational degrees of freedom. This is because unobservable translational degrees of freedom are usually caused by scene structure degradation (parallel plane normals). A circular scan can supplement information on both translation and some rotational dimensions by changing the observation perspective, making it more efficient than performing a three-axis scan alone.
[0112] The following explanation uses two specific application scenarios as examples: Scenario with missing translational degrees of freedom: A robot performs an inventory task in a warehouse environment, surrounded only by vertically placed shelves, with all planar normals parallel to the walls (horizontal direction). Fisher matrix evaluation shows that the translational degrees of freedom along the z-axis (vertical direction) are not considerable (eigenvalue λ = 0.15 < λ). min =1.0). The system triggers a circular scan: the head is controlled to rotate 360° around the base in the xy plane with an angular velocity of 5° / s. After the data collected in each frame is added to the sliding window, the information missing ratio gradually decreases. After scanning 3 times (about 216 seconds, collecting 9 frames of data), λ rises to 1.2, the information missing ratio drops to 0.08<0.2, the scan stops, and the vertical translation uncertainty of the external parameter estimation decreases from ±15mm to ±2mm.
[0113] Scenario with missing rotational degrees of freedom: Due to mechanical limitations, the robot's head can only move within a small range of pitch angles for a long time, resulting in unobservable rotational extrinsic parameters around the y-axis (pitch axis) (eigenvalue λ = 0.3 < λ). min =1.0). The system triggers the pitch scan in the three-axis scan: control the head to slowly rotate within the range of -45° to +45°, acquire 7 frames of data, λ increases to 1.5, the information missing ratio decreases from 0.22 to 0.12, the scan stops, and the estimation error of the pitch axis rotation extrinsic parameter decreases from ±0.2° to ±0.03°.
[0114] It should be noted that after the preliminary matching is completed through semantic map construction and cross-modal association in step S103 above, the resulting geometric primitive pairs still contain erroneous associations due to semantic label ambiguity and plane fitting errors. For example, a camera observation on the left wall may be incorrectly matched with a radar observation on the right wall in the map. Therefore, this application employs a residual-based RANSAC process to refine the preliminary association results. Based on the above embodiments, as an optional embodiment, obtaining geometric primitive pairs of the same semantic geometric primitive in two modalities further includes: Perform at least one round of sampling operation on each geometric primitive pair within the sliding window, select the set of interior points corresponding to the round with the highest number of interior point votes as the valid set of interior points, and delete geometric primitive pairs that do not belong to the valid set of interior points. Each round of sampling includes: At least 3 pairs of geometric primitives that are not coplanar in normal direction are randomly selected from the geometric primitive pairs of the most recent N frames as candidate primitive pairs; Based on the single-frame geometric parameters of the candidate primitive pair, the extrinsic parameters are obtained by using a rigid body transformation model to solve the extrinsic parameters in a closed loop. For each semantic geometric primitive pair of the sliding window, the candidate extrinsic parameters are used to uniformly map the single-frame geometric parameters of the geometric primitive pair in the two modalities to the base coordinate system, and the deviation between the mapped single-frame geometric parameters is determined. The geometric primitive pairs with deviations less than a preset threshold are taken as the inlier set corresponding to this round of sampling, and the geometric primitive pairs in this inlier set are recorded as the inlier votes for this round of sampling.
[0115] In this embodiment, all geometric primitive pairs that have passed the initial matching in the most recent N frames within the sliding window are collected to form a candidate set. S candidate In this embodiment, planar primitives are used as the core, therefore each element in the set is a planar pair. s k ={( ), ( ), T BH [ k ]},in( ) is the first k Planar parameters observed at the camera end, ( ) represents the planar parameters observed by the radar in the k-th frame. T BH [ k ] is the pose transformation from base to head read by the robot joint encoder in the kth frame (a known quantity).
[0116] Set the number of iterations K for RANSAC. ransac(For example, 50), the following sampling operation is performed in each iteration: Three pairs of planes are randomly selected from the candidate set, and this set is denoted as the minimum set. S min ={ s a , s b , s c}
[0117] To calculate the linear independence of the camera-side normal vectors of three pairs of planes, the specific method is as follows: { } Concatenate columns to form a matrix M = [ ], calculate the condition number of the matrix ( M) .like ( M) >10 4 (If the condition number is too large, it indicates that the three normal vectors are nearly coplanar.) In this case, discard the current sample and randomly select three more pairs of planes until the condition that the normals are not coplanar is met. It should be understood that if the normals of the three planes are coplanar, the extrinsic parameters cannot be uniquely solved (there is redundancy in degrees of freedom). For example, if all three planes are vertical walls (their normals are all parallel to the xy plane), then vertical translations and rotations around the z-axis cannot be constrained, leading to the failure of the temporary extrinsic parameter solution.
[0118] Based on minimum set S min The three pairs of planar parameters are solved using a closed-form rigid body transformation model to obtain temporary extrinsic parameters. The specific derivation is as follows: For any pair of planes s k ∈ S min Its global geometric parameters in the base coordinate system ( n B , d B Satisfying cross-modal consistency constraints:
[0119] Substitute into the plane transformation operator The above equation can be split into two components: normal and intercept. Rotational constraint (normal component): = ( k = a , b , c ) Stacking the normal constraints of three pairs of planes yields a system of linear equations concerning the rotation matrix. The initial rotation values are then solved using singular value decomposition (SVD). The matrix is constructed as follows: A =[ ] T , B =[ , , ] T Solve for the optimal rotation matrix that satisfies min|| AR - B || F SVD decomposition A T B =U ,get R = .
[0120] Substitute the obtained rotation matrix into the intercept component constraints: -( ) T +( ) T = - ( ) T ( k = a , b , c ) After simplification, we obtain the translation vector. and The system of linear equations is solved using the least squares method to obtain the temporary translation vector. and The final provisional external parameters for this round of sampling were obtained. =( )and =( ).
[0121] This embodiment utilizes the temporary extrinsic parameters obtained in this round of solving. and For the candidate set S candidate Perform consistency checks on all plane pairs and count the number of interior points: First of all, S candidate Each pair of planes in s j ={( ), ( ), T BH [ j ]}, calculate its cross-modal consistency residuals. r cons,j , r cons,j =r d,j +z×r n,j , where z For normal residual weights, normal residuals r n,j =|| - ||2, Intercept Residual r d,j =| - +( )-( - ).
[0122] In some embodiments, m + 0.1 rad, converting the angular residual to a unit of length: 0.1 rad × 0.5 m (average observation distance) = 0.05 m, then z It is 0.05.
[0123] like r cons,j < If the plane pair is true, then mark it as an interior point in this round and count it in the interior point votes. By recording the interior point votes in each round, retain the set of interior points corresponding to the round with the most interior point votes, denoted as . S inlier .
[0124] Furthermore, the candidate set S candidate China does not belong to S inlier All plane pairs are removed, only the following are retained. S inlier The interior point plane pairs in the set participate in the subsequent optimization of the sliding window factor graph. Simultaneously, the temporary extrinsic parameter corresponding to the round with the most votes in the interior point set is used. and Using these initial values as the basis for factor graph optimization can significantly accelerate the convergence speed.
[0125] This embodiment uses RANSAC random sampling and consensus voting to accurately filter out correctly associated inliers from preliminary matching results containing 30%-50% false associations. Experiments show that in scenarios with dynamic occlusion and semantic label confusion, the accuracy of inlier filtering can reach over 98%, avoiding the divergence of calibration results caused by incorrect constraint injection into the factor graph.
[0126] For example, a robot performs calibration in an office environment, where a white wall exists on the left (normal [1,0,0]). T The white wall on the right (normal [-1,0,0]) T ) and the glass door in front (normal direction [0,0,1) T Because the glass door was transparent, the radar point cloud was incomplete. The camera mistakenly associated the white wall on the left with the global ID of the white wall on the right in the radar map, forming a misassociation pair. After initial association, the candidate set contained 30 planar pairs, of which 10 were misassociated.
[0127] The RANSAC process executes 50 iterations: In one round of sampling, three pairs of correctly associated vertical planes (left wall, right wall, and ground) are randomly selected. Their normals are not coplanar, and a closed-form solution yields accurate temporary extrinsic parameters. When these extrinsic parameters are used to verify all candidate pairs, the residuals of 10 incorrectly associated pairs all exceed 0.05m (the normal deviation between the camera observation on the left wall and the radar observation on the right wall reaches 180°, with a residual of approximately 0.1m), and they are identified as outliers; the residuals of the remaining 20 correctly associated pairs are all less than 0.05m, and they are identified as inliers. These 20 correctly associated pairs are ultimately retained for optimization, successfully eliminating all incorrect associations, and the calibration result rotation error stabilizes within 0.04°.
[0128] In some embodiments, the observation constraint set includes cross-modal plane consistency constraints. In the observation constraint set optimized by the sliding window factor graph, cross-modal plane consistency constraints are a significant improvement over traditional calibration methods. By decomposing the plane consistency residual into normal and intercept components, it achieves decoupling of rotation and translation constraints for the first time, providing a theoretical basis for degradation mechanism analysis and active exploration. In this embodiment, the cross-modal plane consistency constraints are configured as follows: Using the robot base coordinate system as a reference, determine the plane normal vector of the same semantic geometric primitive in the coordinate system of each mode, and the first deviation after rotation transformation; Using the robot base coordinate system as a reference, determine the plane intercept of the same semantic geometric primitive in the coordinate system of each modality, and the second deviation after rotation transformation and translation projection; The extrinsic parameters include the parameters required for the rotation transformation and translation projection, and the residuals corresponding to the cross-modal plane consistency constraints include the first deviation and the second deviation.
[0129] For any given frame, let the camera parameters that observe the same physical plane in that frame be ( , ), radar end parameters are ( , The robot joint encoder reads the pose from the base to the head. T BH [ k ]=( R BH [ k ], t BH [ k The two sets of extrinsic parameters to be optimized are those from the head to the camera. T HC =( R HC , t HC ), base to radar T BL =( R BL , t BL ).
[0130] According to the transformation rules of plane parameters under rigid body transformation: .
[0131] The global geometric parameters of the physical plane in the base coordinate system ( n B,m , d B,m The observation constraints of both the camera and radar terminals should be met simultaneously. Camera-side transformation path: T BH [ k ] T HC ( , )=( n B,m , d B,m ) Radar-side transformation path: T BL ( , )=( n B,m , d B,m ) Therefore, by making the results of the two transformation paths equal, we can obtain the cross-modal plane consistency constraint:T BH [ k ] T HC ( , )= T BL ( , ) On the one hand, this embodiment substitutes the plane transformation operator into the normal part of the consistency constraint to obtain the normal component residual. r n,k,m = R BH [ k ] R HC - R BL .
[0132] The expression for the normal component residual shows that it does not contain any translation vector. t HC and t BL Only with rotation matrix R BH [ k (Known) R HC (To be optimized) R BL (To be optimized) and related to the plane normal vector. The normal component directly constrains the rotational portions of the two sets of extrinsic parameters. The plane normal vector is a directional quantity and is not affected by translation. Therefore, the normal vectors of the same plane observed by the camera and radar must completely coincide after being rotated to the base coordinate system by their respective extrinsic parameters. This provides a highly robust constraint for the estimation of the rotational extrinsic parameters—even if the plane intercept is deviated due to distance measurement errors, as long as the normals are aligned, the estimation of the rotational extrinsic parameters will not be affected.
[0133] On the other hand, this embodiment substitutes the plane transformation operator into the intercept part of the consistency constraint to obtain the intercept component residual. r d,k,m =( - R HC ) –( - R BL )+( R HC R HC ).
[0134] intercept components r d,k,m In this context, translation information is obtained only through scalar projection terms. t T Rn Entering the residual, among which, R HC external parameters for the camera t HC Translation after rotation normal R HC Projection on; R BL Radar extrinsic parameters t BL Translation after rotation normal R BL Projection on; R HC R HC For head translation Normal after rotation R HC R HC The projection on (a known quantity, as it is read from the encoder).
[0135] The intercept component is the signed distance of the plane to the origin of the coordinate system. Translation extrinsic parameters change this distance, which is reflected in the projection of the translation vector onto the normal direction. Therefore, the intercept residual directly reflects the estimation error of the translation extrinsic parameters. The Jacobian matrix of the intercept component residual is non-zero with respect to the translation variable, providing gradient information for the translation extrinsic parameters. However, it should be noted that the strength of the translation constraint depends on the distribution of the plane normals: if all plane normals are parallel (e.g., all wall normals are horizontal in a corridor scene), then... t T Rn It can only constrain the translation component in the normal direction; the translation component perpendicular to the normal direction cannot be observed.
[0136] In this embodiment, the normal component residual and the intercept component residual are combined into a complete cross-modal consistent residual.r cons,k,m =[ r n,k,m , r d,k,m ] T And weighted according to the sensor noise characteristics: The dimension of the normal component residual is radians, and its noise originates from the camera normal estimation error (approximately 0.05 rad) and the radar normal estimation error (approximately 0.02 rad). Therefore, the weight is set to... w n =1 / ≈18.6.
[0137] The dimension of the distance component residual is meters, and its noise originates from camera depth error (approximately 0.01m) and radar ranging error (approximately 0.01m). Therefore, the weight is set to... w d =1 / ≈70.7.
[0138] To suppress the influence of outliers, the Huber robust kernel function is applied to the weighted residuals: When | |≤ hour, =0.5 ; When | |> hour, = .
[0139] Among them, the Huber threshold corresponding to the normal component =0.1rad, Huber threshold corresponding to the intercept component =0.02m.
[0140] It is important to note that the eigenvalue decomposition of the Fisher information matrix in the above embodiment is based on the decomposed residual Jacobian matrix: the Jacobian of the normal component residuals with respect to rotational variables directly contributes to the block in the Fisher matrix corresponding to the rotational degrees of freedom; the Jacobian of the intercept component residuals with respect to translational variables directly contributes to the block in the Fisher matrix related to the translational degrees of freedom associated with the plane normal. Through this decomposition, it can be clearly identified that: if the Jacobian rank of the normal component residuals is insufficient, for example, when all plane normals are coplanar, the eigenvalues corresponding to the rotational degrees of freedom are too small, indicating that the rotational dimension is not considerable; if the eigenvalues of the intercept component residuals are insufficient... t T Rn If the partial derivative of a term with respect to a certain translation direction is zero, such as when the plane normal is perpendicular to the z-axis, then t zIf the partial derivative is zero, the eigenvalue corresponding to that translation degree of freedom is too small, indicating that the translation dimension is not observable.
[0141] For example, when the robot is running in a corridor 1.5m wide, with vertical walls on both sides (normal [1,0,0]), T and [-1,0,0] T The ground is a horizontal plane (normal [0,0,1)). T At this point, the normals of the three planes are mutually perpendicular, and the residuals of the normal components provide complete rotational constraints. The eigenvalues corresponding to the rotational degrees of freedom of the Fisher matrix are all greater than λ. min =0, the rotation dimension is completely observable. All planar normals are perpendicular to the y-axis (wall normal to the x-axis, ground normal to the z-axis), and the terms in the intercept components cannot constrain translation in the y-axis direction (i.e., the corridor width direction). Therefore, the eigenvalue corresponding to the y-axis translation degree of freedom in the Fisher matrix is only 0.1 < λ. min The information loss rate was approximately 1 / 12 ≈ 0.08 (close to the 0.2 threshold). After the system detected that the y-axis translation was not observable, it triggered active exploration: controlling the head to perform a reciprocating motion perpendicular to the y-axis (i.e., moving the head along the width of the corridor), introducing a new observation perspective. At this time, the ground normal was [0,0,1]. T The projection in the radar coordinate system changes with the position of the head along the y-axis. t BL,y y The term is no longer a constant, and the residuals of the intercept components begin to correlate with... t BL,y Sensitive, after 5 frames of data acquisition, the eigenvalue corresponding to the y-axis translation degree of freedom increased to 1.5, the information missing ratio decreased to 0, and the translation extrinsic parameter estimation error decreased from ±10mm to ±1.5mm.
[0142] It should be noted that after determining that the calibration state is stable and pausing iterative optimization through the above embodiments, the system continuously monitors the drift of external parameters. Since collisions, temperature changes, or mechanical vibrations during robot operation may cause irreversible changes in external parameters, this application introduces a drift detection mechanism based on residual mean fluctuations. In some embodiments, this embodiment further includes: The mean residual value corresponding to the observation constraint set when the calibration result of the extrinsic parameters was last updated is used as the baseline mean residual value; When the current frame iterative optimization converges, determine the fluctuation of the mean residual value corresponding to the observation constraint set relative to the mean residual value. If the increase in the mean residual value relative to the mean baseline residual value exceeds the magnitude threshold, then an external parameter drift is determined to have occurred, and iterative optimization is performed again.
[0143] In this embodiment, when the system completes its initial calibration after a cold start or when the extrinsic parameters converge after active exploration, the mean residual value corresponding to the observation constraint set at the current moment is recorded. , which serves as the baseline residual mean.
[0144] = ( ) 2 in, This represents the total number of all residual terms within the sliding window. ( )for The first moment i The residual is a single value. This baseline value represents the system fitting error when the extrinsic parameters are in a steady state.
[0145] During the calibration period when the state is stable (i.e., when iterative optimization is paused), the system continues to synchronously acquire sensor data and update the sliding window at a frequency of 30Hz (to maintain the timeliness of the data within the window), but no longer performs time-consuming nonlinear optimization. Every 1 second, the system uses the calibration results of the last updated extrinsic parameters (i.e., and ), calculate the mean of the current residuals corresponding to the observation constraint set within the current sliding window. .
[0146] = ( ) 2 It is important to note that this embodiment does not update the extrinsic parameters when calculating the residuals; it only verifies the interpretability of the old extrinsic parameters for the current new data. If the extrinsic parameters have not drifted, the old extrinsic parameters should fit the new data well and should remain within a certain range.
[0147] Calculate the increase Δ(t) in the current residual mean relative to the baseline residual mean: Δ(t)=
[0148] In some embodiments, the amplitude threshold is set to 30%. If Δ(t) > 0.3, a significant extrinsic parameter drift is determined to have occurred. Upon determining that extrinsic parameter drift has occurred, the system performs the following operations: Re-trigger the iterative optimization process of the above embodiments; If the drift is caused by slow mechanical deformation, the system can update the new residual mean after reconvergence. This enables adaptive updates to the benchmark, avoiding frequent false alarms caused by benchmark rigidity.
[0149] If, after iterative optimization, the mean residual still cannot fall back to its original value... If the deviation is within 1.3 times, it is determined to be a severe degradation or extreme drift, and the active exploration process described in the above embodiment is automatically triggered to collect new perspective data through motion to re-lock the external parameters.
[0150] In some embodiments, the condition for determining that the calibration state of the extrinsic parameters of the visual sensor is unstable includes that the number of geometric primitive pairs within the sliding window is less than a preset threshold number. .
[0151] This embodiment counts the total number of valid geometric primitive pairs within the sliding window. .like < If the calibration state is unstable, then the calibration state is determined to be unstable.
[0152] As can be seen from the foregoing embodiments, the joint optimization problem in this embodiment has a total of 12+4M degrees of freedom (M is the number of planar primitives). Each pair of planar primitives can provide 6-dimensional constraints (3-dimensional normal + 3-dimensional intercept), but due to the influence of observation noise, the actual effective constraints are approximately 4-5 dimensions per pair.
[0153] To stably solve for 12-dimensional extrinsic parameters, at least 3 pairs of non-degenerate planes are required (3 × 4 = 12-dimensional constraints). Considering the potential loss of interior points and some planes being visible only in a few frames after RANSAC, a safety margin is set: when When the value is less than 8, the rank of the constraint equation is insufficient, the external parameter solution is ambiguous, and the accuracy requirement cannot be met.
[0154] Experiments show that when When = 8, the CRLB (theoretical optimal accuracy) of the extrinsic parameter estimation begins to increase significantly; when When the value is less than 5, optimization is very likely to get trapped in local minima.
[0155] In some embodiments, the condition for determining that the calibration state of the extrinsic parameters of the vision sensor is unstable includes the fact that the included angle between the plane normals of each of all geometric primitive pairs is less than a preset angle threshold. .
[0156] This embodiment extracts the set of camera-side normal vectors of all geometric primitives to the midplane primitives within the sliding window. } (or radar end normal vector { (The two should be consistent after rotation), calculate the angle between any two normal vectors. =arccos( If all < If the calibration state is unstable, then the calibration state is determined to be unstable.
[0157] It should be noted that when the normal angle When <15°, tT Rn The contributions of different normals in the term are approximately linearly correlated, and the eigenvalues of the corresponding translation degrees of freedom in the Fisher matrix begin to decrease significantly (less than) The 15° threshold is an empirical value determined through simulation experiments in this embodiment: when When the angle is ≥15°, the observable information content of the translational degrees of freedom reaches 80% saturation; when At <15°, the information content drops sharply to below 50%, and the calibration accuracy deteriorates significantly.
[0158] In some embodiments, the sensing data of the two modes include camera images and radar point clouds; The step of extracting geometric primitives from each sensing data in the data frame includes: Point clouds in the camera coordinate system are extracted from the camera images. Candidate planes are detected by a plane detection model. Least square fitting is performed on each candidate plane to obtain the plane normal vector and plane intercept in the camera coordinate system. The RANSAC algorithm is used to detect planes in the radar point cloud to obtain the plane normal vector and plane intercept in the radar coordinate system.
[0159] For the synchronized RGB image and depth map output by the RGB-D camera, this embodiment utilizes the camera's factory-calibrated intrinsic parameters to convert each valid pixel in the depth map into a three-dimensional spatial point, obtaining a point cloud in the camera coordinate system. In some embodiments, to improve data quality, median filtering can be applied to the depth map to remove isolated noise points and discard pixels that are too far away (e.g., greater than 3 meters) or have invalid depth, retaining only the valid observation data of the robot's near field.
[0160] Next, the RGB image is input into a pre-trained planar detection network. This network can combine semantic information such as texture and color of the image to automatically identify and segment different types of planar regions such as "ground", "wall", and "desktop", and output the pixel mask of each plane.
[0161] Compared to traditional purely geometric methods, the semantic-assisted approach in this embodiment can effectively handle complex scenarios: for example, even if the table is cluttered, the network can accurately segment the table area using texture cues along the edge of the table, significantly improving the robustness of planar detection in occluded environments. To meet the speed requirements for online operation, this embodiment uses TensorRT to accelerate the network, keeping the single-frame inference time within 30 milliseconds.
[0162] For each segmented plane mask, all corresponding 3D points are further extracted, and the mathematical parameters of the plane are fitted using the least squares method, including the normal vector indicating the orientation of the plane and the intercept indicating the distance from the plane to the camera origin.
[0163] In some embodiments, to improve fitting accuracy, points with closer depths are given higher weights (because closer depth measurements are more accurate), and candidate planes with excessively large average distances from fitted points to the plane are discarded, ensuring that the quality of the observation data input for subsequent optimization is reliable.
[0164] For the 3D point cloud output by LiDAR, this embodiment first downsamples the original LiDAR point cloud, significantly reducing the number of points while preserving the geometric features of large planar surfaces such as walls and the ground, thus improving processing speed. Simultaneously, data from the robot's built-in IMU and odometry are used to compensate for point cloud stretching distortion caused by robot motion, ensuring spatial consistency of the point cloud in each frame. Next, the Random Sample Consensus (RANSAC) algorithm is used to iteratively detect planes. The RANSAC algorithm randomly selects a small number of points to assume a planar model, and then counts how many points conform to this model. After thousands of iterations, the planes with the highest number of interior points are selected as candidates.
[0165] In some embodiments, to avoid missing small-sized planar structures (such as door panels and displays), this embodiment adopts a multi-scale detection strategy: first, a larger voxel mesh is used to detect large planes such as the ground and walls, and then fine downsampling is used to detect small planes such as desktops. At the same time, the normal vector is constrained by combining prior knowledge of the indoor scene (such as the ground is usually horizontal and the walls are usually vertical) to reduce false detections of unexpected planes such as ceilings and slopes.
[0166] For each candidate plane detected by RANSAC, all its interior points are extracted, and the plane's normal vector and intercept are accurately fitted again using the least squares method. Since radar ranging noise increases with distance, points at greater distances are assigned lower weights during fitting. If the proportion of interior points on a candidate plane is too low (e.g., less than 70%), it is considered a failed fit and discarded.
[0167] After the above processing, the system obtained high-quality planar parameters in both the camera and radar coordinate systems. These parameters will serve as the basis for subsequent cross-modal correlation: theoretically, through the extrinsic parameter relationships to be optimized, the same physical plane (like a wall) in the two coordinate systems can be uniformly transformed into a common reference coordinate system, thereby achieving joint calibration of extrinsic parameters.
[0168] In some embodiments, associating the geometric primitives of two modalities in the data frame according to the environment map includes: Each geometric primitive is back-projected onto the base coordinate system and matched with the semantic geometric primitives in the environment map. The matching condition is that the deviation of the geometric parameters is less than a preset association threshold. Geometric primitives of two modalities that are matched by the same semantic identifier are combined into a geometric primitive pair that shares the same semantic geometric primitive.
[0169] This application constructs a globally consistent environment map using semantic SLAM, achieving high-precision automatic association under the dual constraints of unique semantic identifiers and consistent geometric parameters. The specific implementation steps are as follows: S301: Semantic SLAM builds a global environment map This embodiment uses the ORB-SLAM3 semantic SLAM system to build and maintain a global environment map. The map uses the base coordinate system at the moment the robot is powered on as the world coordinate system, and the parameters of all map elements are represented in the base coordinate system {B} to ensure spatiotemporal consistency.
[0170] The map stores all observed semantic geometric primitives, each primitive containing: Unique semantic identifier ID: such as “plane_001” (ground), “plane_002” (left wall), “edge_003” (table edge), is automatically assigned by the SLAM system based on the semantic segmentation results. The ID of the same physical structure remains unchanged throughout its life cycle.
[0171] Global geometric parameters: the planar primitive is ( n B , d B (Normal vector + intercept), edge line primitives are Plücker coordinates ( m B , v B All of these are represented in the base coordinate system {B}.
[0172] Observation history: Records in which keyframes the primitive was observed, the corresponding sensor mode, and the geometric parameters at that time, for use in subsequent prior factor construction.
[0173] In this embodiment, whenever a data frame is inserted into the SLAM system, semantic segmentation and geometric primitive detection are performed on the new data frame. The newly detected primitives are matched and associated with existing primitives in the map. If the match is successful, the global geometric parameters of the primitive are updated (by weighted average of observations from multiple frames). If the match fails, a new semantic ID is assigned to the primitive and it is inserted into the map.
[0174] Step S302: Back-project the geometric primitives to the base coordinate system For all geometric primitives (camera and radar) detected in the current frame within the sliding window, they are back-projected to the base coordinate system {B} through their respective transformation chains to obtain temporary global geometric parameter estimates: For the planar primitives detected by the camera ( , (The m-th plane in the k-th frame), using the known pose of the base to the head position. T BH [ k ]=( R BH [ k ], t BH [ k ]) and head-to-camera extrinsics to be optimized T HC =( R HC , t HC The temporary geometric parameters of the base coordinate system are calculated using the plane transformation operator. , ).
[0175] For the planar primitives detected by the radar end ( , (The m-th plane in the k-th frame), using the base-to-radar extrinsic parameters to be optimized T BL =( R BL , t BL ), calculate its temporary geometric parameters in the base coordinate system ( , ).
[0176] Step S303: Dual matching based on semantic ID and geometric parameters The temporary geometric parameters obtained by backprojection are matched with the existing primitives in the semantic map. The matching must satisfy both semantic consistency and geometric consistency. First, a semantic segmentation network (such as Plane-R-CNN) is used to predict semantic labels (e.g., ground, wall, desktop) for each detected planar primitive. If the camera-side plane... m If the semantic label of a camera plane is "ground", then subsequent geometric matching will only be performed with primitives in the map that also have the semantic label "ground", significantly narrowing down the candidate matching range. If the match is successful, the camera plane inherits the semantic ID of the corresponding primitive in the map (such as "plane_001"); if the match fails, a new semantic ID is assigned to the plane and it is inserted into the map.
[0177] For a pair of candidate primitives that pass semantic consistency matching (such as "ground" on the camera and "plane_001" on the map), calculate the deviation of their geometric parameters and determine whether it is less than a preset association threshold: Normal deviation Δ n =arccos(( ) T ),in This is the global normal vector for "plane_001" on the map. For example, the threshold is set to 0.1 rad.
[0178] Intercept deviation Δ d =| - |, among which This is the global intercept for "plane_001" on the map. For example, the threshold is set to 0.05m.
[0179] If Δ n <0.1rad and Δ d If the value is less than 0.05m, the geometric match is considered successful.
[0180] It should be understood that, through the radar end plane n Perform the exact same semantic and geometric matching process, associating it with semantic primitives in the map.
[0181] Step S304: Construction of bimodal primitive pairs with the same semantic identifier When the camera end element m and radar terminal elements n When two elements are associated with the same semantic geometric primitive in the map (i.e., have the same semantic ID, such as both being "plane_001"), they are considered as a geometric primitive pair sharing the same semantic geometric primitive. Pair p ={( , ),( , ), T BH [ k ID p} ID p This is a shared semantic ID (e.g., "plane_001"). This primitive pair is the core input for constructing the cross-modal consistency factor and RANSAC false association removal.
[0182] In some embodiments, if the camera detects two "ground" planes in the current frame (e.g., due to occlusion segmentation errors), but there is only one "plane_001" in the map, then the one with the smallest geometric deviation from the map is selected as the valid association, and the other is determined to be a false detection and discarded.
[0183] In some embodiments, if there are two "wall" primitives in the map and the camera only detects one in the current frame, then only the closest one is associated, and the other wall is marked as "temporarily invisible".
[0184] Step S305: Validation and optimization of association results and initial value transfer To further improve the robustness of the association, the following verification steps are added in this embodiment: 1. Two-way matching verification: Not only are the primitives of the current frame matched to the map, but the primitives of the map are also back-projected to the coordinate system of the current frame to verify whether the two-way deviations are both less than the threshold, thus avoiding one-way matching errors.
[0185] 2. Timing consistency check: If the primitive is stably associated in 5 consecutive historical frames (matching success rate > 80%), it is determined to be a reliable association and its weight is increased; if it is an isolated match that appears for the first time, its weight is reduced to prevent short-term false detection from interfering with optimization.
[0186] 3. Initial value propagation: After a successful match, the global geometric parameters of the semantic primitive in the map are passed ( As the observed values of the prior factors of the plane parameters, they provide strong prior constraints for factor graph optimization and accelerate the convergence speed.
[0187] In traditional single-frame calibration methods, the geometric primitive parameters of each frame are treated as independent variables, ignoring the inherent consistency of the same physical structure across multiple frames. This embodiment introduces the global geometric parameters of semantic geometric primitives as shared variables across frames, incorporating them into the set of variables to be optimized in the factor graph. Through multi-frame joint optimization, a mutually beneficial symbiosis of extrinsic parameter estimation and environmental map refinement is achieved. In some embodiments, the variables to be optimized further include the global geometric parameters of each semantic geometric primitive in the base coordinate system; the iterative optimization includes: If it is determined that the same semantic geometric primitive exists in multiple consecutive data frames, the global geometric parameters of the semantic geometric primitive are used as shared variables to determine the residuals of the multiple consecutive data frames; The global geometric parameters of the semantic geometric primitives are updated through iterative optimization of multiple data frames.
[0188] Taking a planar primitive as an example, suppose there are a total of M Each independent semantic plane (whose unique ID has been determined through the association steps in the above embodiments) is defined as follows: The global geometric parameters of each plane in the base coordinate system are defined as follows: ={( n B,1 , d B,1 ), ( n B,2 , d B,2 ),…, ( n B,M , d B,M )} in,n B,M ∈S 2 (Unit normal vector) d B,M ∈R (intercept).
[0189] Therefore, the complete vector of variables to be optimized is: =[ T HC , T BL ,( n B,1 , d B,1 ), ( n B,2 , d B,2 ),…, ( n B,M , d B,M )] For any frame within the sliding window k If a certain geometric primitive (such as the camera end plane) is associated with a semantic ID of p In the global plane, when constructing the residual factor for that frame, the observations of that frame are no longer treated as independent variables, but rather combined with the global geometric parameters ( n B,p , d B, p The difference is taken as the residual.
[0190] Taking the cross-modal consistency factor as an example, its residual calculation logic has undergone a fundamental change: Traditional method: Residual measures whether the transformed camera observations are consistent with the transformed radar observations.
[0191] r cons,k,p = T BH [ k ] T HC ( )- T BL ( ) In this embodiment, the residual measures whether the camera observation transformation or radar observation transformation is consistent with the global geometric parameters.
[0192] =T BH [ k ] T HC ( )- ( n B,p , d B,p ) =T BL ( )- ( n B,p , d B,p ) at this time,( n B,p ,d B,p ) is the connection of the first k Frame and all other frames containing that plane (such as the first frame) k -1, k The hub (+1 frame). As long as the observed ID is... p The observations of a plane, after being transformed by extrinsic parameters, must converge to the same global geometric parameter point.
[0193] This design enables global geometric parameters ( n B,p , d B,p It becomes a true shared variable: it has only one instance in the factor graph, but has multiple edges connected to the residual factors of different frames.
[0194] If the semantic ID is p If the plane is stably observed in L consecutive frames within the sliding window, then all the relevant residuals of these frames will jointly constrain the global geometric parameter: In this embodiment, for each frame in L frames, corresponding camera-side residuals and radar-side residuals are constructed. These residuals collectively constitute a set of constraints for the global geometric parameters. During the nonlinear optimization process of the factor graph, both extrinsic parameters and global geometric parameters are adjusted simultaneously to minimize the sum of squares of all residuals. When the optimization converges, the global geometric parameters are updated, and their accuracy is significantly higher than the fitting results of any single frame. The updated parameters are fed back to the semantic SLAM map as the benchmark for the next round of association, forming a closed loop.
[0195] In some embodiments, to further accelerate convergence and prevent drift during the optimization process, a plane parameter prior factor can be added to the factor graph when a certain global geometric parameter has accumulated a high confidence level through multiple frames of observation.
[0196] Suppose that the coarse global geometric parameters of plane p have been estimated using historical data (historical frames outside the sliding window). , ) and its covariance matrix Σprior (Reflecting the uncertainty of the estimate), the prior residual is defined as: r prior,p =
[0197] in It is a constructed rotation matrix. This is a Lie algebra mapping. The weights of the residual are given by... Decide.
[0198] The technical effects of the embodiments of this application are as follows: I. Calibration accuracy is significantly improved Based on CRLB theoretical analysis and experimental verification, S 3 -Calib compared to the KSAC sphere target method: (1) Rotational RMSE decreased from 0.080° to 0.045° (44% improvement); (2) Translational RMSE decreased from 3.5mm to 1.8mm (49% improvement); (3) Rotational RMSE decreased from 0.15° to 0.08° (47% improvement); (4) Translational RMSE decreased from 8.0mm to 4.2mm (a 48% improvement). The fundamental reason for the improved accuracy is that planar observations provide 6-DOF constraints (vs. 4-DOF for spheres), increasing the number of effective constraints per frame from 4-5 to 6-8; the planar parameters are obtained from statistical estimation of multiple points, making them significantly less sensitive to LiDAR single-point noise than single-point sphere center constraints.
[0199] II. Completely target-independent, deployment requires zero preparation. Traditional methods (KSAC requires ~15 minutes to place the ball and measure the radius, Kalibr requires ~20 minutes to prepare the checkerboard and align it precisely) require manual preparation and maintenance of the calibration target. 3 - Calib utilizes naturally occurring semantic geometry structures such as walls, floors, and columns in the scene as calibration references, eliminating deployment preparation time—calibration automatically begins after the robot starts. This advantage is particularly crucial in dynamically deployed scenarios (such as inspection robots that need to frequently change their working environment).
[0200] III. Online lifetime calibration, maintenance-free Traditional calibration methods involve collecting data once and then optimizing offline; freezing external parameters prevents the detection of drift during operation. 3Calib's sliding window factor graph mechanism updates extrinsic parameters while executing tasks throughout the robot's lifecycle. Optimization is paused upon convergence to conserve computational resources, and automatically resumed when significant changes are detected (collisions, temperature changes, or resumption of motion after prolonged inactivity). In practice, each iteration takes 50-100ms, and continuous calibration does not affect task execution.
[0201] IV. Proactively explore ways to improve sample efficiency and robustness Fisher's information-driven active exploration strategy reduces the number of frames required for calibration from 12-15 frames (KSAC passive sampling) to 8-12 frames (a 25-33% improvement in sample efficiency), and can actively compensate for undesirable dimensions, avoiding degradation—traditional methods fail to calibrate when the plane is parallel or head movement is restricted, while S³-Calib automatically detects and plans compensating motion. In a 24-hour long-term test, the expected rotational drift is <0.1° and translational drift is <2cm, with accuracy recovering after 2-3 active exploration triggers.
[0202] V. Robustness Enhancement (1) Huber robust kernel suppresses the influence of out-of-field (LiDAR δ=0.02m, camera δ=1.0px). (2) RANSAC erroneous association removal to prevent erroneous constraint injection optimization (50 iterations, 0.05m+0.1rad threshold). (3) Multi-plane weighted fusion is insensitive to the fit of out-of-points on a single plane; (4) It can still be effectively calibrated when the planar part is visible.
[0203] Robustness testing is expected to achieve a plane detection success rate of ≥95% under scenarios with strong light, dynamic occlusion, sensor contamination, and extreme temperatures of 0-50°C, with a calibration convergence time increase of ≤1.5× and a final accuracy loss of ≤30%.
[0204] Please see Figure 4 This document exemplifies an online joint calibration system for visual-radar extrinsic parameters of a robot platform, based on an embodiment of this application. This system eliminates traditional manual targets, utilizing naturally occurring semantic geometric structures in the environment (such as walls, floors, and columns). Through a closed-loop architecture of "perception-optimization-decision," it achieves intervention-free calibration throughout the entire lifecycle. As shown in the figure, it mainly includes a semantic structure perception module, a sliding window factor graph optimization module, and an active exploration and drift monitoring module. These modules work collaboratively, and the specific implementation process is as follows: The semantic structure perception module is responsible for extracting stable geometric features from heterogeneous sensor data and establishing cross-modal semantic connections, serving as the data entry point for the entire system.
[0205] The semantic structure perception module synchronously accesses color images and depth streams from an RGB-D camera, as well as 3D point cloud streams from a LiDAR scanner. Simultaneously, it accesses pose data from the robot's joint encoders to obtain the real-time transformation relationship from the base to the head.
[0206] For the camera side, a deep learning-based Plane R-CNN network is used to process RGB-D data. Combined with semantic segmentation results, planar regions (such as the ground, tabletops, and walls) in the scene are identified, and the plane normal vector and intercept in the camera coordinate system are fitted. Compared to traditional geometric methods, the deep learning model can effectively overcome the interference of texture loss and lighting variations.
[0207] For the radar side, after voxel downsampling of the lidar point cloud, the RANSAC algorithm is used for plane segmentation to fit the plane parameters in the radar coordinate system. The radar data provides high-precision geometric depth information, compensating for the shortcomings of depth cameras at long distances and in strong light.
[0208] The semantic structure awareness module incorporates the ORB-SLAM3 semantic SLAM engine, maintaining a global semantic map. This map not only contains geometric structures but also assigns a unique semantic ID to each detected physical plane. When new frame data arrives, the module backprojects the planes detected in the current frame onto the base coordinate system and matches them with historical planes in the map. The matching strategy follows the principle of semantic ID consistency and geometric parameter deviation within a threshold, thereby automatically establishing the correspondence between camera observation, radar observation, and physical plane without manual specification.
[0209] The sliding window factor graph optimization module is the core computing engine, responsible for estimating and optimizing extrinsic parameters in real time. It uses a sliding window mechanism to ensure the real-time performance and long-term stability of the computation.
[0210] The sliding window factor graph optimization module constructs a joint optimization problem involving four types of residual factors: Radar point distance factor: Constrains the distance between points in the radar point cloud and its fitted plane to ensure the internal consistency of radar observations.
[0211] Camera reprojection factor: constrains the distance from image pixels to the camera fitting plane profile, ensuring the internal consistency of camera observations.
[0212] Cross-modal consistency factor: This factor unifies camera and radar observations on the same physical plane to the base coordinate system, forcing them to align. Specifically, this factor decomposes the residuals into normal and intercept components. The former is only sensitive to rotational extrinsic parameters, while the latter carries translational extrinsic parameter information, thereby achieving decoupling of rotational and translational constraints.
[0213] Planar prior factors: Utilizing existing historical plane parameters in the semantic map as prior knowledge, providing stable anchor points for optimization and preventing degradation.
[0214] The sliding window factor graph optimization module maintains a sliding window containing the most recent 30 frames of data. When a new data frame arrives, older frames are marginalized and removed. This design keeps the computational complexity within a constant range (50-100ms / frame), meeting the requirements for online operation.
[0215] Ceres Solver was used to perform nonlinear least-squares optimization on the factor map. The variables to be optimized included two sets of extrinsic parameters (head-to-camera, pedestal-to-radar) and global parameters of all active semantic planes. Notably, the global parameters of the planes exist as shared variables within the window. Observations of the same plane in multiple frames jointly constrain this single variable, thereby achieving multi-frame information fusion and significantly improving the accuracy and robustness of parameter estimation.
[0216] The active exploration and drift monitoring module is responsible for monitoring the health of the calibration status and intervening when necessary to restore or improve calibration accuracy, thereby achieving system autonomy.
[0217] Fisher information-driven observability analysis: This module calculates the Fisher information matrix for the current optimization problem in real time and performs eigenvalue decomposition. The eigenvalues of the Fisher matrix reflect the observable information content of each degree of freedom. In this embodiment, the information missing ratio is defined as the proportion of eigenvalues less than a threshold relative to the total dimension.
[0218] When the proportion of missing information exceeds a preset threshold, the system determines that the currently observed geometry has degraded. At this point, the module automatically takes over the robot's head motion control. If the translation dimension is missing, control the head to perform a circular scan and change the observation perspective to introduce new geometric constraints.
[0219] If the rotation dimension is missing, control the head to perform a three-axis scan (Pan / Tilt / Roll) to stimulate rotational motion in each axis.
[0220] During the exploration process, new data is collected at a low speed and fed back to the optimization module in real time until the information loss rate drops below the threshold, after which control is automatically returned to the main task.
[0221] During long-term stable operation, this module continuously monitors the overall residual mean of the factor plot. If the residual mean rises by more than 30% relative to the baseline (possibly due to collisions, mechanical vibrations, or temperature drift), the system determines that external parameter drift has occurred and immediately wakes up the optimization module for recalibration, without the need for manual intervention.
[0222] After the calibration system of this embodiment is started, the semantic perception module continuously extracts and associates planar features from the environment. The optimization module uses these features to perform real-time optimization within a sliding window, outputting high-precision extrinsic parameter results. Simultaneously, the decision module monitors the system's health: if it detects a decline in data quality or parameter accuracy, it automatically triggers the corresponding repair mechanism. The optimized high-precision planar parameters are also fed back to the semantic map, continuously refining it, thus forming a virtuous cycle where perception supports calibration, and calibration enhances perception.
[0223] Through the modular design described above, this system achieves a fundamental shift from "relying on manual targets" to "utilizing the natural environment" and from "one-time offline calibration" to "online calibration throughout the entire lifecycle," making it particularly suitable for intelligent agents such as service robots and autonomous vehicles that operate in complex and dynamic environments for extended periods.
[0224] This application provides an extrinsic parameter calibration device for a visual sensor, such as... Figure 5 As shown, the device may include: a data acquisition module 501, a parameter extraction module 502, a map construction module 503, an association module 504, and an iterative optimization module 505, wherein, The data acquisition module 501 is used to acquire sensor data of two modes synchronously collected by the robot's vision sensor at a preset frame rate, combine the sensor data of the two modes collected in the same frame into a data frame, and construct a sliding window, wherein the sliding window includes the latest N data frames. The parameter extraction module 502 is used to extract the geometric primitives in each sensing data in each data frame within the sliding window, as well as the single-frame geometric parameters of the geometric primitives in the coordinate system of the corresponding mode. The map building module 503 is used to build an environment map based on the data frames in the sliding window. The environment map includes each semantic geometric primitive and the global geometric parameters of the semantic geometric primitive in the robot's base coordinate system. The semantic geometric primitive is a geometric primitive with a semantic identifier. The semantic identifier is used to uniquely indicate the physical structure corresponding to the corresponding geometric primitive. The association module 504 is used to associate the geometric primitives of two modalities in each data frame within the sliding window according to the environment map, so as to obtain the geometric primitive pairs of the same semantic geometric primitives in the two modalities. The iterative optimization module 505 is used to determine the calibration result of the external parameters in the current frame if it is determined that the calibration state of the external parameters of the vision sensor is unstable. The module uses the base coordinate system as the world coordinate system, the most recently updated external parameters as the initial values for iterative optimization in the current frame, and the external parameters and the single-frame geometric parameters of each geometric primitive pair in the sliding window as variables to be optimized. Iterative optimization is performed by minimizing the observation constraint set constructed based on geometric consistency to determine the calibration result of the external parameters in the current frame.
[0225] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.
[0226] Based on the above embodiments, as an optional embodiment, the device further includes a compensation module, which is used for: During the iterative optimization, the Jacobian matrix of each residual corresponding to the observation constraint set with respect to the variable to be optimized is determined, and the Fisher information matrix is determined based on the Jacobian matrix; each eigenvalue of the Fisher information matrix represents the amount of observable information in the direction of the corresponding degree of freedom, and each eigenvector represents the direction of the corresponding degree of freedom in the robot base coordinate system; Count the number of feature values that are less than a preset feature value threshold, and calculate the information missing ratio; If the information missing ratio is greater than a preset information missing ratio threshold, then the feature vector corresponding to the feature value that is less than the preset threshold is determined as the direction of unobservable degrees of freedom. The robot is controlled to perform compensating motion according to the direction of the unobservable degrees of freedom. The data frames collected during the compensating motion are moved into the sliding window to replace the old data frames. The iterative optimization is re-executed according to the updated sliding window until the information missing ratio is not greater than the preset threshold. The compensating motion is then stopped, and a new calibration result of the extrinsic parameters is obtained.
[0227] Based on the above embodiments, as an optional embodiment, the unobservable degrees of freedom include at least one of rotational degrees of freedom and translational degrees of freedom; The step of controlling the robot to perform compensatory motion according to the direction of the unobservable degrees of freedom includes: If the unobservable degrees of freedom include translational degrees of freedom, then control the robot to perform a circular scanning action; If the unobservable degrees of freedom include rotational degrees of freedom, then the robot is controlled to sequentially perform scanning actions involving yaw, pitch, and roll rotations along the three axes.
[0228] Based on the above embodiments, as an optional embodiment, obtaining geometric primitive pairs of the same semantic geometric primitive in two modalities further includes: Perform at least one round of sampling operation on each geometric primitive pair within the sliding window, select the set of interior points corresponding to the round with the highest number of interior point votes as the valid set of interior points, and delete geometric primitive pairs that do not belong to the valid set of interior points. Each round of sampling includes: At least 3 pairs of geometric primitives that are not coplanar in normal direction are randomly selected from the geometric primitive pairs of the most recent N frames as candidate primitive pairs; Based on the single-frame geometric parameters of the candidate primitive pair, the extrinsic parameters are obtained by using a rigid body transformation model to solve the extrinsic parameters in a closed loop. For each semantic geometric primitive pair of the sliding window, the candidate extrinsic parameters are used to uniformly map the single-frame geometric parameters of the geometric primitive pair in the two modalities to the base coordinate system, and the deviation between the mapped single-frame geometric parameters is determined. The geometric primitive pairs with deviations less than a preset threshold are taken as the inlier set corresponding to this round of sampling, and the geometric primitive pairs in this inlier set are recorded as the inlier votes for this round of sampling.
[0229] Based on the above embodiments, as an optional embodiment, the geometric primitive includes a plane, and the geometric parameters corresponding to the plane include the plane normal vector and the plane intercept; The observation constraint set includes cross-modal plane consistency constraints, which are configured as follows: Using the robot base coordinate system as a reference, determine the plane normal vector of the same semantic geometric primitive in the coordinate system of each mode, and the first deviation after rotation transformation; Using the robot base coordinate system as a reference, determine the plane intercept of the same semantic geometric primitive in the coordinate system of each modality, and the second deviation after rotation transformation and translation projection; The extrinsic parameters include the parameters required for the rotation transformation and translation projection, and the residuals corresponding to the cross-modal plane consistency constraints include the first deviation and the second deviation.
[0230] Based on the above embodiments, as an optional embodiment, the device further includes an offset detection module, used for: The mean residual value corresponding to the observation constraint set when the calibration result of the extrinsic parameters was last updated is used as the baseline mean residual value; When the current frame iterative optimization converges, determine the fluctuation of the mean residual value corresponding to the observation constraint set relative to the mean residual value. If the increase in the mean residual value relative to the mean baseline residual value exceeds the magnitude threshold, then an external parameter drift is determined to have occurred, and iterative optimization is performed again.
[0231] Based on the above embodiments, as an optional embodiment, the conditions under which the calibration state of the extrinsic parameters of the visual sensor is determined to be unstable include at least one of the following: The number of geometric primitive pairs within the sliding window is less than a preset threshold. The included angle of each plane normal in all geometric primitive pairs is less than a preset angle threshold.
[0232] Based on the above embodiments, as an optional embodiment, the sensing data of the two modes include camera images and radar point clouds; The step of extracting geometric primitives from each sensing data in the data frame includes: Point clouds in the camera coordinate system are extracted from the camera images. Candidate planes are detected by a plane detection model. Least square fitting is performed on each candidate plane to obtain the plane normal vector and plane intercept in the camera coordinate system. The radar point cloud is subjected to a random sampling consensus algorithm to detect the plane, and the plane normal vector and plane intercept in the radar coordinate system are obtained.
[0233] Based on the above embodiments, as an optional embodiment, the step of associating the geometric primitives of two modalities in the data frame according to the environment map includes: Each geometric primitive is back-projected onto the base coordinate system and matched with the semantic geometric primitives in the environment map. The matching condition is that the deviation of the geometric parameters is less than a preset association threshold. Geometric primitives of two modalities that are matched by the same semantic identifier are combined into a geometric primitive pair that shares the same semantic geometric primitive.
[0234] Based on the above embodiments, as an optional embodiment, the sensing data of the two modes include camera images and radar point clouds; The observation constraint set also includes: Point cloud planar distance constraints are used to constrain the distance from points belonging to the same plane to the fitted plane of that plane in the radar coordinate system. Camera plane reprojection constraint is used to constrain the distance from image plane pixels to the contour line obtained by backprojecting the plane in the camera coordinate system; The prior constraints on plane parameters are used to constrain the prior estimates and covariance of the global plane parameters of the semantic geometric primitives in the base coordinate system.
[0235] Based on the above embodiments, as an optional embodiment, the variable to be optimized also includes the global geometric parameters of each semantic geometric primitive in the base coordinate system; The iterative optimization includes: If it is determined that the same semantic geometric primitive exists in multiple consecutive data frames, the global geometric parameters of the semantic geometric primitive are used as shared variables to determine the residuals of the multiple consecutive data frames; The global geometric parameters of the semantic geometric primitives are updated through iterative optimization of multiple data frames.
[0236] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of a visual sensor extrinsic parameter calibration method. Compared with related technologies, this embodiment can achieve the following: This embodiment dynamically maintains the bimodal sensor data synchronously acquired in the most recent N frames through a sliding window, avoiding the computational power explosion caused by the infinite accumulation of historical data while ensuring the information redundancy required for optimization; by extracting the geometric primitives of each frame's bimodal sensor and combining them with semantic SLAM to construct a global environment map with unique semantic identifiers, it utilizes the naturally existing stable structures in the scene. Using the structure as a calibration reference, it eliminates the reliance on manual calibration targets; it automatically completes cross-modal geometric primitive association through semantic identification, saving the tedious process of manually labeling correspondences; it only uses the most recent extrinsic parameters as initial values when the calibration state is unstable, and jointly optimizes the single-frame geometric parameters of each geometric primitive pair within the sliding window with the extrinsic parameters. This avoids the ineffective consumption of computing power in the stable state, and suppresses single-frame noise interference through the statistical characteristics of multi-frame observations. It takes into account the high accuracy, low latency and low power consumption of online calibration, and solves the problems of traditional one-time offline calibration that cannot detect extrinsic parameter drift during operation and the high cost of relying on manual target deployment.
[0237] In one alternative embodiment, an electronic device is provided, such as Figure 6 As shown, Figure 6 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0238] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0239] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, bus 4002 is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus.
[0240] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0241] The memory 4003 is used to store computer programs that execute the embodiments of this application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer programs stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0242] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.
[0243] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0244] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the illustrations or text descriptions.
[0245] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.
[0246] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A method for calibrating the extrinsic parameters of a vision sensor, characterized in that, include: The robot's vision sensor synchronously collects sensor data from two modalities at a preset frame rate. The sensor data from the two modalities collected in the same frame are combined into a data frame, and a sliding window is constructed. The sliding window includes the latest N data frames. For each data frame within the sliding window, extract the geometric primitives in each sensing data in the data frame and the single-frame geometric parameters of the geometric primitives in the coordinate system of the corresponding mode; An environment map is constructed based on the data frames within the sliding window. The environment map includes each semantic geometric primitive and the global geometric parameters of the semantic geometric primitive in the robot's base coordinate system. The semantic geometric primitive is a geometric primitive with a semantic identifier. The semantic identifier is used to uniquely indicate the physical structure corresponding to the corresponding geometric primitive. For each data frame within the sliding window, the geometric primitives of the two modalities in the data frame are associated according to the environment map to obtain geometric primitive pairs of the same semantic geometric primitive in the two modalities; If it is determined that the calibration state of the extrinsic parameters of the vision sensor is unstable, then the base coordinate system is used as the world coordinate system, the most recently updated extrinsic parameters are used as the initial values for the current frame iterative optimization, the extrinsic parameters and the single-frame geometric parameters of each geometric primitive pair within the sliding window are used as variables to be optimized, and iterative optimization is performed by minimizing the observation constraint set constructed based on geometric consistency to determine the calibration result of the extrinsic parameters in the current frame.
2. The method according to claim 1, characterized in that, Also includes: Determine the Jacobian matrix of each residual pair of the observation constraint set corresponding to the variable to be optimized during the iterative optimization, and determine the Fisher information matrix based on the Jacobian matrix; Each eigenvalue of the Fisher information matrix represents the amount of observable information in the direction of the corresponding degree of freedom, and each eigenvector represents the direction of the corresponding degree of freedom in the robot base coordinate system; Count the number of feature values that are less than a preset feature value threshold, and calculate the information missing ratio; If the information missing ratio is greater than a preset information missing ratio threshold, then the feature vector corresponding to the feature value that is less than the preset threshold is determined as the direction of unobservable degrees of freedom. The robot is controlled to perform compensating motion according to the direction of the unobservable degrees of freedom. The data frames collected during the compensating motion are moved into the sliding window to replace the old data frames. The iterative optimization is re-executed according to the updated sliding window until the information missing ratio is not greater than the preset threshold. The compensating motion is then stopped, and a new calibration result of the extrinsic parameters is obtained.
3. The method according to claim 2, characterized in that, The unobservable degrees of freedom include at least one of rotational degrees of freedom and translational degrees of freedom; The step of controlling the robot to perform compensatory motion according to the direction of the unobservable degrees of freedom includes: If the unobservable degrees of freedom include translational degrees of freedom, then control the robot to perform a circular scanning action; If the unobservable degrees of freedom include rotational degrees of freedom, then the robot is controlled to sequentially perform scanning actions involving yaw, pitch, and roll rotations along the three axes.
4. The method according to claim 1, characterized in that, The method of obtaining geometric primitive pairs of the same semantic geometric primitive in two modalities also includes: Perform at least one round of sampling operation on each geometric primitive pair within the sliding window, select the set of interior points corresponding to the round with the highest number of interior point votes as the valid set of interior points, and delete geometric primitive pairs that do not belong to the valid set of interior points. Each round of sampling includes: At least 3 pairs of geometric primitives that are not coplanar in normal direction are randomly selected from the geometric primitive pairs of the most recent N frames as candidate primitive pairs; Based on the single-frame geometric parameters of the candidate primitive pairs, the extrinsic parameters are obtained by using a rigid body transformation model to solve for them in a closed loop. For each semantic geometric primitive pair of the sliding window, the candidate extrinsic parameters are used to uniformly map the single-frame geometric parameters of the geometric primitive pair in the two modalities to the base coordinate system, and the deviation between the mapped single-frame geometric parameters is determined. The geometric primitive pairs with deviations less than a preset threshold are taken as the inlier set corresponding to this round of sampling, and the geometric primitive pairs in this inlier set are recorded as the inlier votes for this round of sampling.
5. The method according to claim 2, characterized in that, The geometric primitives include planes, and the geometric parameters corresponding to the planes include plane normals and plane intercepts; The observation constraint set includes cross-modal plane consistency constraints, which are configured as follows: Using the robot base coordinate system as a reference, determine the plane normal vector of the same semantic geometric primitive in the coordinate system of each mode, and the first deviation after rotation transformation; Using the robot base coordinate system as a reference, determine the plane intercept of the same semantic geometric primitive in the coordinate system of each modality, and the second deviation after rotation transformation and translation projection; The extrinsic parameters include the parameters required for the rotation transformation and translation projection, and the residuals corresponding to the cross-modal plane consistency constraints include the first deviation and the second deviation.
6. The method according to claim 1, characterized in that, Also includes: The mean residual value corresponding to the observation constraint set when the calibration result of the extrinsic parameters was last updated is used as the baseline mean residual value; When the current frame iterative optimization converges, determine the fluctuation of the mean residual value corresponding to the observation constraint set relative to the mean residual value. If the increase in the mean residual value relative to the mean baseline residual value exceeds the magnitude threshold, then an external parameter drift is determined to have occurred, and iterative optimization is performed again.
7. The method according to claim 1, characterized in that, The conditions under which the calibration state of the extrinsic parameters of the vision sensor is determined to be unstable include at least one of the following: The number of geometric primitive pairs within the sliding window is less than a preset threshold. The included angle of each plane normal in all geometric primitive pairs is less than a preset angle threshold.
8. The method according to claim 1, characterized in that, The sensing data for the two modes include camera images and radar point clouds; The step of extracting geometric primitives from each sensing data in the data frame includes: Point clouds in the camera coordinate system are extracted from the camera images. Candidate planes are detected by a plane detection model. Least square fitting is performed on each candidate plane to obtain the plane normal vector and plane intercept in the camera coordinate system. The radar point cloud is subjected to a random sampling consensus algorithm to detect the plane, and the plane normal vector and plane intercept in the radar coordinate system are obtained.
9. The method according to claim 1, characterized in that, Associating the geometric primitives of two modalities in the data frame based on the environment map includes: Each geometric primitive is back-projected onto the base coordinate system and matched with the semantic geometric primitives in the environment map. The matching condition is that the deviation of the geometric parameters is less than a preset association threshold. Geometric primitives of two modalities that are matched by the same semantic identifier are combined into a geometric primitive pair that shares the same semantic geometric primitive.
10. The method according to claim 1, characterized in that, The sensing data for the two modes include camera images and radar point clouds; The observation constraint set also includes: Point cloud planar distance constraints are used to constrain the distance from points belonging to the same plane to the fitted plane of that plane in the radar coordinate system. Camera plane reprojection constraint is used to constrain the distance from image plane pixels to the contour line obtained by backprojecting the plane in the camera coordinate system; The prior constraints on plane parameters are used to constrain the prior estimates and covariance of the global plane parameters of the semantic geometric primitives in the base coordinate system.
11. The method according to claim 1, characterized in that, The variables to be optimized also include the global geometric parameters of each semantic geometric primitive in the base coordinate system; The iterative optimization includes: If it is determined that the same semantic geometric primitive exists in multiple consecutive data frames, the global geometric parameters of the semantic geometric primitive are used as shared variables to determine the residuals of the multiple consecutive data frames; The global geometric parameters of the semantic geometric primitives are updated through iterative optimization of multiple data frames.
12. A extrinsic parameter calibration device for a vision sensor, characterized in that, include: The data acquisition module is used to acquire sensor data from two modalities synchronously collected by the robot's vision sensor at a preset frame rate, combine the sensor data from the two modalities collected in the same frame into a data frame, and construct a sliding window, wherein the sliding window includes the latest N data frames. The parameter extraction module is used to extract the geometric primitives in each sensing data in each data frame within the sliding window, as well as the single-frame geometric parameters of the geometric primitives in the coordinate system of the corresponding mode. The map building module is used to build an environment map based on the data frames in the sliding window. The environment map includes each semantic geometric primitive and the global geometric parameters of the semantic geometric primitive in the robot's base coordinate system. The semantic geometric primitive is a geometric primitive with a semantic identifier. The semantic identifier is used to uniquely indicate the physical structure corresponding to the corresponding geometric primitive. The association module is used to associate the geometric primitives of two modalities in each data frame within the sliding window according to the environment map, so as to obtain the geometric primitive pairs of the same semantic geometric primitive in the two modalities. The iterative optimization module is used to determine the calibration result of the extrinsic parameters in the current frame if the calibration state of the extrinsic parameters of the visual sensor is determined to be unstable. The module uses the base coordinate system as the world coordinate system, the most recently updated extrinsic parameters as the initial values for iterative optimization in the current frame, and the extrinsic parameters and the single-frame geometric parameters of each geometric primitive pair within the sliding window as variables to be optimized. Iterative optimization is performed by minimizing the observation constraint set constructed based on geometric consistency to determine the calibration result of the extrinsic parameters in the current frame.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the extrinsic parameter calibration method for the visual sensor according to any one of claims 1-11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the extrinsic parameter calibration method for the visual sensor according to any one of claims 1-11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the extrinsic parameter calibration method for the visual sensor according to any one of claims 1-11.