Robot calibration method and apparatus

By constructing a global factor graph to integrate visual, inertial, and kinematic information for online autonomous calibration, the problems of robot sensor extrinsic parameter drift and high existing calibration costs are solved, achieving efficient online calibration and accurate sensor parameter updates.

CN122408830APending Publication Date: 2026-07-17BEIJING ACCELERATED EVOLUTION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ACCELERATED EVOLUTION TECH CO LTD
Filing Date
2026-05-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing robot calibration technologies are prone to sensor extrinsic parameters drift in highly dynamic interactive scenarios, leading to accumulated errors in visual inertial odometry, which affects spatial perception accuracy and autonomous control safety. Furthermore, existing online calibration methods are prone to parameter drift or failure to converge when the robot's motion mode is singular. Traditional offline calibration is costly and disrupts continuous operation capabilities.

Method used

By collecting multi-dimensional data to construct a global factor map, integrating visual, inertial, and kinematic information, online autonomous calibration is achieved. This includes responding to calibration commands to collect environmental feature points and inertial data, controlling the robot to perform translational and rotational movements, determining acceleration and extrinsic parameter data, and constructing a global factor map for calibration.

Benefits of technology

It enables online calibration without downtime at the deployment site, reducing maintenance costs, ensuring continuous robot operation, improving calibration robustness and accuracy, and ensuring parameter convergence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122408830A_ABST
    Figure CN122408830A_ABST
Patent Text Reader

Abstract

This specification provides a robot calibration method and apparatus. The robot calibration method is applied to a target robot and includes: in response to a calibration command triggered for the target robot, acquiring environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot; controlling the target robot to perform translational and rotational movements, and determining acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the movement execution results; acquiring kinematic data corresponding to the kinematic dimension of the target robot during the translational and rotational movements; constructing a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and calibrating the target robot based on the global factor map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of robot control technology, and in particular to robot calibration methods and apparatus. Background Technology

[0002] In the field of robotics, visual inertial odometry (VIO), comprised of binocular vision and an inertial measurement unit (IMU), is a core module for achieving three-dimensional spatial perception and autonomous navigation. The calibration accuracy of the spatial geometric relationships between sensors, including rotation matrices, translation vectors, and time synchronization errors, directly determines the accuracy of state estimation and mapping. However, current calibration techniques face multiple severe challenges in practical applications. In highly dynamic interactive scenarios, such as running, jumping, or navigating complex terrain, and when encountering unforeseen physical collisions, the mechanical mounting structures of sensors like the binocular camera and IMU are highly susceptible to slight deformation due to high-frequency vibrations or hardware stress release, causing deviations in the factory-preset extrinsic parameters. This extrinsic parameter drift directly introduces significant cumulative errors in the VIO, severely jeopardizing the robot's spatial perception accuracy and the safety of its autonomous control decisions. To address the issue of extrinsic parameter drift, traditional methods rely on downtime and return to the factory for rigorous offline calibration using dedicated targets. This process is not only time-consuming and labor-intensive, with extremely high logistics and maintenance costs, but also severely undermines the robot's continuous operation capability and business availability. Therefore, the industry urgently needs a solution that can achieve online calibration on-site without manual intervention. Furthermore, existing targetless online self-calibration algorithms heavily depend on the random motion of the device in three-dimensional space. When the robot is in a stable walking state or with a single motion pattern, the system's translational extrinsic parameters and global scale are easily rendered unobservable, leading to calibration parameter drift or failure to converge. Simultaneously, existing self-calibration frameworks often treat the carrier as a rigid black box, failing to effectively integrate the rich multi-degree-of-freedom joints of the robot body and the kinematic prior information provided by its built-in high-precision encoders, thus limiting the robustness and adaptability of the calibration process. Therefore, an effective solution is urgently needed to address these problems. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a robot calibration method. One or more embodiments of this specification also relate to a robot calibration apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a robot calibration method is provided, applied to a target robot, comprising: In response to a calibration command triggered for the target robot, environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot are collected. The target robot is controlled to perform translational and rotational movements, and acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension are determined based on the movement execution results. During the translational and rotational movements, kinematic data of the target robot in the kinematic dimension are collected. A global factor map is constructed based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and the target robot is calibrated according to the global factor map.

[0005] According to a second aspect of the embodiments of this specification, another robot calibration method is provided, applied to a target robot, comprising: In response to a calibration command submitted by a user through a robot controller for the target robot in the working environment, the system collects environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. The target robot is controlled to perform translational and rotational movements, and acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension are determined based on the movement execution results. During the translational and rotational movements, kinematic data of the target robot in the kinematic dimension are collected. A global factor map is constructed based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and the target robot is calibrated according to the global factor map. The calibrated target robot continues to perform its work tasks in the work environment.

[0006] According to a third aspect of the embodiments of this specification, a robot calibration device is provided, applied to a target robot, comprising: The triggering module is configured to, in response to a calibration command triggered for the target robot, collect environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. The control module is configured to control the target robot to perform translational and rotational movements, and to determine acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the movement execution results. The acquisition module is configured to acquire kinematic data of the target robot in the kinematic dimension during the execution of the translational and rotational movements. The construction module is configured to construct a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and to calibrate the target robot according to the global factor map.

[0007] According to a fourth aspect of the embodiments of this specification, another robot calibration apparatus is provided, applied to a target robot, comprising: The trigger command module is configured to, in response to a calibration command submitted by a user through a robot controller for the target robot in the working environment, collect environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. The control motion module is configured to control the target robot to perform translational and rotational movements, and to determine acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the movement execution results. The data acquisition module is configured to acquire kinematic data of the target robot in the kinematic dimension during the execution of the translational and rotational movements. The subgraph construction module is configured to construct a global factor graph based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and to calibrate the target robot according to the global factor graph; The task execution module is configured to continue performing work tasks in the work environment via the calibrated target robot.

[0008] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-described robot calibration method.

[0009] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the above-described robot calibration method.

[0010] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described robot calibration method.

[0011] The robot calibration method provided in this embodiment collects multi-dimensional initial data in response to calibration commands, controls the robot's movement to obtain corresponding parameters, and then constructs a global factor graph to complete the calibration. By integrating visual, inertial, and kinematic multi-dimensional information, it completes online autonomous calibration, solving the problems of high cost of existing offline calibration affecting continuous robot operation and the inaccuracy of existing online calibration when the robot's motion mode is stable. It can eliminate manual intervention and complete the online calibration of robot sensor extrinsic parameters on the deployment site without stopping the robot to return to the factory for offline calibration, reducing maintenance costs and ensuring the robot's continuous operation capability. At the same time, it effectively integrates the robot's own kinematic prior information, improving the robustness and accuracy of calibration, and ensuring the convergence of calibration parameters even when the robot's motion mode is single and stable. Attached Figure Description

[0012] Figure 1 This is a flowchart illustrating a robot calibration method provided in one embodiment of this specification; Figure 2 This is a flowchart of another robot calibration method provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the structure of a robot calibration device provided in one embodiment of this specification; Figure 4 This is a schematic diagram of the structure of another robot calibration device provided in one embodiment of this specification; Figure 5 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0016] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0018] An IMU (Inertial Measurement Unit) is a sensor combination device that uses an accelerometer, a gyroscope (and a magnetometer) to measure the three-axis acceleration and three-axis angular velocity of a carrier in real time, and calculates its attitude, velocity, and displacement accordingly.

[0019] VIO (Visual-Inertial Odometry) is an algorithm framework that tightly / loosely couples and fuses camera (visual) and IMU (inertial) data to achieve high-precision, high-robust real-time pose estimation (6-DOF), and is also often referred to as VINS (Visual-Inertial System).

[0020] This specification provides a method for calibrating a robot, and also relates to a robot calibration device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.

[0021] In the fields of embodied intelligence and legged robots, the visual inertial odometry (VIO), composed of binocular vision and an inertial measurement unit (IMU), serves as a core module for achieving three-dimensional spatial perception and autonomous navigation. Its calibration accuracy is determined by the accurate calibration of the spatial geometric relationships and time synchronization errors between sensors. The problem of hardware extrinsic parameters drifting due to physical vibrations is particularly prominent. When the robot runs, jumps, or walks on complex terrain during highly dynamic interactions, minute deformations of the sensor mounting structure occur, causing extrinsic parameters to shift, introducing cumulative errors into the VIO, and jeopardizing the accuracy of spatial perception and the safety of autonomous control decisions. Furthermore, existing calibration techniques have been found to have limitations. Offline calibration processes interrupt continuous operation and are unsuitable for adapting to dynamic environments. Simultaneously, online targetless calibration methods are affected by state degradation under conditions of single robot motion modes, leading to translational extrinsic parameters and global scale parameters being affected by unobservable states, resulting in drifting or non-convergence of calibration results. In addition, the sensor fusion dimension is limited, and the kinematic prior information provided by the robot's multi-degree-of-freedom joints and their encoders is not effectively utilized, hindering improvements in calibration accuracy. For example, when a legged robot is performing a field exploration task traversing a gravel slope, it is required to jump to cross obstacles. This causes vibrations, and the mechanical connection between the binocular camera and the inertial measurement unit is affected by minute displacements, leading to deviations in extrinsic parameters. The position trajectory output by the visual-inertial odometry system drifts, the judgment of terrain height is misjudged, and a fall is possible. In this scenario, because the robot is in a stable walking phase, its motion patterns lack diversity, existing online calibration algorithms cannot effectively enhance system observability, translational extrinsic parameters continuously drift, and navigation performance deteriorates.

[0022] In view of this, see Figure 1 , Figure 1 A flowchart of a robot calibration method according to an embodiment of this specification is shown, applied to a target robot, and specifically includes the following steps.

[0023] Step S102: In response to the calibration command triggered for the target robot, collect the environmental feature point data corresponding to the visual dimension and the inertial data corresponding to the inertial dimension of the target robot.

[0024] Step S104: Control the target robot to perform translational and rotational movements, and determine the acceleration data associated with the inertial dimension and the extrinsic parameter data associated with the visual dimension based on the movement execution results.

[0025] Step S106: During the execution of the translational and rotational movements, kinematic data of the target robot in the kinematic dimension is collected.

[0026] Step S108: Construct a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and calibrate the target robot according to the global factor map.

[0027] The robot calibration method provided in this embodiment can be applied to any robot, such as a unibody robot, to enable the robot to perform calibration operations on its own without external intervention in any environment, thereby effectively reducing costs and improving the robot's operational stability.

[0028] Specifically, the target robot refers to the robotic entity that requires calibration, typically equipped with various sensors such as vision sensors, inertial sensors, and kinematic sensors. The calibration command refers to the command or signal that triggers the robot to begin the calibration process. This command can be manually input by the user or automatically triggered by the system based on preset conditions. The visual dimension refers to the way environmental information is acquired through optical sensors (such as cameras). In this dimension, the robot can perceive visual features in the environment, such as textures, edges, and corners. Environmental feature point data refers to visual information acquired from vision sensors that describes specific points in the environment. This data typically includes the pixel coordinates, three-dimensional coordinates, and their correspondence in different image frames of the feature points. The inertial dimension refers to the way the robot's own motion state information is acquired through an inertial measurement unit (IMU). An IMU typically includes accelerometers and gyroscopes, capable of measuring the robot's linear acceleration and angular velocity. Inertial data refers to the raw data collected by the inertial measurement unit (IMU), including accelerometer data and gyroscope data. This data reflects the robot's motion trends and attitude changes in space.

[0029] Correspondingly, translational motion refers to the displacement of the robot body along a straight line in space, without involving changes in posture. For example, the robot's legs or torso can be controlled to perform smooth reciprocating movements along the Z-axis (direction of gravity) of the absolute coordinate system, such as squatting and standing up. By introducing pure and significant Z-axis acceleration, visual feature loss caused by violent rotation can be avoided, breaking the scale ambiguity of visual-inertial fusion and forcibly activating the observability of the translational component in the extrinsic parameters. Rotational motion refers to the posture change of the robot body around a certain axis in space, without involving changes in position, such as controlling the robot's neck to perform orthogonal axis rotation scanning movements (such as yaw and pitch angle scans). Pure rotational motion provides sufficient angular velocity and angular acceleration excitation, driving visual natural feature points to move over a wide range within the camera's field of view, used for high-precision calculation of the rotational extrinsic parameters of the camera and IMU, as well as the binocular relative extrinsic parameters. Acceleration data refers to data describing the rate of change of the robot's motion velocity. During calibration, this data is used to correlate the inertial dimension to reflect the dynamic characteristics of the robot during translation or rotation. Extrinsic parameter data refers to data describing the spatial geometric relationships between different sensors or between a sensor and the robot body, typically including rotation matrices and translation vectors. This data is crucial for achieving multi-sensor data fusion. The kinematic dimension refers to the way information such as the robot's joint position and velocity is obtained through kinematic sensors such as joint encoders on the robot body. This dimension provides precise measurements of the robot's own motion. Kinematic data refers to data collected by the robot's joint encoders or other kinematic sensors, reflecting the motion state of each joint and the overall pose changes of the robot. A global factor graph is a graphical model used to represent and optimize multi-sensor data fusion problems. In this graph, nodes represent state variables to be estimated (such as extrinsic parameters and time offsets), and edges represent sensor measurements or prior information. Optimal state variables can be solved by optimizing this graph.

[0030] Based on this, in response to a calibration command triggered for the target robot, environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension are collected. Specifically, the triggering method for the calibration command can be varied. For example, the operator can manually input the command through the user interface to start the calibration program. Alternatively, the robot system can automatically trigger it according to preset conditions, such as when the robot starts up, during routine maintenance, or when abnormal sensor data is detected. When collecting environmental feature point data, basic image processing techniques can be used, such as edge detection or corner detection of images captured by the camera to extract salient feature points. These feature points can be used as environmental feature point data for subsequent visual information processing. When collecting inertial data, the raw accelerometer and gyroscope outputs of the inertial measurement unit (IMU) over a period of time can be simply read. This raw data directly reflects the robot's instantaneous motion state.

[0031] Secondly, the target robot is controlled to perform translational and rotational movements, and the acceleration data associated with the inertial dimension and the extrinsic parameter data associated with the visual dimension are determined based on the movement results. To achieve translational movements, a series of straight or curved trajectories can be pre-set, and the robot can be controlled to move along these trajectories. For example, the robot can be controlled to perform a squatting motion along a straight line on the ground. To achieve rotational movements, the robot's head can be controlled to rotate around its own axis or an external point. During the movement execution, the robot's acceleration data can be estimated by performing simple integration or difference processing on the raw inertial data acquired by the IMU. Simultaneously, by analyzing images captured by the visual sensor in different postures and combining simple geometric transformation principles, the extrinsic parameter data between the visual and inertial sensors can be preliminarily estimated.

[0032] Secondly, during the translational and rotational movements, kinematic data of the target robot in the kinematic dimensions is collected. This kinematic data can be acquired by reading encoder data from each joint of the robot. For example, in a multi-joint robot, each joint may be equipped with an encoder to measure the joint's rotation angle or displacement. As the robot performs translational and rotational movements, these encoders continuously output data, recording the real-time motion state of each joint. This data provides precise information about the robot's overall motion.

[0033] Finally, a global factor map is constructed based on the environmental feature point data, extrinsic parameter data, inertial data, acceleration data, and kinematic data, and the target robot is calibrated according to this global factor map. When constructing the global factor map, the various types of data collected above are used as observations in the map, and the extrinsic parameters to be calibrated and time offsets are used as state variables in the map. For example, environmental feature point data and preliminarily estimated extrinsic parameter data are used to construct visual observation factors, inertial data and acceleration data are used to construct inertial measurement factors, and kinematic data are used to construct kinematic prior factors. After construction, an iterative optimization algorithm, such as the nonlinear least squares method, can be used to solve the global factor map. By optimizing the state variables in the map, more accurate sensor extrinsic parameters and time offsets can be obtained, thereby completing the calibration of the target robot.

[0034] For example, suppose user A deploys a legged robot at location A. This robot is equipped with binocular cameras, an inertial measurement unit (IMU), and encoders for each joint. Due to potential vibrations during transport or minor deformations in the hardware structure after long-term operation, the robot's factory-set sensor extrinsic parameters and time synchronization may deviate, affecting its perception accuracy. To address this issue, user A sends a calibration command to the legged robot via a remote control system. Upon receiving the command, the robot first enters the data acquisition phase. In this phase, the robot uses its binocular cameras to continuously acquire image sequences of the surrounding environment and extracts a large amount of environmental feature point data, such as ground texture and wall edges. Simultaneously, the IMU begins to synchronously acquire the robot's raw inertial data, including three-axis acceleration and three-axis angular velocity. Subsequently, the robot is controlled to execute a series of preset translational and rotational movements. For example, the robot first performs a squatting motion and then rotates its head. During these movements, the robot continuously acquires IMU data and calculates the acceleration data associated with the inertial dimension based on this data. Meanwhile, by analyzing images captured by the binocular cameras in different poses and combining them with robot motion information, the extrinsic parameters associated with the visual dimension were preliminarily determined, such as the rotation matrix and translation vector between the camera and the IMU. Notably, during the entire translation and rotation process, the encoders at each joint of the robot also worked synchronously, accurately recording the kinematic data of each joint, such as joint angles and angular velocities.

[0035] After all necessary data acquisition is completed, the robot system begins constructing a global factor graph. In this graph, previously acquired environmental feature point data, preliminarily determined extrinsic parameter data, inertial data, acceleration data, and kinematic data are added as different observation factors or prior factors. For example, visual feature point observations constitute the visual factor, inertial data measurements constitute the inertial factor, and kinematic encoder measurements constitute the kinematic factor. These factors collectively constrain the state variables to be optimized, including the precise extrinsic parameters between the camera and IMU and the time offset between them. The system uses a nonlinear optimization algorithm to solve for this global factor graph, iteratively adjusting the state variables to minimize the residuals between all observations and predictions. Finally, the optimization process converges, yielding a relatively accurate set of camera-IMU extrinsic parameter data and time offset data. This data is used to update the robot's sensor configuration, thus completing the high-precision calibration of the legged robot.

[0036] In summary, this system collects multi-dimensional initial data in response to calibration commands, controls robot movement to obtain corresponding parameters, and constructs a global factor graph to complete calibration. By integrating visual, inertial, and kinematic information, it achieves online autonomous calibration, solving the problems of high cost of existing offline calibration affecting continuous robot operation and the inaccuracy of existing online calibration when the robot's motion mode is stable. It can eliminate manual intervention and complete the online calibration of robot sensor extrinsic parameters on-site, without the need for downtime and return to the factory for offline calibration, reducing maintenance costs and ensuring the robot's continuous operation capability. At the same time, it effectively integrates the robot's own kinematic prior information, improving the robustness and accuracy of calibration, and ensuring the convergence of calibration parameters even when the robot's motion mode is simple and stable.

[0037] In practice, the robot may be in a slightly wobbly or unstable state when receiving commands, which will cause the initial inertial data to contain motion noise, affecting the accuracy of the inertial sensor's zero-bias estimation and initial attitude. Simultaneously, images acquired by the vision system in an unstable state may also lead to errors in feature point extraction and tracking, thus affecting subsequent calibration accuracy. Therefore, in this embodiment, the step of collecting environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot in response to a calibration command triggered for the target robot includes: In response to a calibration command triggered for the target robot, the target robot is controlled to maintain a stationary state for a set time; if the duration of the stationary state is greater than the set time, the target robot's environmental feature point data in the visual dimension is collected, as well as the target robot's initial inertial data, gyroscope data, and gravity data are collected; the initial inertial data, the gyroscope data, and the gravity data are used as the inertial data corresponding to the inertial dimension.

[0038] Specifically, controlling the target robot to remain stationary for a set period of time aims to ensure that the robot is in a stable, motion-undisturbed initial condition before data acquisition begins. This stationary state is crucial for zero-bias estimation of inertial sensors and initial feature point extraction of visual sensors, effectively avoiding noise and errors introduced by motion and laying the foundation for subsequent accurate calibration. This stationary state can be achieved by the robot's motion controller receiving calibration commands, immediately stopping all joint movement, locking joint positions, and simultaneously monitoring the output of the inertial measurement unit (IMU) to ensure that its acceleration and angular velocity readings fluctuate within a preset small threshold range, thus confirming that the robot is indeed stationary. Alternatively, external sensors, such as LiDAR or depth cameras, can be used to monitor the robot's pose changes relative to the environment; when the pose change is less than a certain minimum value within a set time, the robot is considered to be stationary.

[0039] Correspondingly, the set time refers to the minimum duration for which the robot needs to remain stationary. This time length is set to ensure that the inertial sensor has sufficient time for self-calibration or to stabilize its output, and to provide the vision system with stable image frames for initial feature point extraction. The set time can be empirically set based on the characteristics of the inertial sensor used (such as startup time and stabilization time) and the requirements of the vision algorithm for the image sequence, for example, set to 2 seconds, 5 seconds, or 10 seconds. Furthermore, the set time can also be dynamically adjusted. For example, the system can determine when a stable state is reached based on the stability indicators (such as variance) of the inertial sensor output. Once a stable state is reached, even if the preset fixed time has not been reached, the set time requirement can be considered met.

[0040] Accordingly, when the duration of the static state exceeds the set time, environmental feature point data corresponding to the target robot in the visual dimension is collected, as well as the initial inertial data, gyroscope data, and gravity data of the target robot. This technical feature specifies the triggering conditions and specific content of data acquisition. Data acquisition only begins after the robot confirms that it is in a stable static state for a sufficiently long time, ensuring that the acquired data has high initial quality. The acquired data includes visual environmental feature point data (used for visual localization and environmental modeling), initial inertial data (usually referring to the raw acceleration and angular velocity data of the IMU), gyroscope data (angular velocity data), and gravity data (used to determine the initial attitude and gravity direction). When the system detects that the duration of the static state meets the set time, the visual sensor (such as a camera) begins to capture image sequences and extract environmental feature points from them; simultaneously, the inertial measurement unit (IMU) begins to record its raw accelerometer and gyroscope outputs, and can estimate the gravity vector in the static state based on the accelerometer readings. As another implementation, after the static state has lasted for a set period of time, a dedicated data acquisition module can initiate a synchronized data stream between the vision and inertial sensors. The vision data acquisition module is responsible for identifying and storing feature point information from the image, while the inertial data acquisition module is responsible for reading and storing the raw acceleration, angular velocity, and gravity direction information calculated by the accelerometer in the static state from the IMU.

[0041] Accordingly, the initial inertial data, the gyroscope data, and the gravity data are used as the inertial data corresponding to the inertial dimension. This technical feature clarifies that specific inertial-related data collected in a static state will be integrated and used as the data input for the inertial dimension in subsequent calibration processes. This integration ensures that the inertial data not only contains the original motion information but also the key information obtained in a static state for attitude estimation and gravity alignment, thereby improving the integrity and accuracy of the inertial data. The system can package or format the raw acceleration and angular velocity readings directly obtained from the IMU sensor (i.e., initial inertial data and gyroscope data) with the gravity vector estimated from the accelerometer readings in a static state to form a unified inertial data package for subsequent factor map construction. Alternatively, the initial inertial data, gyroscope data, and gravity data can be fused during the data preprocessing stage. For example, gravity data can be used to roughly align the initial attitude, and then these preprocessed and fused data can be input into the calibration algorithm as the inertial data corresponding to the inertial dimension.

[0042] Based on this, in response to a calibration command triggered for the target robot, the system first controls the target robot to remain stationary for a set time. This step ensures that the robot is in a stable, motion-undisturbed initial condition before data acquisition begins. By monitoring the robot's motion state, for example through the output of the inertial measurement unit or an external pose sensor, the system can determine whether the robot has truly reached a stationary state and maintained it for a sufficient duration. Only when the duration of the stationary state exceeds the set time does the system begin acquiring environmental feature point data corresponding to the target robot in the visual dimension, as well as the initial inertial data, gyroscope data, and gravity data of the target robot. This conditional acquisition mechanism effectively avoids the problem of acquiring low-quality data when the robot is unstable. Specifically, in a stationary state, the accelerometer of the inertial measurement unit can accurately measure the direction of gravity, thereby providing reliable gravity data, which is crucial for subsequent attitude initialization and inertial data pre-integration. At the same time, gyroscope data in a stationary state can be used to more accurately estimate its zero bias. Images acquired by the visual sensor in a stationary state are also more stable, which is beneficial for extracting high-quality environmental feature points. Finally, the initial inertial data, the gyroscope data, and the gravity data are integrated as the inertial data corresponding to the inertial dimension, providing high-quality initial inertial information for the subsequent construction of the global factor map. In this way, the solution of this application ensures the quality of the initial data in the calibration process, providing accurate and reliable input for subsequent translational movements, rotational movements, and the construction of the global factor map, thereby significantly improving the overall accuracy and robustness of robot calibration.

[0043] For example, when a user sends calibration commands to the target robot through the user interface, the robot first stops all current movement and enters a static waiting mode. In this mode, the robot's motion controller locks all joints and continuously monitors the output of the robot's internal inertial measurement unit (IMU). For instance, the system continuously checks whether the IMU's accelerometer readings fluctuate within ±0.05g and whether the gyroscope readings fluctuate within ±0.1 degrees / second. If these readings remain continuously within a specified threshold for more than a preset 5 seconds, the system considers the robot to have reached a stable static state. At this point, the robot's stereo camera system begins to capture a series of image frames and extracts natural texture features such as corners and edges from the environment, forming environmental feature point data corresponding to the visual dimension. Simultaneously, the IMU begins to record its raw accelerometer and gyroscope data, which is labeled as initial inertial data and gyroscope data. Furthermore, when the robot is stationary, the accelerometer readings primarily reflect gravitational acceleration, and the system uses these readings to calculate the gravity vector as gravity data. Ultimately, the initial inertial data, gyroscope data, and gravity data collected in a stable, stationary state will be packaged and integrated as the inertial data input corresponding to the inertial dimension in subsequent calibration algorithms.

[0044] In summary, by forcing the robot to remain stationary for a set period before data acquisition and then collecting data in this stable state, data noise and errors caused by the robot's initial movement or unstable state can be effectively avoided. This results in cleaner and more accurate initial inertial, gyroscope, and gravity data, providing a reliable foundation for subsequent inertial data pre-integration and attitude initialization. Simultaneously, visual environment feature point data acquired in a stationary state is also more stable and accurate, which is beneficial for improving the quality of feature point extraction and matching. This high-quality initial data input significantly improves the accuracy of subsequent global factor map construction, thus making the calibration results for the target robot more accurate and robust, effectively solving the problem of decreased calibration accuracy caused by data acquisition under unstable initial conditions. Furthermore, accurately and effectively determining the acceleration data associated with the inertial dimension and the extrinsic parameter data associated with the visual dimension from the robot's motion execution results is a key issue affecting the final calibration accuracy and efficiency. If this crucial data is not accurately obtained, it will directly affect the subsequent construction of the factor map and the reliability of the calibration results. Therefore, in this embodiment, controlling the target robot to perform translational and rotational movements, and determining the acceleration data associated with the inertial dimension and the extrinsic parameter data associated with the visual dimension based on the motion execution results, includes: The target robot is controlled to perform a translational movement, and acceleration data associated with the inertial dimension is determined based on the translational movement execution result; the target robot is controlled to perform a rotational movement, and rotational extrinsic data and relative extrinsic data associated with the visual dimension are determined based on the rotational movement execution result, and the rotational extrinsic data and the relative extrinsic data are used as extrinsic data.

[0045] Specifically, controlling a target robot to perform translational movements aims to simplify the motion model by guiding the robot through pure translational motion, thereby more accurately separating and extracting acceleration information related to the inertial dimension. Translational movements can be achieved through pre-planned path planning, such as guiding the robot to move at a constant or variable speed on a straight track, or performing non-rotational linear reciprocating motion in a plane. Another approach is to use an external positioning system (such as a laser tracker or optical motion capture system) to monitor the robot's position in real time and adjust the robot's motion based on feedback to ensure that it only performs translational movements. Determining the acceleration data associated with the inertial dimension based on the translational movement results means obtaining the acceleration information in the inertial coordinate system by analyzing the target robot's motion state during the translational movement. Methods for determining acceleration data can include: directly measuring acceleration using the robot's internal inertial measurement unit (IMU) and calibrating and denoising based on the characteristics of translational motion; or calculating acceleration by performing a second derivative on the position data collected during the robot's translation, where the position data can come from a visual odometry, lidar odometry, or an external high-precision positioning system. Controlling a target robot to perform rotational movements aims to simplify the motion model by guiding the robot through pure rotational motion, thereby more accurately separating and extracting extrinsic information related to the visual dimension. Rotational movements can be achieved through preset rotational trajectories, such as allowing the robot to rotate at multiple angles from a fixed point or rotate continuously around an axis. Another approach is to utilize an external vision-assisted system to precisely control its rotational posture by identifying specific marker points on the robot, ensuring that it only rotates. Determining the rotational extrinsic data and relative extrinsic data associated with the visual dimension based on the rotational movement execution results refers to determining the rotational relationship (rotational extrinsic data) between the vision sensors and the robot body, and the relative positional relationship (relative extrinsic data) between the vision sensors (such as cameras) during the target robot's rotational movement by analyzing data captured by its vision sensors (such as cameras). Determining these extrinsic parameters can be achieved through methods such as: using visual SLAM (Simultaneous Localization and Mapping) techniques to estimate the camera pose and 3D positions of feature points during rotation via feature point matching and bundle adjustment, thereby deriving the extrinsic parameters; or, by observing images of a specific calibration plate at different rotation angles, using the Perspective-n-Point (PnP) algorithm or multi-view geometry methods to calculate the relative pose between the camera and the robot body. Using rotational and relative extrinsic parameter data as extrinsic parameters clarifies that during calibration, the rotational and relative extrinsic parameter data determined through rotational actions are combined or uniformly represented as the final extrinsic parameter data.This means that in the subsequent construction of the global factor graph, all external parameters related to the visual dimension will be based on this data, which is precisely acquired through rotational movements. This integration ensures the integrity and consistency of the geometric relationship between the vision system and the robot body.

[0046] Based on this, by decomposing complex robot motion into independent translational and rotational movements, precise determination of inertial acceleration data and visual extrinsic parameters can be achieved. Specifically, during translational movements, the target robot is controlled to only change its position without rotating its posture. In this process, the acceleration sensed by the inertial measurement unit (IMU) primarily originates from the robot's translational motion. By acquiring and processing this acceleration data, acceleration data associated with the inertial dimension can be effectively separated, avoiding interference from rotational motion on acceleration measurements. Subsequently, during rotational movements, the target robot is controlled to only change its posture without translating its position. During this process, the image sequences captured by visual sensors (such as cameras) primarily reflect its own rotational changes relative to the environment. By analyzing this visual data, the rotational extrinsic parameters between the visual sensors and the robot, as well as the relative extrinsic parameters between different visual sensors, can be accurately calculated. This step-by-step, decoupled motion control strategy allows each motion mode to focus on extracting specific types of calibration parameters, thereby improving the purity of data acquisition and the accuracy of parameter determination. In this way, high-quality acceleration and extrinsic parameter data are provided for the subsequent construction of a global factor map based on environmental feature point data, inertial data, and kinematic data, which significantly improves the robustness and accuracy of the overall calibration method.

[0047] For example, when controlling a target robot to perform a translational movement, a corresponding squat command can be pre-planned. Based on this squat command, the robot is driven to perform a squatting motion, resulting in a translational effect. During this translation, the robot's internal IMU collects three-axis acceleration data at a frequency of 200Hz. To determine the acceleration data associated with the inertial dimension, the raw acceleration data collected by the IMU can be low-pass filtered to eliminate high-frequency noise. Combined with the robot's kinematic model, the acceleration can be estimated and optimized using Kalman filtering or extended Kalman filtering. When controlling the target robot to perform a rotational movement, the robot's head can be set to rotate at a set angle. Simultaneously, the robot's first and second image acquisition devices synchronously acquire image sequences at a frame rate of 30 frames per second. To determine the rotational and relative extrinsic parameters associated with the visual dimension, visual odometry technology can be used to match and track feature points in consecutive frames during rotation. The relative attitude changes between cameras can be calculated using multi-view geometric methods (e.g., essential matrix or fundamental matrix estimation). Furthermore, a calibration board detection algorithm can be combined to include a pre-placed checkerboard calibration board within the robot's field of view during rotation. By identifying images of the calibration board from different perspectives, the rotational and relative extrinsic parameters between the camera and the robot body can be accurately calculated using a calibration method or its variants. Finally, the rotational and relative extrinsic parameter data obtained through these methods are integrated as complete extrinsic parameter data for subsequent global factor graph construction.

[0048] In summary, by decomposing the robot's complex motion into independent translational and rotational movements, the system can focus on accurately acquiring inertial acceleration data during translational movements and accurately acquiring visual extrinsic parameter data during rotational movements. This decoupled motion control and data acquisition strategy effectively avoids data interference between different motion modes, significantly improving the purity and accuracy of acceleration and extrinsic parameter data. Therefore, it provides high-quality input data for the subsequent construction of a global factor graph, enabling more accurate estimation of the robot's state variables and ultimately achieving more reliable and precise calibration of the target robot.

[0049] In practice, to ensure that no degradation occurs during the calibration of external parameters, squatting and head-shaking movements can be used to activate the observability of parameters from a physical dynamics perspective. Taking the translational external parameter between the camera and the IMU as an example, the acceleration data can be determined by the following formula (1): (1) in, This represents the acceleration data (linear acceleration vector) determined in the inertial local coordinate system. This represents the rotation matrix from the camera coordinate system to the inertial coordinate system. This represents the linear acceleration data in the camera coordinate system. This represents the angular acceleration data in the camera coordinate system. This represents the angular velocity data in the camera coordinate system. This represents the translation data from the origin of the inertial coordinate system to the origin of the camera coordinate system. This represents the cross product operation between vectors.

[0050] It should be noted that the translation extrinsic parameters Only with angular acceleration and angular velocity Coupling (i.e., cross product) occurs if the robot is in a smooth motion ( Then it contains the quantity to be determined. The terms will approach the zero vector, and the corresponding dimension of the system's Jacobian matrix will generate a null space, leading to... Unable to solve. This embodiment actively injects high-intensity linear acceleration through a preset action (breaking visual scale blur to extract precise...). ) and angular acceleration with high signal-to-noise ratio and angular velocity This fundamentally guarantees the transfer of external parameters. The full rank is considerable in global optimization.

[0051] Right now: This refers to the linear acceleration felt by the robot in its own inertial coordinate system. This data is one of the core parameters for describing the robot's motion state and is crucial for subsequent kinematic and dynamic analysis. It can be measured directly by an inertial measurement unit (IMU) or estimated using multi-sensor fusion algorithms. This matrix represents the rotational relationship between the camera coordinate system and the inertial coordinate system. It describes the camera's attitude relative to the inertial measurement unit (IMU) and is crucial for coordinate system transformation. It can be pre-acquired using external calibration tools or estimated in real-time through optimization algorithms during system operation. This refers to linear acceleration data observed in the camera coordinate system. This refers to the angular acceleration data observed in the camera coordinate system. This is the translation data from the origin of the inertial coordinate system to the origin of the camera coordinate system. This vector describes the camera's mounting position relative to the inertial measurement unit (IMU). It can be obtained through precise mechanical measurement or estimated through methods such as hand-eye calibration. The above formula is used to transform the kinematic parameters (angular acceleration, angular velocity, translation) in the camera coordinate system to the inertial coordinate system, thereby accurately calculating the acceleration data in the inertial coordinate system. This formula considers the relative pose relationship between the camera and the inertial measurement unit, and through rotation matrix and cross product operations, it eliminates measurement errors caused by differences in sensor mounting position and attitude, ensuring the accuracy of the acceleration data.

[0052] In summary, the above processing method accurately transforms the kinematic parameters in the camera coordinate system to the inertial coordinate system, thereby obtaining high-precision acceleration data in the inertial local coordinate system. This effectively solves the problem of inaccurate acceleration measurement caused by the installation offset between the camera and the inertial measurement unit. High-precision acceleration data, as a key input for constructing the global factor graph, can significantly improve the estimation accuracy of robot state variables (such as extrinsic parameter data and time offset data) during factor graph optimization, thereby improving the overall calibration accuracy and robustness of the target robot. Furthermore, when collecting environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot in response to a calibration command triggered for the target robot, simply collecting environmental feature point data may encounter problems such as uneven distribution of feature points, a large amount of noise or mismatched points, and a lack of depth information. These problems will directly affect the accuracy and robustness of the subsequent global factor map construction, thereby reducing the accuracy of robot calibration. Therefore, in this embodiment, the collection of environmental feature point data corresponding to the visual dimension of the target robot includes: A first image sequence associated with a first image acquisition device is acquired, wherein the target robot includes a first image acquisition device and a second image acquisition device. The first image sequence is divided into grids, and natural texture feature points are extracted from the divided image sequence according to a feature point extraction algorithm to obtain an initial feature point set. Each feature point in the initial feature point set is continuously tracked in the temporal dimension, and the first feature point determined after tracking is matched to the second image acquisition device to obtain initial depth data. Based on the initial depth data, a feature point matching set associated with the spatiotemporal dimension is determined. Anomaly detection is performed on the feature point matching set, and abnormal feature points are removed according to the anomaly detection results to obtain a target feature point set. Key feature points are determined in the target feature point set, and the feature point data corresponding to the key feature points are used as the environmental feature point data corresponding to the visual dimension.

[0053] Specifically, the first and second image acquisition devices are sensors, typically cameras, used to capture visual information about the environment. They can be configured as a stereo camera pair to provide depth information; for example, they can be any two from a monocular, binocular, or multi-camera system. The first image sequence refers to a series of image frames continuously captured by the first image acquisition device over a period of time, providing visual information about the environment during the target robot's movement and serving as the basis for subsequent feature point extraction and tracking. Mesh partitioning refers to spatially dividing each image frame into several small, regular regions, such as rectangular grids. This helps to extract feature points evenly across different regions of the image, avoiding the concentration of feature points in certain areas while other areas lack feature points, thereby improving the uniformity and coverage of feature point distribution. Feature point extraction algorithms are algorithms used to identify points with salientity, repeatability, and discriminability from an image, such as SIFT, SURF, ORB, or FAST. These algorithms can effectively identify stable feature points from images. Natural texture feature points refer to feature points extracted from naturally existing textures in the environment, rather than artificially set markers, and have environmental universality. The initial feature point set is the collection of all initially identified feature points, which may include some unstable, repetitive, or low-quality feature points. Continuous inter-frame tracking in the temporal dimension refers to tracking the same feature point across consecutive image frames in the first image sequence to determine its position at different time points. This is typically achieved through optical flow or descriptor matching. Matching the tracked first feature point to the second image acquisition device means finding the corresponding matching point in the image captured by the second image acquisition device at the same time as the feature point tracked in the first image acquisition device. This is typically achieved through stereo matching algorithms. Initial depth data refers to the distance information from the feature point to the camera plane calculated through stereo matching. The feature point matching set associated with the spatiotemporal dimension refers to the set that integrates the feature points obtained after temporal tracking and stereo matching, along with their corresponding depth information, timestamps, and image coordinates. Anomaly detection is a technique for identifying points in a dataset that do not conform to expected patterns or behaviors. In feature point matching, anomalies may include mismatched points, motion-blurred points, occluded points, or points with inaccurate depth calculations. Anomaly detection algorithms can be based on statistical methods, geometric constraints, or machine learning methods. Anomaly removal refers to removing inaccurate or unreliable feature points identified by anomaly detection algorithms from the feature point matching set. The target feature point set is the set of feature points remaining after anomaly detection and removal, which are considered reliable and of high quality. Identifying key feature points involves further filtering from the target feature point set to select feature points with the highest value for robot calibration and localization. These key feature points typically have good spatial distribution, stable tracking performance, high disparity variation, or are located in important regions within the robot's field of view.The feature point data corresponding to the key feature point refers to all relevant information of the selected key feature point, including but not limited to its pixel coordinates in the image, its corresponding three-dimensional spatial coordinates, timestamp, descriptor information, and its tracking ID between different frames.

[0054] Based on this, when collecting environmental feature point data corresponding to the visual dimension of the target robot, the first and second image acquisition devices equipped on the target robot are used to continuously capture a first image sequence through the first image acquisition device. To ensure the uniform distribution of feature points in the images and improve extraction efficiency, the first image sequence is divided into grids, and a feature point extraction algorithm is used within each grid to identify natural texture feature points, thereby obtaining an initial feature point set. Subsequently, to establish the temporal and spatial correlation of feature points, each feature point in the initial feature point set is continuously tracked between frames in the temporal domain, and the tracking results are matched with the images captured by the second image acquisition device. Initial depth data is calculated using the principle of stereo vision, thereby constructing a feature point matching set associated with the spatiotemporal dimensions. To ensure the quality and reliability of the feature point data, this application performs anomaly detection on the obtained feature point matching set, identifying and eliminating abnormal feature points that are mismatched, unstable, or have inaccurate depth calculations, thereby obtaining a more reliable target feature point set. Finally, key feature points crucial for robot calibration are further selected from the target feature point set. These key feature points typically exhibit good spatial distribution and stable tracking performance, and their corresponding feature point data are used as the final visual dimension environmental feature point data. Through the above refined feature point acquisition and selection process, this application effectively solves the problems of high noise, uneven distribution, and numerous mismatches that may exist in environmental feature point data in traditional methods. This method not only ensures the quality of visual data, providing high-confidence visual observations for subsequent robot calibration based on global factor maps, but also significantly improves the accuracy and robustness of environmental feature point data through multi-dimensional optimization (temporal tracking, stereo matching, anomaly detection, and key point selection), thus laying a solid foundation for the accuracy and stability of the entire robot calibration process.

[0055] In practice, to ensure the visual reprojection residuals in the factor plot The high confidence and high accuracy can be achieved through the following processing at the visual front end: 1. Feature Extraction and Homogenization: During the execution of the excitation action, the left eye image sequence is divided into grids, and the Shi-Tomasi algorithm is used to extract natural texture feature points to ensure that the feature points are uniformly distributed within the camera's field of view (FOV).

[0056] 2. Spatiotemporal dual-dimensional tracking and matching: (1) Temporal domain: The multi-layer pyramid LK optical flow method (Lucas-Kanade Optical Flow) is used to track feature points in continuous video frames. (2) Spatial domain: The left eye feature points are accurately matched to the right eye using the binocular epipolar geometric constraint, and the initial depth is calculated.

[0057] 3. Dynamic and Abnormal Feature Removal: During rapid motion excitation, the RANSAC algorithm combined with the Fundamental Matrix is ​​used to identify and remove outliers caused by motion blur or dynamic obstacles, retaining only static inliers.

[0058] 4. Keyframe Filtering Strategy: A threshold is set based on the average disparity change of feature points. When an action causes a change in the field of view that exceeds the threshold, keyframe extraction is triggered, and the 3D coordinates and pixel coordinates of the visual features contained therein are used as observation values ​​and input into the global factor graph optimizer for processing.

[0059] For example, the following steps can be used to collect environmental feature point data corresponding to the visual dimension of a target robot. First, the target robot can be equipped with a pair of synchronously triggered global shutter cameras as the first and second image acquisition devices. The first image acquisition device continuously captures images at a frequency of 30 frames per second, forming a first image sequence. For each frame, it can be divided into a 16x16 grid, and feature points are extracted within each grid using the ORB feature point extraction algorithm to obtain an initial feature point set. Next, for each feature point in the initial feature point set, the Lucas-Kanade optical flow method can be used to track it in 5 consecutive frames to establish its temporal association. Simultaneously, the tracked feature points are matched with the images captured by the second image acquisition device at the same time using BRIEF descriptors, and filtered using epipolar geometric constraints to calculate the initial depth data. Based on these tracking and matching results, a feature point matching set with associated spatiotemporal dimensions, including feature point pixel coordinates, 3D coordinates, timestamps, and tracking IDs, can be constructed. Subsequently, to improve data quality, the RANSAC algorithm can be applied to the feature point matching set for anomaly detection, removing mismatched points that do not conform to epipolar geometry constraints and points with large depth calculation errors, thereby obtaining the target feature point set. Finally, in the target feature point set, key feature points can be determined based on the tracking time of the feature points, their uniformity of distribution in the image, and the magnitude of their corresponding disparity changes. For example, feature points with a tracking time of more than 10 frames, distributed in the central region of the image, and with large disparity changes are preferentially selected as key feature points, and their three-dimensional coordinates and pixel coordinates are used as the environmental feature point data corresponding to the visual dimension.

[0060] In summary, the above-described processing effectively addresses the problem of decreased calibration accuracy and robustness caused by low-quality environmental feature point data (such as high noise, uneven distribution, and numerous mismatches) during robot calibration. Through a refined process of feature point acquisition, tracking, matching, anomaly detection, and key point screening, this application can obtain high-quality, highly reliable visual environmental feature point data with good spatial distribution. This high-quality visual data serves as crucial input for constructing the global factor graph, significantly improving the convergence speed and accuracy of the factor graph optimization process. This results in more accurate target extrinsic parameter data and time offset data when calibrating the target robot, thereby enhancing the overall accuracy and stability of the robot's localization, navigation, and task execution. Furthermore, the target feature point set may contain a large number of redundant or insufficiently informative feature points. Directly using all feature points for subsequent calibration calculations would increase the computational burden and may lead to a decrease in calibration accuracy due to uneven distribution or insufficient information of feature points. Therefore, in this embodiment, the step of determining key feature points in the target feature point set and using the feature point data corresponding to the key feature points as the environmental feature point data corresponding to the visual dimension includes: A field-of-view threshold is constructed based on the average disparity change information corresponding to each target feature point in the target feature point set; when the field-of-view change of the target robot is greater than the field-of-view threshold, a key image frame is determined; key feature points corresponding to the key image frame are extracted from the target feature point set, and the three-dimensional coordinates and pixel coordinates corresponding to the key feature points are used as environmental feature point data corresponding to the visual dimension.

[0061] Specifically, a field-of-view threshold is constructed based on the average disparity change information corresponding to each target feature point in the target feature point set, aiming to provide a quantitative standard for subsequent selection of key image frames. The average disparity change information reflects the degree of positional change of the feature point under different viewpoints and is closely related to the depth of the feature point and camera motion. By analyzing these changes, the amount of geometric information contained in the feature point can be evaluated. The field-of-view threshold can be constructed in various ways. For example, the pixel displacement of the target feature point between consecutive frames can be calculated, and its disparity change in three-dimensional space can be estimated by combining camera intrinsic and extrinsic parameters. Then, statistical analysis (such as calculating the average, median, or setting percentiles) can be performed on the disparity changes of all target feature points to determine the threshold. Alternatively, the field-of-view threshold can be dynamically adjusted by analyzing the motion vector length of the feature point on the image plane and combining it with a preset motion model or empirical value.

[0062] Accordingly, when the target robot's field of view changes more than a field of view threshold, key image frames are identified. This step aims to filter out image frames containing rich motion information or significant geometric changes. When the robot's field of view change (i.e., the change in camera viewpoint) is sufficiently large, it usually means that the frame has captured new, valuable environmental information, or provides a sufficiently large baseline to improve the accuracy of depth estimation and feature point localization. Key image frames can be identified, for example, by calculating the norm of the camera pose change between the current and previous frames (such as the length of the translation vector or the rotation angle) and comparing it to the field of view threshold; alternatively, they can be determined by evaluating the number of newly detected feature points or the success rate of feature point tracking in the current frame, combined with the field of view threshold.

[0063] Accordingly, key feature points corresponding to key image frames are extracted from the target feature point set. The 3D coordinates and pixel coordinates of the key feature points are used as environmental feature point data corresponding to the visual dimension. This step is the final selection of feature points for calibration. By focusing on key image frames, it is ensured that the selected feature points have high information content and reliability, avoiding the use of redundant or unstable feature points. 3D coordinates and pixel coordinates are key information required for visual calibration and optimization. Once the key image frame is determined, the most representative feature points can be further selected from all tracked target feature points in the frame based on their distribution in the image, the magnitude of disparity change, or tracking stability, etc., as key feature points. Alternatively, all target feature points in the key image frame that meet certain quality standards (such as tracking length and small reprojection error) can be directly used as key feature points, and their 3D coordinates (obtained through triangulation or depth mapping) and pixel coordinates in the image can be extracted.

[0064] Based on this, a field-of-view threshold is first constructed based on the average disparity change information corresponding to each target feature point in the target feature point set, thereby quantifying the amount of geometric information contained in the feature points. Then, when the field-of-view change of the target robot exceeds the constructed field-of-view threshold, the system can intelligently determine key image frames, ensuring that only those image frames that provide sufficient new information or have significant geometric changes are considered. Finally, key feature points corresponding to these key image frames are extracted from the target feature point set, and their 3D coordinates and pixel coordinates are used as environmental feature point data corresponding to the visual dimension. This series of operations allows the subsequent construction of the global factor map to be based on a more refined and information-rich set of visual feature point observations, thus significantly reducing computational complexity and improving the efficiency of the calibration algorithm. Simultaneously, because the selected key feature points have greater disparity changes and stronger geometric constraints, they can provide more accurate pose estimation and depth information, thereby improving the accuracy and robustness of robot calibration and avoiding calibration errors caused by insufficient or unevenly distributed feature point information.

[0065] For example, suppose the target robot is equipped with a first image acquisition device and a second image acquisition device. After obtaining the set of target feature points, the system continuously calculates the average pixel displacement of each target feature point across consecutive frames and estimates its corresponding average disparity change by combining camera intrinsic parameters and robot motion information. For example, a threshold can be set whereby a feature point is considered to provide sufficient information when its average disparity change exceeds a preset value (e.g., moving more than 5 pixels on the image plane and corresponding 3D depth change exceeding 10 cm). Based on this average disparity change information, the system can construct a field of view threshold, for example, by calculating the median or 90th percentile of the average disparity changes of all target feature points. During robot movement, the system monitors the image sequence captured by the first image acquisition device in real time and calculates the camera pose change between the current frame and the previous frame. If the camera pose change of the current frame (e.g., translation distance exceeding 5 mm or rotation angle exceeding 0.5 degrees) is greater than the previously constructed field of view threshold, the current frame is marked as a key image frame. Subsequently, from the target feature point set, feature points that were observed in these key image frames and met certain quality standards (such as reprojection error less than 1 pixel) were selected. These selected feature points, along with their 3D coordinates obtained through triangulation or depth estimation and their pixel coordinates in the key image frames, were determined as the environmental feature point data corresponding to the visual dimensions for subsequent calibration.

[0066] In summary, the above processing effectively selects the most valuable visual feature points for robot calibration, avoiding the waste of computational resources and decreased calibration accuracy caused by processing a large number of redundant or low-quality feature points. This makes the subsequent calibration process based on the global factor map more efficient and accurate, thereby improving the robot's overall localization and perception capabilities. In practical implementation, when constructing the global factor map and performing calibration operations, effectively integrating data from different dimensions such as vision, inertia, and kinematics, and accurately determining the key parameters required for calibration, is a crucial issue affecting calibration accuracy and robustness. To address this, in this embodiment, the construction of a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and the calibration of the target robot based on the global factor map, includes: Based on the environmental feature point data and the extrinsic parameter data, a set of visual feature point observations associated with the visual dimension is constructed; based on the inertial data and the acceleration data, a set of inertial data measurements associated with the inertial dimension is constructed; and based on the kinematic data, a set of kinematic encoder measurements associated with the kinematic dimension is constructed. Visual factors are constructed based on the set of visual feature point observations; inertial factors are constructed based on the set of inertial data measurements; and kinematic factors are constructed based on the set of kinematic encoder measurements. A global factor graph is constructed based on the visual factors, the inertial factors, and the kinematic factors; and target extrinsic parameter data and time offset data are determined based on the global factor graph as an operation for calibrating the target robot.

[0067] Specifically, constructing a set of visual feature point observations associated with the visual dimension refers to converting environmental feature point data and extrinsic parameter data (e.g., the relative pose of the camera and the robot body) collected by visual sensors (e.g., cameras) into observation values ​​that can be used for factor graph optimization. This can be achieved, for example, by extracting feature points from image sequences using feature point detection and tracking algorithms (e.g., SIFT, ORB), and combining these with camera intrinsic and extrinsic parameters to obtain the positions of these feature points in three-dimensional space or their projected coordinates on the image plane as observation values. Alternatively, multi-view geometry principles can be utilized, such as triangulation, to convert feature point matching information from different viewpoints into three-dimensional point cloud data, and then combining this with camera pose information to form an observation set. Constructing a set of inertial data measurements associated with the inertial dimension refers to converting inertial data (e.g., angular velocity, linear acceleration) and acceleration data (determined by translational motion) collected by the inertial measurement unit (IMU) into measurement values ​​that can be used for factor graph optimization. This can be achieved, for example, by pre-integrating the raw IMU data, integrating the angular velocity and acceleration over a period of time to obtain the relative pose change, which is then used as the inertial measurement value. Alternatively, raw angular velocity and acceleration readings can be used directly, or filtered data can be combined with timestamps to form a sequence of measurement values. Constructing a set of kinematic encoder measurements associated with the kinematic dimensions refers to converting kinematic data (e.g., joint angles, joint velocities) collected by kinematic sensors such as robot joint encoders into a measurement form that can be used for factor graph optimization. This can, for example, calculate the relative displacement or velocity of robot joints based on encoder readings, and then deduce the relative pose change of the robot's end effector or base through the robot's forward kinematics model. Alternatively, the changes in joint angles or joint velocities can be directly used as measurement values, or they can be converted into pose changes in the robot's body coordinate system.

[0068] Accordingly, constructing visual factors based on the set of visual feature point observations involves creating a node in the factor graph that associates the visual observations with the state variables to be optimized (e.g., camera pose, 3D position of feature points, extrinsic parameters, etc.) and defines a residual function to measure the error between the observed and predicted values. This can be achieved, for example, by using reprojection error as the residual function, i.e., the difference between the observed pixel coordinates and the predicted pixel coordinates calculated based on the current state variables (camera pose, 3D position of feature points, extrinsic parameters). Alternatively, distance error based on the 3D position of feature points can be used, i.e., the distance between the estimated 3D point and the actual 3D point based on the visual observations. Constructing inertial factors based on the set of inertial data measurements involves creating a node in the factor graph that associates the inertial measurements with the state variables to be optimized (e.g., robot pose, velocity, IMU bias, etc.) and defines a residual function to measure the error between the measured and predicted values. This can be achieved, for example, by using a pre-integrated residual function to compare the relative pose change obtained through pre-integration with the relative pose change predicted by state variables (initial pose, velocity, IMU bias). Alternatively, the raw IMU data can be used directly to construct a residual function based on the IMU dynamics model. Constructing kinematic factors based on the set of kinematic encoder measurements involves creating a node in the factor graph that associates the kinematic encoder measurements with the state variables to be optimized (e.g., robot joint pose, robot body pose, etc.) and defines a residual function to measure the error between the measured and predicted values. This can be achieved, for example, by using a robot forward kinematics model to convert the encoder measurements into a relative pose change of the robot body and compare it with the relative pose change predicted by the state variables. Alternatively, residual functions based on joint angles or joint velocities can be constructed to directly constrain the joint states.

[0069] Accordingly, a global factor graph is constructed based on the visual factors, inertial factors, and kinematic factors. The target extrinsic parameters and time offset data are then determined based on this global factor graph. As a calibration operation for the target robot, this involves connecting the constructed visual factors, inertial factors, and kinematic factors to form a unified global factor graph. By optimizing this factor graph, the state variables that minimize all residuals are solved, including target extrinsic parameters (e.g., the relative pose between the camera and IMU) and time offset data (e.g., the time synchronization error between sensors). This can be achieved, for example, by using nonlinear optimization algorithms (such as Levenberg-Marquardt, Gauss-Newton, etc.) to solve the factor graph, iteratively adjusting the state variables until convergence, thereby obtaining the optimal extrinsic parameters and time offset data. Alternatively, open-source libraries based on graph optimization (such as GTSAM, Ceres Solver, etc.) can be used to construct and optimize the factor graph.

[0070] Based on this, data from different dimensions such as vision, inertia, and kinematics are first transformed into structured sets of observations or measurements, and then corresponding factors are constructed for each data type. These factors capture the mathematical relationships and error models between the data from each sensor and the parameters to be calibrated. Subsequently, these independent visual, inertial, and kinematic factors are organically integrated into a unified global factor graph. This integration method allows data from different sensors to be tightly coupled within a common optimization framework, thus simultaneously considering the constraint information provided by all sensors. By performing nonlinear optimization on this global factor graph, the system can iteratively solve for the state variables that minimize the residuals of all factors, including the key parameters required for robot calibration, namely target extrinsic parameters and time offset data. This multi-sensor data fusion strategy effectively utilizes the complementary advantages of each sensor, overcomes the limitations that may exist with a single sensor, and thus improves the accuracy and robustness of the calibration results.

[0071] For example, suppose the target robot is equipped with a stereo camera system, an inertial measurement unit (IMU), and joint encoders. During calibration, firstly, the stereo camera acquires a sequence of environmental images. Through feature point detection and matching, combined with triangulation principles, the three-dimensional coordinates of feature points in the environment and their projections onto the images can be obtained. These data, together with the camera's initial extrinsic parameters, constitute the set of visual feature point observations. Simultaneously, the IMU continuously outputs angular velocity and linear acceleration data. These raw inertial data undergo pre-integration processing to obtain the relative pose changes between adjacent time points, forming a set of inertial data measurements. Furthermore, the robot's joint encoders provide real-time feedback of joint angles. Through the robot's forward kinematics model, changes in joint angles are converted into relative pose changes of the robot's base or end effector, forming a set of kinematic encoder measurements.

[0072] After data preparation, the system begins constructing a factor graph. For visual data, a reprojection error factor is constructed for each visual feature point, comparing the observed pixel coordinates with the predicted pixel coordinates calculated based on the currently estimated camera pose, the 3D position of the feature point, and the camera-IMU extrinsic parameters. For inertial data, an inertial residual factor is constructed for each pre-integration segment, comparing the relative pose change obtained from pre-integration with the relative pose change predicted based on the currently estimated IMU pose, velocity, and IMU bias. For kinematic data, a kinematic residual factor is constructed for each kinematic measurement, comparing the relative pose change derived from the encoder with the relative pose change predicted based on the currently estimated robot body pose. Finally, all these visual, inertial, and kinematic factors are concatenated to form a global factor graph. This global factor graph is iteratively optimized using nonlinear optimization algorithms such as Levenberg-Marquardt, adjusting the state variables in the factor graph, including the camera-IMU extrinsic parameters (rotation and translation) and the time offset between the sensors, until the residuals converge to a minimum. The final optimized extrinsic parameter data and time offset data are the results of calibrating the target robot.

[0073] In summary, the above processing methods enable the structured integration of heterogeneous sensor data from different dimensions such as vision, inertial, and kinematics into a unified global factor graph optimization framework. This refined data processing and factor construction method allows the system to fully leverage the unique advantages and complementary information of each sensor, thereby achieving accurate estimation of target extrinsic parameters and time offset data during multi-sensor fusion calibration. This tightly coupled optimization strategy significantly improves the accuracy and robustness of robot calibration, effectively solving the problem of inaccurate parameter determination in multi-sensor data fusion, and providing a solid foundation for high-precision robot positioning, navigation, and control. In practical applications, the target extrinsic data and the time offset data can be determined by the following formula (2): (2) in, This indicates that the state variables are adjusted using iterative algorithms (such as the Gauss-Newton or Levenberg-Marquardt algorithms). This minimizes the total substitution value within the parentheses. This represents the set of inertial data measurements. This represents the set of observations of the visual feature points. Let represent the set of kinematic encoder measurements, k represent the index information corresponding to the inertial data observations in the set of inertial data measurements, c represent the index information corresponding to the visual feature point observations in the set of visual feature point observations, and m represent the index information corresponding to the kinematic encoder measurements in the set of kinematic encoder measurements. Represents the inertial pre-integral residual function. The projection residual function represents the visual feature reconstruction. Represents the kinematic prior residual function. The integral of the actual measured value. This represents the actual extracted pixel observation coordinates. This represents the measured value of relative pose change. Indicates The squared Mahalanobis distance of the covariance matrix. This represents the noise covariance matrix corresponding to the inertia dimension. The noise covariance matrix corresponding to the visual dimension. This represents the noise covariance matrix corresponding to the kinematic dimension.

[0074] Based on this, a unified nonlinear least-squares optimization problem is constructed by modeling the residual functions of visual, inertial, and kinematic data, and weighting the measurement uncertainties of different sensor modes using Mahalanobis distance. This method can effectively fuse data from heterogeneous sensors, overcoming the suboptimal or robust calibration results that may result from simply combining data. Through iterative optimization, the system can find a set of optimal calibration parameters (including target extrinsic data and time offset data) that statistically best explain all sensor observations while fully considering the noise characteristics of each sensor.

[0075] It should be noted that in the above nonlinear factor graph optimization framework, the global state vector that needs to be estimated... It can be defined as the following formula (3): (3) in, This represents the set of global state variables that the factor graph optimizer needs to solve. This indicates the total number of keyframes extracted within the sliding window or in global optimization. Indicates the first The IMU's ontological state vector in the world coordinate system at each keyframe moment. Its internal structure is specifically expanded as follows: ,in Represents the position vector. Represents the velocity vector. Represents a quaternion that characterizes rotation. This represents the accelerometer bias. This represents zero bias in the gyroscope. Indicates from the IMU base coordinate system ( To the left eye camera coordinate system The extrinsic parameters (extrinsic parameter matrix) of the extrinsic parameter matrix, including rotation extrinsic parameters. With translational external parameters . Indicates from the left eye camera coordinate system ( To the right eye camera coordinate system The binocular extrinsic parameters. This represents the time synchronization error (timestamp offset) between camera sensor data and IMU sensor data.

[0076] See Figure 2 , Figure 2 A flowchart of another robot calibration method according to an embodiment of this specification is shown, applied to a target robot, and specifically includes the following steps.

[0077] Step S202: In response to the calibration command submitted by the user through the robot controller for the target robot in the working environment, the environmental feature point data corresponding to the visual dimension and the inertial data corresponding to the inertial dimension of the target robot are collected.

[0078] Step S204: Control the target robot to perform translational and rotational movements, and determine the acceleration data associated with the inertial dimension and the extrinsic parameter data associated with the visual dimension based on the movement execution results.

[0079] Step S206: During the execution of the translational and rotational movements, kinematic data of the target robot in the kinematic dimension is collected.

[0080] Step S208: Construct a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and calibrate the target robot according to the global factor map.

[0081] Step S210: The calibrated target robot continues to perform its work tasks in the work environment.

[0082] This embodiment provides another robot calibration method, which belongs to the same inventive concept as the robot calibration method described above. Any details not described in detail can be found in the same or corresponding descriptions in the above embodiments, and will not be elaborated upon here. Specifically, the controller refers to a control terminal used to control the target robot. The user can actively submit calibration commands to the target robot through the control terminal. Correspondingly, the work task refers to the task that the target robot needs to perform, such as a handling task (e.g., handling heavy goods), a movement task (e.g., kicking a ball, doing gymnastics), or an interactive task (e.g., dancing, providing companionship), etc. This embodiment does not impose any limitations on these tasks.

[0083] Corresponding to the above method embodiments, this specification also provides embodiments of a robot calibration device. Figure 3 A schematic diagram of a robot calibration device according to one embodiment of this specification is shown. Figure 3 As shown, the device is applied to the target robot and includes: The triggering module 302 is configured to, in response to a calibration command triggered for the target robot, collect environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. Control module 304 is configured to control the target robot to perform translational and rotational movements, and to determine acceleration data associated with the inertial dimension and extrinsic data associated with the visual dimension based on the movement execution results; The acquisition module 306 is configured to acquire kinematic data of the target robot in the kinematic dimension during the execution of the translational and rotational movements. The construction module 308 is configured to construct a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and to calibrate the target robot according to the global factor map.

[0084] In an optional embodiment, the step of collecting environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot in response to a calibration command triggered for the target robot includes: In response to a calibration command triggered for the target robot, the target robot is controlled to maintain a stationary state for a set time; if the duration of the stationary state is greater than the set time, environmental feature point data corresponding to the visual dimension of the target robot is collected, as well as initial inertial data, gyroscope data and gravity data of the target robot are collected; the initial inertial data, the gyroscope data and the gravity data are used as the inertial data corresponding to the inertial dimension.

[0085] In an optional embodiment, controlling the target robot to perform translational and rotational movements, and determining acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the movement execution results, includes: The target robot is controlled to perform a translational movement, and acceleration data associated with the inertial dimension is determined based on the translational movement execution result; the target robot is controlled to perform a rotational movement, and rotational extrinsic data and relative extrinsic data associated with the visual dimension are determined based on the rotational movement execution result, and the rotational extrinsic data and the relative extrinsic data are used as extrinsic data.

[0086] In an optional embodiment, the acceleration data is determined by the following formula:

[0087] in, This represents acceleration data determined in an inertial local coordinate system. This represents the rotation matrix from the camera coordinate system to the inertial coordinate system. This represents the linear acceleration data in the camera coordinate system. This represents the angular acceleration data in the camera coordinate system. This represents the angular velocity data in the camera coordinate system. This represents the translation data from the origin of the inertial coordinate system to the origin of the camera coordinate system.

[0088] In an optional embodiment, the acquisition of environmental feature point data corresponding to the target robot in the visual dimension includes: A first image sequence associated with a first image acquisition device is acquired, wherein the target robot includes a first image acquisition device and a second image acquisition device. The first image sequence is divided into grids, and natural texture feature points are extracted from the divided image sequence according to a feature point extraction algorithm to obtain an initial feature point set. Each feature point in the initial feature point set is continuously tracked in the temporal dimension, and the first feature point determined after tracking is matched to the second image acquisition device to obtain initial depth data. Based on the initial depth data, a feature point matching set associated with the spatiotemporal dimension is determined. Anomaly detection is performed on the feature point matching set, and abnormal feature points are removed according to the anomaly detection results to obtain a target feature point set. Key feature points are determined in the target feature point set, and the feature point data corresponding to the key feature points are used as the environmental feature point data corresponding to the visual dimension.

[0089] In an optional embodiment, determining key feature points in the target feature point set and using the feature point data corresponding to the key feature points as the environmental feature point data corresponding to the visual dimension includes: A field-of-view threshold is constructed based on the average disparity change information corresponding to each target feature point in the target feature point set; when the field-of-view change of the target robot is greater than the field-of-view threshold, a key image frame is determined; key feature points corresponding to the key image frame are extracted from the target feature point set, and the three-dimensional coordinates and pixel coordinates corresponding to the key feature points are used as environmental feature point data corresponding to the visual dimension.

[0090] In an optional embodiment, the step of constructing a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and calibrating the target robot according to the global factor map, includes: Based on the environmental feature point data and the extrinsic parameter data, a set of visual feature point observations associated with the visual dimension is constructed; based on the inertial data and the acceleration data, a set of inertial data measurements associated with the inertial dimension is constructed; and based on the kinematic data, a set of kinematic encoder measurements associated with the kinematic dimension is constructed. Visual factors are constructed based on the set of visual feature point observations; inertial factors are constructed based on the set of inertial data measurements; and kinematic factors are constructed based on the set of kinematic encoder measurements. A global factor graph is constructed based on the visual factors, the inertial factors, and the kinematic factors; and target extrinsic parameter data and time offset data are determined based on the global factor graph as an operation for calibrating the target robot.

[0091] In an optional embodiment, the target extrinsic data and the time offset data are determined by the following formula:

[0092] in, This indicates that the state variables are adjusted through an iterative algorithm. This minimizes the total substitution value within the parentheses. This represents the set of inertial data measurements. This represents the set of observations of the visual feature points. Let represent the set of kinematic encoder measurements, k represent the index information corresponding to the inertial data observations in the set of inertial data measurements, c represent the index information corresponding to the visual feature point observations in the set of visual feature point observations, and m represent the index information corresponding to the kinematic encoder measurements in the set of kinematic encoder measurements. Represents the inertial pre-integral residual function. The projection residual function represents the visual feature reconstruction. Represents the kinematic prior residual function. The integral of the actual measured value. This represents the actual extracted pixel observation coordinates. This represents the measured value of relative pose change. Indicates The squared Mahalanobis distance of the covariance matrix. This represents the noise covariance matrix corresponding to the inertia dimension. The noise covariance matrix corresponding to the visual dimension. This represents the noise covariance matrix corresponding to the kinematic dimension.

[0093] The above is a schematic scheme of a robot calibration device according to this embodiment. It should be noted that the technical solution of this robot calibration device and the technical solution of the robot calibration method described above belong to the same concept. For details not described in detail in the technical solution of the robot calibration device, please refer to the description of the technical solution of the robot calibration method described above.

[0094] Corresponding to the above method embodiments, this specification also provides another embodiment of a robot calibration device. Figure 4 A schematic diagram of another robot calibration device provided in one embodiment of this specification is shown. Figure 4 As shown, the device is applied to the target robot and includes: The trigger instruction module 402 is configured to, in response to a calibration instruction submitted by a user through a robot controller for the target robot in the working environment, collect environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. The control action module 404 is configured to control the target robot to perform translational and rotational actions, and to determine acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the action execution results; The data acquisition module 406 is configured to acquire kinematic data of the target robot in the kinematic dimension during the execution of the translational and rotational movements. The subgraph construction module 408 is configured to construct a global factor graph based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and to calibrate the target robot according to the global factor graph; The task execution module 410 is configured to continue performing work tasks in the work environment via the calibrated target robot.

[0095] The above is an illustrative scheme of another robot calibration device according to this embodiment. It should be noted that the technical solution of this robot calibration device and the technical solution of the robot calibration method described above belong to the same concept. For details not described in detail in the technical solution of the robot calibration device, please refer to the description of the technical solution of the robot calibration method described above.

[0096] Figure 5 A structural block diagram of a computing device 500 according to one embodiment of this specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.

[0097] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0098] In one embodiment of this specification, the above-described components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0099] The computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 500 can also be a mobile or stationary server.

[0100] The processor 520 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described robot calibration method.

[0101] The above is a schematic representation of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the robot calibration method described above belong to the same concept. Details not described in detail in the technical solution of the computing device can be found in the description of the technical solution of the robot calibration method described above.

[0102] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described robot calibration method.

[0103] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the robot calibration method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the robot calibration method described above.

[0104] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described robot calibration method.

[0105] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described robot calibration method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described robot calibration method.

[0106] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0107] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0108] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0109] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0110] The preferred embodiments disclosed above are merely illustrative of this specification. Optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described in this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification.

Claims

1. A method for calibrating a robot, characterized in that, Applied to target robots, including: In response to a calibration command triggered for the target robot, environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot are collected. The target robot is controlled to perform translational and rotational movements, and acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension are determined based on the movement execution results. During the translational and rotational movements, kinematic data of the target robot in the kinematic dimension are collected. A global factor map is constructed based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and the target robot is calibrated according to the global factor map.

2. The robot calibration method according to claim 1, characterized in that, The step of collecting environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot in response to a calibration command triggered for the target robot includes: In response to a calibration command triggered on the target robot, the target robot is controlled to remain stationary for a set time. If the duration of the static state is greater than the set time, the environmental feature point data corresponding to the visual dimension of the target robot is collected, as well as the initial inertial data, gyroscope data and gravity data of the target robot are collected. The initial inertial data, the gyroscope data, and the gravity data are used as the inertial data corresponding to the inertial dimension.

3. The robot calibration method according to claim 1, characterized in that, The process of controlling the target robot to perform translational and rotational movements, and determining acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the movement execution results, includes: The target robot is controlled to perform a translational movement, and the acceleration data associated with the inertial dimension is determined based on the result of the translational movement. The target robot is controlled to perform a rotation action. Based on the result of the rotation action, rotation extrinsic data and relative extrinsic data associated with the visual dimension are determined, and the rotation extrinsic data and the relative extrinsic data are used as extrinsic data.

4. The robot calibration method according to claim 3, characterized in that, The acceleration data is determined by the following formula: in, This represents acceleration data determined in an inertial local coordinate system. This represents the rotation matrix from the camera coordinate system to the inertial coordinate system. This represents the linear acceleration data in the camera coordinate system. This represents the angular acceleration data in the camera coordinate system. This represents the angular velocity data in the camera coordinate system. This represents the translation data from the origin of the inertial coordinate system to the origin of the camera coordinate system.

5. The robot calibration method according to claim 2, characterized in that, The acquisition of environmental feature point data corresponding to the target robot in the visual dimension includes: Acquire a first image sequence associated with the target robot and a first image acquisition device, wherein the target robot includes a first image acquisition device and a second image acquisition device; The first image sequence is divided into grids, and natural texture feature points are extracted from the divided image sequence according to the feature point extraction algorithm to obtain an initial feature point set; For each feature point in the initial feature point set, continuous inter-frame tracking is performed in the temporal dimension. The first feature point determined after tracking is matched to the second image acquisition device to obtain initial depth data. Based on the initial depth data, a feature point matching set with associated spatiotemporal dimensions is determined. Anomaly detection is performed on the feature point matching set, and abnormal feature points are removed based on the anomaly detection results to obtain the target feature point set; Key feature points are determined from the target feature point set, and the feature point data corresponding to the key feature points are used as the environmental feature point data corresponding to the visual dimension.

6. The robot calibration method according to claim 5, characterized in that, The step of determining key feature points in the target feature point set and using the feature point data corresponding to the key feature points as the environmental feature point data corresponding to the visual dimension includes: A field of view threshold is constructed based on the average disparity change information corresponding to each target feature point in the target feature point set. When the change in the field of view of the target robot is greater than the field of view threshold, a key image frame is determined; Extract the key feature points corresponding to the key image frames from the target feature point set, and use the three-dimensional coordinates and pixel coordinates corresponding to the key feature points as the environmental feature point data corresponding to the visual dimension.

7. The robot calibration method according to claim 1, characterized in that, The process of constructing a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and calibrating the target robot according to the global factor map, includes: Based on the environmental feature point data and the extrinsic parameter data, a set of visual feature point observations associated with the visual dimension is constructed; based on the inertial data and the acceleration data, a set of inertial data measurement values ​​associated with the inertial dimension is constructed; and based on the kinematic data, a set of kinematic encoder measurement values ​​associated with the kinematic dimension is constructed. Visual factors are constructed based on the set of visual feature point observations, inertial factors are constructed based on the set of inertial data measurements, and kinematic factors are constructed based on the set of kinematic encoder measurements. A global factor map is constructed based on the visual factor, the inertial factor, and the kinematic factor, and target extrinsic parameter data and time offset data are determined based on the global factor map as an operation for calibrating the target robot.

8. The robot calibration method according to claim 7, characterized in that, The target extrinsic data and the time offset data are determined by the following formula: in, This indicates that the state variables are adjusted through an iterative algorithm. This minimizes the total substitution value within the parentheses. This represents the set of inertial data measurements. This represents the set of observations of the visual feature points. Let represent the set of kinematic encoder measurements, k represent the index information corresponding to the inertial data observations in the set of inertial data measurements, c represent the index information corresponding to the visual feature point observations in the set of visual feature point observations, and m represent the index information corresponding to the kinematic encoder measurements in the set of kinematic encoder measurements. Represents the inertial pre-integral residual function. The projection residual function represents the visual feature reconstruction. Represents the kinematic prior residual function. The integral of the actual measured value. This represents the actual extracted pixel observation coordinates. This represents the measured value of relative pose change. Indicated by The squared Mahalanobis distance of the covariance matrix. This represents the noise covariance matrix corresponding to the inertia dimension. The noise covariance matrix corresponding to the visual dimension. This represents the noise covariance matrix corresponding to the kinematic dimension.

9. A method for calibrating a robot, characterized in that, Applied to target robots, including: In response to a calibration command submitted by a user through a robot controller for the target robot in the working environment, the system collects environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. The target robot is controlled to perform translational and rotational movements, and acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension are determined based on the movement execution results. During the translational and rotational movements, kinematic data of the target robot in the kinematic dimension are collected. A global factor map is constructed based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and the target robot is calibrated according to the global factor map. The calibrated target robot continues to perform its work tasks in the work environment.

10. A calibration device for a robot, characterized in that, Applied to target robots, including: The triggering module is configured to, in response to a calibration command triggered for the target robot, collect environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. The control module is configured to control the target robot to perform translational and rotational movements, and to determine acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the movement execution results. The acquisition module is configured to acquire kinematic data of the target robot in the kinematic dimension during the execution of the translational and rotational movements. The construction module is configured to construct a global factor map based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and to calibrate the target robot according to the global factor map.

11. A calibration device for a robot, characterized in that, Applied to target robots, including: The trigger command module is configured to, in response to a calibration command submitted by a user through a robot controller for the target robot in the working environment, collect environmental feature point data corresponding to the visual dimension and inertial data corresponding to the inertial dimension of the target robot. The control motion module is configured to control the target robot to perform translational and rotational movements, and to determine acceleration data associated with the inertial dimension and extrinsic parameter data associated with the visual dimension based on the movement execution results. The data acquisition module is configured to acquire kinematic data of the target robot in the kinematic dimension during the execution of the translational and rotational movements. The subgraph construction module is configured to construct a global factor graph based on the environmental feature point data, the extrinsic parameter data, the inertial data, the acceleration data, and the kinematic data, and to calibrate the target robot according to the global factor graph; The task execution module is configured to continue performing work tasks in the work environment via the calibrated target robot.

12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 9.

13. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.

14. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.