Humanoid robot multi-mode visual calibration system and operation method
By constructing a closed-loop control system, efficient coordination of multimodal visual sensing information and automatic maintenance of hand-eye relationship are achieved, solving the accuracy and robustness problems of multimodal vision systems in complex environments and improving the accuracy and consistency of robot vision-guided operations.
Patent Information
- Application Number
- CN202511574883.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2025-12-23
Smart Images

Figure CN121179433A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, and in particular to a multimodal visual calibration system and operating method for humanoid robots. Background Technology
[0002] Robotics, especially humanoid robots with human-like form and operational capabilities, represents a cutting-edge area in automation and artificial intelligence. Among these technologies, the vision system, as a core component enabling robots to perceive their environment and interact with the physical world, is of paramount importance. With the increasing complexity of application scenarios, single-type vision sensors are no longer sufficient to meet the robot's need for comprehensive and robust perception of its surroundings. Therefore, vision systems integrating multiple modalities such as visible light, depth, and infrared thermal imaging have emerged, aiming to enhance the robot's perception capabilities under challenging conditions such as changing lighting, occlusion, and cluttered backgrounds by leveraging the complementary advantages of heterogeneous sensors. A key foundation of such multimodal vision systems is accurate hand-eye calibration, which involves determining the spatial transformation relationship between the vision sensor coordinate system and the robot's actuator (such as a robotic arm) coordinate system. The accuracy of this relationship directly determines the precision of the robot's vision-guided operations.
[0003] In existing technologies, hand-eye calibration of multimodal vision systems typically relies on calibrating each sensor separately and then connecting them in series through fixed spatial relationships, or on offline calibration using specific, high-contrast calibration objects under strictly controlled environments. This approach has a significant drawback: due to the large differences in physical characteristics and inconsistent data features of heterogeneous sensors, and the fact that factors such as vibration and temperature drift during robot operation can cause changes in the pre-defined spatial geometric relationships between sensors, traditional offline, step-by-step calibration methods struggle to maintain long-term, high-precision calibration results. This leads to a decrease in the positioning and manipulation accuracy of the robot based on multimodal vision during long-term operation or in dynamic environments, resulting in insufficient system robustness. Summary of the Invention
[0004] In view of this, the present invention provides a multimodal visual calibration system for humanoid robots. One or more embodiments of this specification also relate to a multimodal visual calibration method for humanoid robots, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the present invention, a humanoid robot multimodal visual calibration system is provided, comprising: A multimodal vision sensing and synchronous acquisition unit is used to synchronously acquire image data of a target through multiple rigidly mounted heterogeneous vision sensors; The multi-source data preprocessing and feature alignment unit is used to receive image data, perform coordinate unification and data augmentation processing, and generate aligned multi-source data. The multimodal contour fusion extraction unit is used to receive aligned multi-source data and generate high-precision target contour information and its three-dimensional spatial coordinates by fusing the feature information therein. The multi-sensor collaborative calibration and online correction unit is used to calculate and maintain the spatial geometric relationship between multiple heterogeneous vision sensors based on robot motion and target contour information, and to correct them in real time. The hand-eye calibration calculation unit is used to construct and solve the hand-eye transformation relationship based on the target contour information and its three-dimensional spatial coordinates, robot motion information, and spatial geometric relationships, using a weighted least squares algorithm. The calibration quality assessment and feedback unit is used to assess the quality of the hand-eye transition relationship and generate feedback signals based on the assessment results to optimize the data acquisition strategy. The feedback signals are transmitted to the multimodal visual sensing and synchronous acquisition unit to adjust its acquisition process and form a closed-loop control.
[0006] In some implementations, the multimodal visual sensing and synchronous acquisition unit includes an RGB camera, a depth camera, and an infrared thermal imager; synchronous acquisition is achieved by generating a hardware trigger signal through a synchronization controller to ensure that the exposure timestamps of the RGB camera, depth camera, and infrared thermal imager are aligned.
[0007] In some implementations, the operations performed by the multi-source data preprocessing and feature alignment unit include: mapping point cloud data acquired by the depth camera to the image coordinate system of the RGB camera, and resampling the thermal image acquired by the infrared thermal imager to the same resolution as the RGB image, thereby achieving pixel-level spatial alignment.
[0008] In some implementations, the multimodal contour fusion extraction unit includes a deep learning network that integrates a cross-modal attention mechanism to adaptively fuse features from the RGB modality, depth modality, and infrared modality, and output the target contour points and their corresponding contour extraction confidence scores.
[0009] In some implementations, the objective function constructed by the hand-eye calibration solution unit to solve the hand-eye transformation relationship is:
[0010] Where N represents the total number of valid robot pose pairs used for calibration; This represents the translational state weight of the i-th pose pair. Represents the rotational dynamic weights of the i-th pose pair; Let represent the homogeneous transformation matrix of the robotic arm end effector at the i-th pose relative to the base coordinate system, provided by the robot controller. This represents the homogeneous transformation matrix of the camera coordinate system relative to a fixed world coordinate system, calculated by the multimodal contour fusion extraction unit when observing the target in the i-th pose. This represents the hand-eye transformation relationship to be solved; It is the rotation matrix extracted from the corresponding homogeneous transformation matrix, and t is the rotation matrix and translation vector extracted from the corresponding homogeneous transformation matrix; The identity matrix is denoted by log(), which maps the rotation matrix to its Lie algebra vector. Denotes the Euclidean norm; It is a dimensionless equilibrium factor.
[0011] In some implementations, the formulas for calculating the translational dynamic weights and rotational dynamic weights include:
[0012]
[0013] in, The average contour extraction confidence score of the target contour points in the i-th pose reflects the model's confidence in the accuracy of target contour recognition in this observation. The average infrared feature stability of the target contour points in the i-th pose is obtained by taking the reciprocal of the standard deviation of the temperature value and normalizing it. This represents the average depth measurement uncertainty of the target contour points in the i-th pose, and the quantization represents the measurement noise or error range of the depth camera at these points. This is a preset constant used for depth value normalization. It is used to normalize depth uncertainty. Transform into a dimensionless ratio / This facilitates exponential calculations and the integration of parameters with different dimensions; , and All are positive zero exponential adjustment coefficients, used to adjust the sensitivity and contribution of the corresponding quality indicators to the weights. Controlling the confidence level of the average profile extraction The influence, when When <1, the penalty for low-confidence data is reduced; when When the value is greater than 1, the reward for high-confidence data is enhanced; Controlling infrared feature stability The influence, the way it works and same; Controlled average depth measurement uncertainty Translational weights The degree of inhibition, The larger the value, the greater the reduction in the weight of pose, which is unreliable in depth measurement, in translation calculations. It is a very small positive number, and it is added to the denominator to prevent the denominator from being 0.
[0014] In some implementations, the contour extraction confidence level C i The activation values of the output layer of the multimodal contour fusion extraction unit are obtained after normalization; infrared feature stability S i The depth measurement uncertainty D is obtained by normalizing the reciprocal of the standard deviation of the temperature values of corresponding contour points in multiple consecutive infrared images; i It is obtained by estimating the measurement variance of the depth camera at the corresponding contour points.
[0015] In some implementations, the multi-sensor collaborative calibration and online correction unit periodically detects the consistency of the projection of common feature points in the multimodal data during system operation, and fine-tunes the spatial geometric relationship parameters between multiple heterogeneous vision sensors based on the detection results.
[0016] In some implementations, the evaluation dimensions of the calibration quality assessment and feedback unit include accuracy, consistency, and robustness; wherein accuracy is evaluated by calculating the reprojection error; consistency is evaluated by cross-validating the measurement results of multiple heterogeneous vision sensors; and robustness is evaluated by injecting simulated disturbances into the system and performing stress tests.
[0017] According to a second aspect of the present invention, a multimodal visual calibration method for a humanoid robot is provided, the method being applied to the system of the aforementioned claim, the method comprising: Image data of the target is acquired synchronously through a multimodal visual sensing and synchronous acquisition unit; The acquired image data is processed by a multi-source data preprocessing and feature alignment unit to generate aligned multi-source data; High-precision target contour information and its three-dimensional spatial coordinates are extracted from aligned multi-source data using a multimodal contour fusion extraction unit. The spatial relationship between visual sensors is maintained and online correction is performed through a multi-sensor collaborative calibration and online correction unit; The hand-eye transformation relationship is solved by using a hand-eye calibration solution unit. The quality of the hand-eye transition relationship is evaluated through a calibration quality assessment and feedback unit, and a feedback signal is generated. The feedback signal is transmitted to the multimodal vision sensing and synchronous acquisition unit to adjust the subsequent data acquisition process, forming a closed-loop calibration process.
[0018] In some embodiments, the synchronous acquisition of target image data via the multimodal visual sensing and synchronous acquisition unit includes: A hardware trigger signal is generated by a synchronization controller to synchronously control an RGB camera, a depth camera, and an infrared thermal imager to acquire images, ensuring that the exposure timestamps of the RGB camera, the depth camera, and the infrared thermal imager are aligned.
[0019] In some embodiments, the process of processing the acquired image data through the multi-source data preprocessing and feature alignment unit to generate aligned multi-source data includes: The point cloud data acquired by the depth camera is mapped to the image coordinate system of the RGB camera, and the thermal image acquired by the infrared thermal imager is resampled to the same resolution as the RGB image captured by the RGB camera to achieve pixel-level spatial alignment.
[0020] In some implementations, the step of extracting high-precision target contour information and its three-dimensional spatial coordinates from the aligned multi-source data using a multimodal contour fusion extraction unit includes: Using a deep learning network with an integrated cross-modal attention mechanism, features from RGB, deep, and infrared modalities are adaptively fused, and the target contour points and their corresponding contour extraction confidence scores are output.
[0021] In some embodiments, solving the hand-eye transformation relationship through the hand-eye calibration calculation unit includes: Construct and solve the following objective function to obtain the hand-eye transformation relationship X:
[0022] Where N represents the total number of valid robot pose pairs used for calibration; This represents the translational state weight of the i-th pose pair. Represents the rotational dynamic weights of the i-th pose pair; Let represent the homogeneous transformation matrix of the robotic arm end effector at the i-th pose relative to the base coordinate system, provided by the robot controller. This represents the homogeneous transformation matrix of the camera coordinate system relative to a fixed world coordinate system, calculated by the multimodal contour fusion extraction unit when observing the target in the i-th pose. This represents the hand-eye transformation relationship to be solved; It is the rotation matrix extracted from the corresponding homogeneous transformation matrix, and t is the rotation matrix and translation vector extracted from the corresponding homogeneous transformation matrix; The identity matrix is denoted by log(), which maps the rotation matrix to its Lie algebra vector. Denotes the Euclidean norm; It is a dimensionless equilibrium factor.
[0023] In some implementations, the formulas for calculating the translational dynamic weights and the rotational dynamic weights include:
[0024]
[0025] in, The average contour extraction confidence score of the target contour points in the i-th pose reflects the model's confidence in the accuracy of target contour recognition in this observation. The average infrared feature stability of the target contour points in the i-th pose is obtained by taking the reciprocal of the standard deviation of the temperature value and normalizing it. This represents the average depth measurement uncertainty of the target contour points in the i-th pose, and the quantization represents the measurement noise or error range of the depth camera at these points. This is a preset constant used for depth value normalization. It is used to normalize depth uncertainty. Transform into a dimensionless ratio / This facilitates exponential calculations and the integration of parameters with different dimensions; , and All are positive zero exponential adjustment coefficients, used to adjust the sensitivity and contribution of the corresponding quality indicators to the weights. Controlling the confidence level of the average profile extraction The influence, when When <1, the penalty for low-confidence data is reduced; when When the value is greater than 1, the reward for high-confidence data is enhanced; Controlling infrared feature stability The influence, the way it works and same; Controlled average depth measurement uncertainty Translational weights The degree of inhibition, The larger the value, the greater the reduction in the weight of pose, which is unreliable in depth measurement, in translation calculations. It is a very small positive number, and it is added to the denominator to prevent the denominator from being 0.
[0026] In some implementations, maintaining the spatial relationship between visual sensors and performing online correction through the multi-sensor collaborative calibration and online correction unit includes: periodically detecting the projection consistency of common feature points in the environment in multimodal data during system operation, and fine-tuning the spatial geometric relationship parameters between the multiple heterogeneous visual sensors based on the detection results.
[0027] At least one embodiment of the present invention constructs a closed-loop control system that integrates synchronous acquisition, data alignment, feature fusion, online correction, hand-eye relationship calculation, and quality assessment feedback. This system achieves efficient coordination of multimodal visual sensing information and automatic maintenance and optimization of hand-eye relationships, thereby effectively overcoming the accuracy attenuation problem caused by inaccurate relationships between sensors in the prior art. It significantly improves the accuracy, consistency, and long-term reliability of humanoid robots in perception and operation based on multimodal vision in complex and variable environments. Attached Figure Description
[0028] Figure 1 This is a simplified structural diagram of a humanoid robot multimodal vision calibration system provided by the present invention; Figure 2 This is a flowchart of the operation method of a humanoid robot multimodal vision calibration system provided by the present invention. Detailed Implementation
[0029] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0030] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the one or more embodiments of this specification. The singular forms “a” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items. The modifications “a” and “a plurality” as used in this disclosure are illustrative and not restrictive, and those skilled in the art will understand that they should be understood as “one or more” unless the context clearly indicates otherwise.
[0031] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0032] See Figure 1 , Figure 1This diagram illustrates a simplified structural schematic of a humanoid robot multimodal vision calibration system according to some embodiments of this specification. Specifically, it includes: a multimodal vision sensing and synchronous acquisition unit for synchronously acquiring image data of a target using multiple rigidly mounted heterogeneous vision sensors; a multi-source data preprocessing and feature alignment unit for receiving image data and performing coordinate unification and data augmentation processing to generate aligned multi-source data; a multimodal contour fusion and extraction unit for receiving the aligned multi-source data and generating high-precision target contour information and its three-dimensional spatial coordinates by fusing the feature information therein; and multi-sensor collaborative calibration and online correction. The unit is used to calculate and maintain the spatial geometric relationship between multiple heterogeneous vision sensors based on robot motion and target contour information, and to correct it in real time; the hand-eye calibration solution unit is used to construct and solve the hand-eye transformation relationship based on the target contour information and its three-dimensional spatial coordinates, robot motion information, and spatial geometric relationship using a weighted least squares algorithm; the calibration quality assessment and feedback unit is used to assess the quality of the hand-eye transformation relationship and generate feedback signals based on the assessment results to optimize the data acquisition strategy. The feedback signals are transmitted to the multimodal vision sensing and synchronous acquisition unit to adjust its acquisition process and form a closed-loop control.
[0033] The multimodal vision sensing and synchronous acquisition unit refers to the hardware component and its control logic responsible for synchronously acquiring image data from multiple different types of sensors in the system. This unit generates hardware trigger signals through a synchronization controller to strictly align the exposure times of the RGB camera, depth camera, and infrared thermal imager, ensuring strict temporal synchronization of data from different modalities and laying the foundation for subsequent fusion processing. Multiple heterogeneous vision sensors can refer to a group of image acquisition devices that differ in physical principles and output data types. For example, in this system, it specifically refers to the RGB camera, depth camera, and infrared thermal imager rigidly mounted at the end of the robotic arm, capable of capturing target information from different physical dimensions.
[0034] The multi-source data preprocessing and feature alignment unit can refer to a software module that performs preliminary processing and spatial registration of raw data from different sensors. For example, this unit maps depth point clouds to the RGB image coordinate system and resamples thermal images to the same resolution as RGB images to achieve pixel-level spatial alignment.
[0035] The multimodal contour fusion extraction unit is a processing module that uses a deep learning model to extract target contour information from aligned multi-source data. This unit employs a convolutional neural network with an integrated cross-modal attention mechanism to adaptively fuse RGB, depth, and infrared features to generate high-precision target contours and their 3D coordinates. High-precision target contour information refers to the accurate geometric representation describing the edges of the target object's shape, consisting of a series of continuous 3D spatial point coordinates, along with the contour extraction confidence score for each point, used to accurately characterize the target's shape and position. 3D spatial coordinates can refer to the position representation of target contour points in a specific 3D coordinate system, such as the (X, Y, Z) coordinate values in the camera coordinate system in this system, used for subsequent calculation of the target's absolute pose in space.
[0036] The multi-sensor collaborative calibration and online correction unit refers to a functional module used to establish and maintain the spatial geometric relationship between multiple sensors and to correct this relationship in real time. This unit periodically detects the projection consistency of common feature points and performs parameter fine-tuning based on errors to ensure the long-term accuracy of multimodal data fusion. Spatial geometric relationship refers to the rotation and translation transformation parameters between different visual sensor coordinate systems. For example, initial values are obtained through offline calibration and continuously optimized by the online correction module during system operation to maintain accurate relative poses between sensors. Real-time correction refers to the continuous fine-tuning of sensor parameters during system operation. For example, by analyzing the projection deviation of feature points between consecutive frames, filtering algorithms are used to incrementally update extrinsic parameters to compensate for calibration parameter drift caused by vibration and temperature changes.
[0037] The hand-eye calibration solution unit can refer to an algorithm module specifically designed to calculate the relative pose transformation between the robot arm's end effector and the camera. For example, based on a series of robot poses and corresponding target observation data, this unit constructs a weighted least squares problem to solve the hand-eye transformation matrix X, which is used to establish the coordinate relationship between visual perception and robot motion control. The weighted least squares algorithm can refer to a mathematical method that assigns different weights to different observation data in least squares optimization. For example, in this system, the weights are dynamically calculated based on contour extraction confidence, infrared feature stability, and depth uncertainty, to make the calibration results more reliable with high-quality data. The hand-eye transformation relationship can refer to a homogeneous transformation matrix describing the fixed transformation relationship between the robot's end effector coordinate system and the camera coordinate system. It is represented by a 4x4 matrix containing rotation and translation components and is a core parameter for the robot to accurately perform visual guidance tasks.
[0038] The calibration quality assessment and feedback unit can refer to a module that performs multi-dimensional performance evaluation of the calculated hand-eye calibration results and generates control commands. For example, this unit evaluates accuracy, consistency, and robustness by calculating reprojection error, performing cross-validation, and stress testing to determine whether the calibration was successful or needs to be repeated. Feedback signals can refer to control commands output by the calibration quality assessment unit that guide adjustments to the data acquisition strategy. For instance, when evaluation indicators are not met, this signal instructs the system to acquire more data in specific postures to form a closed-loop optimization, gradually improving calibration quality. The data acquisition strategy can refer to a set of rules or parameters that control how the multimodal vision sensing unit acquires data, including trigger conditions, robot pose selection, and acquisition frequency, which can be dynamically adjusted based on feedback signals. Closed-loop control can refer to an automatic control system structure that includes execution, measurement, evaluation, feedback, and adjustment stages. For example, in this system, the quality assessment information of the calibration results is fed back to the acquisition unit, driving a new round of data acquisition and calibration optimization to automate and continuously improve the entire calibration process.
[0039] The present invention will be further described below through a detailed embodiment: The following describes in detail the workflow of the humanoid robot multimodal vision calibration system of the present invention, using a specific application scenario. This embodiment simulates the humanoid robot performing the task of identifying and locating a target object on a specific workbench in an indoor environment with changing lighting and some heat interference.
[0040] The multimodal vision sensing and synchronous acquisition unit employs a rigid mechanical structure to mount an RGB camera, a depth camera, and an infrared thermal imager. This unit uses a field-programmable gate array (FPGA)-based synchronization controller to generate precise hardware trigger signals. The rising edge of this signal simultaneously triggers all three sensors to perform image exposure and data acquisition, ensuring that the acquired RGB color image, depth point cloud data, and infrared thermal image are strictly aligned in time, with a time synchronization error of less than 1 millisecond. The unit outputs a stream of multimodal raw image data with perfectly matched timestamps.
[0041] The multi-source data preprocessing and feature alignment unit receives the raw data stream acquired synchronously. This unit first uses the intrinsic parameters of the depth camera and its initial extrinsic parameters with the RGB camera to map the depth point cloud data to the image coordinate system of the RGB camera through perspective projection transformation, generating a depth map that corresponds one-to-one with the pixels of the RGB image. Simultaneously, this unit uses a bicubic interpolation algorithm to resample the low-resolution thermal image acquired by the infrared thermal imager to the same 1920×1080 pixel resolution as the RGB image. Based on this, the unit also performs automatic white balance and contrast-limited adaptive histogram equalization (CLAHE) processing on the RGB image, bilateral filtering denoising on the depth map, and non-uniformity correction (NUC) based on fixed-pattern noise (FPN) on the thermal image. Finally, the unit outputs RGB-DT (red-green-blue-depth-temperature) multi-source data blocks aligned in both spatial coordinates and pixel resolution.
[0042] The core of the multimodal contour fusion extraction unit is a pre-trained deep learning network employing an encoder-decoder architecture and integrating a cross-modal attention mechanism. This unit receives aligned RGB-DT data blocks as input. In the network encoder, the RGB image, depth map, and thermal image are processed through independent convolutional neural network (CNN) branches to extract modality-specific features. Subsequently, the cross-modal attention mechanism adaptively calculates the weights of depth and thermal features on the RGB feature channels, achieving weighted fusion at the feature level. For example, in regions where the object's temperature differs significantly from the background, thermal features are given higher weights. The fused feature map is then used in the network decoder to progressively recover spatial details through upsampling and skip connections, ultimately outputting a probability map indicating that each pixel belongs to the target contour. Binarization is performed by setting a probability threshold of 0.5, generating a high-precision binary target contour mask that integrates color, geometric, and thermal information.
[0043] The multi-sensor collaborative calibration and online correction unit operates continuously during system operation. This unit uses the target contour information generated in the preceding steps as input and calculates the end effector pose using the robotic arm joint angle data provided by the robot controller. This unit periodically (e.g., every 30 seconds) detects the consistency of stable, easily identifiable feature points in multimodal data (such as table corners or corners of specific markers) across different sensor observations. By comparing the differences in the three-dimensional coordinates of the same physical feature point calculated based on RGB image features, depth point cloud features, and thermal image features in the sensor joint coordinate system, this unit fine-tunes using a variant of the Iterative Closest Point (ICP) algorithm, calculating and updating the spatial geometric transformation parameters (i.e., rotation matrices and translation vectors) between the sensors to compensate for small displacements caused by vibration or temperature changes, maintaining the accuracy of the calibration relationship between the sensors.
[0044] The core task of the hand-eye calibration solution unit is to solve for the hand-eye transformation matrix from the robot wrist coordinate system to the multi-sensor joint coordinate system. This unit utilizes the classic AX=XB equation model under an "eye-to-hand" configuration. Its inputs include: the 3D spatial corner coordinates of the same target object detected in different robot poses, provided by the multimodal contour fusion extraction unit (obtained by fusing corner detection from RGB images, 3D reconstruction from depth maps, and region centroids from thermal images); and the homogeneous transformation matrix of the robotic arm end effector relative to the robot's base coordinate system at the corresponding moment, provided by the robot controller. The unit employs a weighted least squares algorithm for solving the matrix, where the weight of each feature point pair is dynamically calculated based on its contour extraction confidence, depth measurement uncertainty, and infrared feature stability. This suppresses the influence of unreliable observations during the solution process, ultimately outputting the optimal hand-eye transformation matrix.
[0045] The calibration quality assessment and feedback unit performs multi-dimensional evaluation of the calculated hand-eye transformation matrix. This unit calculates the reprojection error after projecting known 3D feature points onto the image plane using the current hand-eye transformation matrix and robot pose matrix, as an accuracy indicator; it compares the dispersion of the 3D coordinates of the same target point calculated by the RGB camera, depth camera, and infrared thermal imager according to the current calibration relationship, as a consistency indicator; and it recalculates the hand-eye matrix by simulating small-amplitude sensor data perturbations (such as adding Gaussian noise) and observes the magnitude of the change, as a robustness indicator. If the evaluation results (e.g., an average reprojection error exceeding 3 pixels) indicate a decrease in calibration quality, this unit generates a feedback signal.
[0046] This feedback signal is transmitted to the multimodal vision sensing and synchronous acquisition unit. This feedback signal may contain instructions, such as a request for the system to switch to a higher-frequency "calibration-optimized" data acquisition mode. In this mode, the synchronous controller increases the acquisition trigger frequency from the usual 30 Hz to 60 Hz and controls the robotic arm to execute a series of specific motion trajectories designed specifically for calibration, with larger rotational and translational baselines, to acquire richer, geometrically more constrained observational data, thus providing a higher-quality data foundation for the next hand-eye calibration calculation. This forms a closed-loop control process from perception, processing, calculation to evaluation and then feedback back to perception, ensuring the stability of the system's calibration accuracy during long-term operation.
[0047] The beneficial effects of one of the embodiments in this specification include at least the following: by constructing a closed-loop control system that integrates synchronous acquisition, data alignment, feature fusion, online correction, hand-eye relationship calculation, and quality assessment feedback, efficient coordination of multimodal visual sensing information and automatic maintenance and optimization of hand-eye relationship are realized. This effectively overcomes the accuracy decay problem caused by inaccurate relationships between sensors in the prior art, and significantly improves the accuracy, consistency, and long-term reliability of humanoid robots in perception and operation based on multimodal vision in complex and variable environments.
[0048] In some implementations, the multimodal visual sensing and synchronous acquisition unit includes an RGB camera, a depth camera, and an infrared thermal imager; synchronous acquisition is achieved by generating a hardware trigger signal through a synchronization controller to ensure that the exposure timestamps of the RGB camera, depth camera, and infrared thermal imager are aligned.
[0049] An RGB camera can refer to a sensor that generates color images by capturing the three primary colors of red, green, and blue in the visible light band. For example, in this system, a global shutter type industrial camera is used to provide high-resolution color and texture information of the target. A depth camera can refer to an active vision sensor that can directly acquire distance information from each point in the scene to the camera. For example, a depth camera using the time-of-flight principle calculates depth by emitting modulated infrared light and measuring its returning phase difference, used to generate 3D point cloud data of the target. An infrared thermal imager can refer to a sensor that generates temperature distribution images by detecting infrared radiation emitted from the surface of an object. For example, a thermal imager using uncooled microbolometer technology is used to capture the thermal characteristics of the target, assisting in identification under low light or occlusion conditions. A synchronization controller can refer to a hardware unit that generates precise timing signals to coordinate the simultaneous operation of multiple devices. For example, this controller generates TTL level hardware trigger pulses and sends them simultaneously to the RGB camera, depth camera, and infrared thermal imager to ensure that all sensors begin exposure at the same physical moment. Hardware trigger signals can refer to precisely timed level or pulse signals generated by dedicated hardware circuits, such as a rising-edge valid digital signal with edge precision on the microsecond level. These signals are used to directly control the sensor's exposure start time, achieving frame-level synchronization. Exposure timestamps can refer to the precise moment when the sensor's photoelectric conversion elements begin accumulating photocharge, such as a time value latched by the camera's internal clock or external synchronization signal. These are used to mark the absolute or relative time of each frame's acquisition at the data level. Exposure timestamp alignment refers to ensuring a high degree of consistency in the acquisition start times of image data recorded by different sensors. For example, by sharing the same hardware trigger source, the difference in exposure start times between cameras can be less than a very small threshold, eliminating motion blur or registration errors caused by timing differences.
[0050] By employing a multimodal visual sensing and synchronous acquisition unit that includes an RGB camera, a depth camera, and an infrared thermal imager, and using a synchronization controller to generate hardware trigger signals to achieve exposure timestamp alignment, it is possible to ensure that the target information perceived from different physical dimensions has strict temporal consistency. This effectively avoids the data fusion misalignment problem caused by asynchronous sensor acquisition times, laying a reliable data foundation for subsequent high-precision multimodal feature fusion and target contour extraction.
[0051] In some implementations, the operations performed by the multi-source data preprocessing and feature alignment unit include: mapping point cloud data acquired by the depth camera to the image coordinate system of the RGB camera, and resampling the thermal image acquired by the infrared thermal imager to the same resolution as the RGB image, thereby achieving pixel-level spatial alignment.
[0052] Point cloud data refers to a set of coordinates representing a series of three-dimensional points on the surface of an object, acquired by a depth sensor. Each point contains X, Y, and Z coordinates, and may also include information such as reflection intensity, used to describe the three-dimensional geometry of the target object. An RGB camera's image coordinate system refers to a two-dimensional planar coordinate system with the top-left corner of the RGB image as the origin and pixels as the unit. For example, the horizontal axis U points to the right, and the vertical axis V points downwards, used to define the position of each pixel in the image. A thermal image refers to a grayscale image acquired by an infrared thermal imager that reflects the temperature distribution at various points in a scene. For example, the grayscale value of each pixel in the image corresponds to the infrared radiation intensity at that point, and can be converted into a temperature value after calibration. Resampling refers to the process of transforming an image from one resolution or grid to another and recalculating the pixel values. For example, using bilinear interpolation or cubic convolution interpolation algorithms, the value of a pixel at a new location is estimated based on the surrounding pixel values. Pixel-level spatial alignment refers to the state where, after geometric transformation, the same physical point in the scene corresponding to image data from different sources falls on the same pixel coordinate position in all images. For example, the edge of a target object is located in the exact same pixel row and column in both the RGB image and the resampled thermal image, which is used to achieve accurate fusion of information from different modalities.
[0053] By mapping point cloud data acquired by a depth camera to the image coordinate system of an RGB camera, and resampling the thermal image acquired by an infrared thermal imager to the same resolution as the RGB image, accurate spatial registration of visual data from different modalities can be achieved. This provides the necessary conditions for subsequent cross-modal feature fusion and target contour extraction based on pixel-level operations, effectively improving the accuracy and reliability of multi-source information collaborative processing.
[0054] In some implementations, the multimodal contour fusion extraction unit includes a deep learning network that integrates a cross-modal attention mechanism to adaptively fuse features from the RGB modality, depth modality, and infrared modality, and output the target contour points and their corresponding contour extraction confidence scores.
[0055] Cross-modal attention mechanisms can refer to neural network components that dynamically calculate and integrate feature weights from different modal data sources. For example, this mechanism generates an attention weight map by calculating the correlation score between RGB feature maps, depth feature maps, and infrared feature maps, guiding the network to focus on regions where information from different modalities is complementary. RGB modality can refer to data provided by RGB color cameras, containing scene color and texture information. For instance, RGB image data input to the network is typically normalized to a specific numerical range and processed as a three-channel tensor to provide the appearance features of the target. Depth modality can refer to data provided by depth cameras, representing scene geometry and object distance information. For example, a single-channel depth map, where pixel values represent distance, is normalized and used as an additional input channel to provide the three-dimensional structural information of the target. Infrared modality can refer to data provided by infrared thermal imagers, reflecting the surface temperature distribution characteristics of objects. For example, a single-channel thermal image, where pixel values are temperature-dependent, is preprocessed and input to the network to provide the thermal radiation characteristics of the target, aiding in identification under complex conditions. Adaptive fusion refers to the process of dynamically adjusting the combination of different information sources based on the specific content of the input data. For example, cross-modal attention mechanisms generate different weight coefficients for each spatial location and feature channel, which are used to weight and sum feature maps from different modalities, achieving data-driven feature integration. Target contour points refer to a series of discrete two-dimensional or three-dimensional coordinate points that describe the boundary of a target object, predicted by the model. For example, the model outputs a single-channel probability map with the same resolution as the input image. Post-processing (such as non-maximum suppression) extracts pixel positions with probabilities higher than a threshold as contour points to accurately delineate the shape of the target. It should be noted that these contour points... Contour extraction confidence refers to a quantitative indicator of the model's certainty about the accuracy of the extracted contour point location. For example, a value between zero and one generated by the network output layer (such as the sigmoid activation function) is used as a weighting basis in subsequent hand-eye calibration calculations to measure the reliability of the observation data at that point.
[0056] By employing a deep learning network with an integrated cross-modal attention mechanism in the multimodal contour fusion extraction unit, feature information from RGB, depth, and infrared modalities can be adaptively balanced and fused. This allows the target contour extraction process to fully utilize the complementary advantages of different modal data, thereby generating high-precision and high-reliability target contour information and its confidence level even under complex lighting or partial occlusion conditions. This provides a high-quality data foundation for the subsequent high-precision solution of hand-eye calibration.
[0057] In some implementations, the objective function constructed by the hand-eye calibration solution unit to solve the hand-eye transformation relationship is:
[0058] Where N represents the total number of valid robot pose pairs used for calibration; This represents the translational state weight of the i-th pose pair. Represents the rotational dynamic weights of the i-th pose pair; Let represent the homogeneous transformation matrix of the robotic arm end effector at the i-th pose relative to the base coordinate system, provided by the robot controller. This represents the homogeneous transformation matrix of the camera coordinate system relative to a fixed world coordinate system, calculated by the multimodal contour fusion extraction unit when observing the target in the i-th pose. This represents the hand-eye transformation relationship to be solved; It is the rotation matrix extracted from the corresponding homogeneous transformation matrix, and t is the rotation matrix and translation vector extracted from the corresponding homogeneous transformation matrix; The identity matrix is denoted by log(), which maps the rotation matrix to its Lie algebra vector. Denotes the Euclidean norm; It is a dimensionless equilibrium factor.
[0059] As a concrete example: after the system has acquired N sets of valid robot pose pairs (e.g., N=20), the hand-eye calibration unit begins operation. Its input data includes: 20 homogeneous transformation matrices of the robotic arm end effectors relative to the base coordinate system, read from the robot controller. And the homogeneous transformation matrix of the camera in 20 poses relative to a world coordinate system fixed on a calibration plate, calculated by the multimodal contour fusion extraction unit. The unit first starts from each and Extract the rotation matrix part ( , Translation vector part () , Then, according to the aforementioned formula, the translational state weights are calculated for each pose pair. and rotational dynamic weights Next, construct a matrix X (containing the rotation matrix) for the hand-eye transformation to be determined. Translation vector The nonlinear least squares objective function is given. This function consists of two parts: a translation error term, which measures the difference in translation vectors at both ends of the equation AX=XB; and a rotation error term, which measures the difference in rotation matrices using Lie algebras. A balancing factor λ (e.g., set to 1) is used to adjust the relative importance of the two terms. Finally, the Levenberg-Marquardt algorithm is used to iteratively optimize the objective function, minimizing the overall error until convergence, outputting the optimal hand-eye transformation matrix X.
[0060] By constructing and solving the objective function based on the weighted least squares method through the hand-eye calibration solution unit, the hand-eye calibration problem can be transformed into an optimization problem that considers the differences in data quality. By using translational dynamic weights and rotational dynamic weights to adaptively adjust the influence of different poses on the observation data in the solution, the influence of noise and outliers can be effectively suppressed, thereby improving the calibration accuracy and robustness of the hand-eye transformation matrix.
[0061] In some implementations, the formulas for calculating the translational dynamic weights and rotational dynamic weights include:
[0062]
[0063] in, The average contour extraction confidence score of the target contour points in the i-th pose reflects the model's confidence in the accuracy of target contour recognition in this observation. The average infrared feature stability of the target contour points in the i-th pose is obtained by taking the reciprocal of the standard deviation of the temperature value and normalizing it. This represents the average depth measurement uncertainty of the target contour points in the i-th pose, and the quantization represents the measurement noise or error range of the depth camera at these points. This is a preset constant used for depth value normalization. It is used to normalize depth uncertainty. Transform into a dimensionless ratio / This facilitates exponential calculations and the integration of parameters with different dimensions; , and All are positive zero exponential adjustment coefficients, used to adjust the sensitivity and contribution of the corresponding quality indicators to the weights. Controlling the confidence level of the average profile extraction The influence, when When <1, the penalty for low-confidence data is reduced; when When the value is greater than 1, the reward for high-confidence data is enhanced; Controlling infrared feature stability The influence, the way it works and same; Controlled average depth measurement uncertainty Translational weights The degree of inhibition, The larger the value, the greater the reduction in the weight of pose, which is unreliable in depth measurement, in translation calculations. It is a very small positive number, and it is added to the denominator to prevent the denominator from being 0.
[0064] By employing a dynamic weighting calculation method based on contour extraction confidence, infrared feature stability, and depth measurement uncertainty, a weight value that accurately reflects the quality of each robot pose pair's observation data can be assigned. This allows the hand-eye calibration process to adaptively rely on high-quality data while mitigating the impact of low-quality data, thereby significantly improving the accuracy and robustness of the calibration results and overcoming the shortcomings of traditional methods that treat all data equally.
[0065] In some implementations, the multi-sensor collaborative calibration and online correction unit periodically detects the consistency of the projection of common feature points in the multimodal data during system operation, and fine-tunes the spatial geometric relationship parameters between multiple heterogeneous vision sensors based on the detection results.
[0066] Shared feature points refer to points that can be clearly identified and matched in data simultaneously acquired by multiple heterogeneous vision sensors, representing the same physical location in a scene. Examples include a high-contrast corner, a specific texture patch, or a temperature-significant region, used as a benchmark for multi-sensor data association. Multimodal data refers to datasets of the same scene acquired by different types of sensors, differing in data format and physical meaning. Examples include coexisting RGB images, depth point clouds, and infrared thermal images, providing complementary information about the scene. Projection consistency refers to the degree to which the image features (such as position and descriptors) of the same 3D spatial point can be correctly matched when projected onto their respective image planes in different sensor coordinate systems according to their respective imaging models. For example, a corner point can be detected and its position matches in the edge maps of both the RGB and depth images, used to verify the accuracy of calibration parameters between sensors. Projection error refers to the difference between the observed feature point image coordinates and the image coordinates predicted based on sensor calibration parameters and the 3D point position, usually expressed as pixel distance, such as reprojection error, used to quantify calibration accuracy or the degree of alignment between sensors. Spatial geometric relationship parameters can refer to mathematical parameters that describe the relative position and orientation between different sensor coordinate systems, such as rotation matrices and translation vectors. These parameters define rigid body transformations between coordinate systems.
[0067] By periodically detecting the projection consistency of common feature points in the environment in multimodal data during system operation, and fine-tuning the spatial geometric relationship parameters between sensors based on the projection error, it is possible to compensate for the attenuation of calibration parameters caused by environmental interference or mechanical deformation in real time, maintain the accurate spatial alignment necessary for multi-sensor data fusion, and thus ensure the stability and reliability of the entire vision system in long-term operation.
[0068] Corresponding to the above system embodiments, this specification also provides embodiments of the operation method of a humanoid robot multimodal vision calibration system. Figure 2 A flowchart illustrating the operation method of a humanoid robot multimodal vision calibration system provided in some embodiments of this specification is shown. Figure 2 As shown, the specific steps include: Step 201: Simultaneously acquire image data of the target through the multimodal visual sensing and synchronous acquisition unit; Step 202: The acquired image data is processed by the multi-source data preprocessing and feature alignment unit to generate aligned multi-source data; Step 203: Extract high-precision target contour information and its three-dimensional spatial coordinates from aligned multi-source data using a multimodal contour fusion extraction unit; Step 204: Maintain the spatial relationship between visual sensors and perform online calibration through a multi-sensor collaborative calibration and online correction unit; Step 205: Solve the hand-eye transformation relationship using the hand-eye calibration solution unit; Step 206: Evaluate the quality of the hand-eye transition relationship through a calibration quality assessment and feedback unit, and generate a feedback signal; Step 207: The feedback signal is transmitted to the multimodal visual sensing and synchronous acquisition unit to adjust the subsequent data acquisition process and form a closed-loop calibration process.
[0069] In some embodiments, the synchronous acquisition of target image data via the multimodal visual sensing and synchronous acquisition unit includes: A hardware trigger signal is generated by a synchronization controller to synchronously control an RGB camera, a depth camera, and an infrared thermal imager to acquire images, ensuring that the exposure timestamps of the RGB camera, the depth camera, and the infrared thermal imager are aligned.
[0070] In some embodiments, the process of processing the acquired image data through the multi-source data preprocessing and feature alignment unit to generate aligned multi-source data includes: The point cloud data acquired by the depth camera is mapped to the image coordinate system of the RGB camera, and the thermal image acquired by the infrared thermal imager is resampled to the same resolution as the RGB image captured by the RGB camera to achieve pixel-level spatial alignment.
[0071] In some implementations, the step of extracting high-precision target contour information and its three-dimensional spatial coordinates from the aligned multi-source data using a multimodal contour fusion extraction unit includes: Using a deep learning network with an integrated cross-modal attention mechanism, features from RGB, deep, and infrared modalities are adaptively fused, and the target contour points and their corresponding contour extraction confidence scores are output.
[0072] In some embodiments, solving the hand-eye transformation relationship through the hand-eye calibration calculation unit includes: Construct and solve the following objective function to obtain the hand-eye transformation relationship X:
[0073] Where N represents the total number of valid robot pose pairs used for calibration; This represents the translational state weight of the i-th pose pair. Represents the rotational dynamic weights of the i-th pose pair; Let represent the homogeneous transformation matrix of the robotic arm end effector at the i-th pose relative to the base coordinate system, provided by the robot controller. This represents the homogeneous transformation matrix of the camera coordinate system relative to a fixed world coordinate system, calculated by the multimodal contour fusion extraction unit when observing the target in the i-th pose. This represents the hand-eye transformation relationship to be solved; It is the rotation matrix extracted from the corresponding homogeneous transformation matrix, and t is the rotation matrix and translation vector extracted from the corresponding homogeneous transformation matrix; The identity matrix is denoted by log(), which maps the rotation matrix to its Lie algebra vector. Denotes the Euclidean norm; It is a dimensionless equilibrium factor.
[0074] In some implementations, the formulas for calculating the translational dynamic weights and the rotational dynamic weights include:
[0075]
[0076] in, The average contour extraction confidence score of the target contour points in the i-th pose reflects the model's confidence in the accuracy of target contour recognition in this observation. The average infrared feature stability of the target contour points in the i-th pose is obtained by taking the reciprocal of the standard deviation of the temperature value and normalizing it. This represents the average depth measurement uncertainty of the target contour points in the i-th pose, and the quantization represents the measurement noise or error range of the depth camera at these points. This is a preset constant used for depth value normalization. It is used to normalize depth uncertainty. Transform into a dimensionless ratio / This facilitates exponential calculations and the integration of parameters with different dimensions; , and All are positive zero exponential adjustment coefficients, used to adjust the sensitivity and contribution of the corresponding quality indicators to the weights. Controlling the confidence level of the average profile extraction The influence, when When <1, the penalty for low-confidence data is reduced; when When the value is greater than 1, the reward for high-confidence data is enhanced; Controlling infrared feature stability The influence, the way it works and same; Controlled average depth measurement uncertainty Translational weights The degree of inhibition, The larger the value, the greater the reduction in the weight of pose, which is unreliable in depth measurement, in translation calculations. It is a very small positive number, and it is added to the denominator to prevent the denominator from being 0.
[0077] In some implementations, maintaining the spatial relationship between visual sensors and performing online correction through the multi-sensor collaborative calibration and online correction unit includes: periodically detecting the projection consistency of common feature points in the environment in multimodal data during system operation, and fine-tuning the spatial geometric relationship parameters between the multiple heterogeneous visual sensors based on the detection results.
[0078] The above is an illustrative scheme of a humanoid robot multimodal visual calibration method according to this embodiment. It should be noted that the technical solution of this humanoid robot multimodal visual calibration method and the technical solution of the humanoid robot multimodal visual calibration system described above belong to the same concept. For details not described in detail in the technical solution of the humanoid robot multimodal visual calibration method, please refer to the description of the technical solution of the humanoid robot multimodal visual calibration system described above.
[0079] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0080] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this invention. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A multimodal visual calibration system for humanoid robots, characterized in that, include: A multimodal vision sensing and synchronous acquisition unit is used to synchronously acquire image data of a target through multiple rigidly mounted heterogeneous vision sensors; A multi-source data preprocessing and feature alignment unit is used to receive the image data, perform coordinate unification and data augmentation processing, and generate aligned multi-source data; The multimodal contour fusion extraction unit is used to receive the aligned multi-source data and generate high-precision target contour information and its three-dimensional spatial coordinates by fusing the feature information therein. A multi-sensor collaborative calibration and online correction unit is used to calculate and maintain the spatial geometric relationship between the multiple heterogeneous vision sensors based on the robot's motion and the target contour information, and to correct them in real time. The hand-eye calibration calculation unit is used to construct and solve the hand-eye transformation relationship based on the target contour information and its three-dimensional spatial coordinates, the robot motion information, and the spatial geometric relationship, using a weighted least squares algorithm. The calibration quality assessment and feedback unit is used to assess the quality of the hand-eye transition relationship and generate a feedback signal based on the assessment result to optimize the data acquisition strategy. The feedback signal is transmitted to the multimodal visual sensing and synchronous acquisition unit to adjust its acquisition process and form a closed-loop control.
2. The system according to claim 1, characterized in that, The multimodal visual sensing and synchronous acquisition unit includes an RGB camera, a depth camera, and an infrared thermal imager; the synchronous acquisition is achieved by a synchronization controller generating a hardware trigger signal to ensure that the exposure timestamps of the RGB camera, the depth camera, and the infrared thermal imager are aligned.
3. The system according to claim 2, characterized in that, The operations performed by the multi-source data preprocessing and feature alignment unit include: mapping the point cloud data acquired by the depth camera to the image coordinate system of the RGB camera, and resampling the thermal image acquired by the infrared thermal imager to the same resolution as the RGB image captured by the RGB camera, thereby achieving pixel-level spatial alignment.
4. The system according to claim 1, characterized in that, The multimodal contour fusion extraction unit includes a deep learning network that integrates a cross-modal attention mechanism to adaptively fuse features from RGB, depth and infrared modes and output target contour points and their corresponding contour extraction confidence scores.
5. The system according to claim 1, characterized in that, The objective function constructed by the hand-eye calibration solution unit for solving the hand-eye transformation relationship is: Where N represents the total number of valid robot pose pairs used for calibration; This represents the translational state weight of the i-th pose pair. Represents the rotational dynamic weights of the i-th pose pair; Let represent the homogeneous transformation matrix of the robotic arm end effector at the i-th pose relative to the base coordinate system, provided by the robot controller. This represents the homogeneous transformation matrix of the camera coordinate system relative to a fixed world coordinate system, calculated by the multimodal contour fusion extraction unit when observing the target in the i-th pose. This represents the hand-eye transformation relationship to be solved; It is the rotation matrix extracted from the corresponding homogeneous transformation matrix, and t is the rotation matrix and translation vector extracted from the corresponding homogeneous transformation matrix; The identity matrix is denoted by log(), which maps the rotation matrix to its Lie algebra vector. Denotes the Euclidean norm; It is a dimensionless equilibrium factor.
6. The system according to claim 5, characterized in that, The formulas for calculating the translational dynamic weights and the rotational dynamic weights include: in, The average contour extraction confidence score of the target contour points in the i-th pose reflects the model's confidence in the accuracy of target contour recognition in this observation. The average infrared feature stability of the target contour points in the i-th pose is obtained by taking the reciprocal of the standard deviation of the temperature value and normalizing it. This represents the average depth measurement uncertainty of the target contour points in the i-th pose, and the quantization represents the measurement noise or error range of the depth camera at these points. A preset constant is used for normalizing depth values; used to reduce depth uncertainty. Transform into a dimensionless ratio / This facilitates exponential calculations and the integration of parameters with different dimensions; , and All are positive zero exponential adjustment coefficients, used to adjust the sensitivity and contribution of the corresponding quality indicators to the weights. Controlling the confidence level of the average profile extraction The influence, when When <1, the penalty for low-confidence data is reduced; when When the value is greater than 1, the reward for high-confidence data is enhanced; Controlling infrared feature stability The influence, the way it works and same; Controlled average depth measurement uncertainty Translational weights The degree of inhibition, The larger the value, the greater the reduction in the weight of pose, which is unreliable in depth measurement, in translation calculations. It is a very small positive number, and it is added to the denominator to prevent the denominator from being 0.
7. The system according to claim 1, characterized in that, During system operation, the multi-sensor collaborative calibration and online correction unit periodically detects the consistency of the projection of common feature points in the multimodal data, and fine-tunes the spatial geometric relationship parameters between the multiple heterogeneous visual sensors based on the detection results.
8. A method for operating a multimodal vision calibration system for a humanoid robot, characterized in that, The method is applied to the system according to any one of claims 1 to 7, and the method comprises: The multimodal visual sensing and synchronous acquisition unit synchronously acquires image data of the target; The acquired image data is processed by the multi-source data preprocessing and feature alignment unit to generate aligned multi-source data; The multimodal contour fusion extraction unit extracts high-precision target contour information and its three-dimensional spatial coordinates from the aligned multi-source data. The multi-sensor collaborative calibration and online correction unit maintains the spatial relationship between visual sensors and performs online correction. The hand-eye calibration calculation unit is used to solve the hand-eye transformation relationship; The calibration quality assessment and feedback unit evaluates the quality of the hand-eye transformation relationship and generates a feedback signal. The feedback signal is transmitted to the multimodal visual sensing and synchronous acquisition unit to adjust the subsequent data acquisition process, forming a closed-loop calibration process.
9. The method according to claim 8, characterized in that, The method of synchronously acquiring image data of the target through a multimodal visual sensing and synchronous acquisition unit includes: A hardware trigger signal is generated by a synchronization controller to synchronously control an RGB camera, a depth camera, and an infrared thermal imager to acquire images, ensuring that the exposure timestamps of the RGB camera, the depth camera, and the infrared thermal imager are aligned.
10. The method according to claim 9, characterized in that, The process of processing the acquired image data through the multi-source data preprocessing and feature alignment unit to generate aligned multi-source data includes: The point cloud data acquired by the depth camera is mapped to the image coordinate system of the RGB camera, and the thermal image acquired by the infrared thermal imager is resampled to the same resolution as the RGB image captured by the RGB camera to achieve pixel-level spatial alignment.
Citation Information
Cited By
Magnesium alloy electric drive main shell mold wear detection method and system based on machine vision
CN121544592A
Control method of industrial robot and related device
CN122008260A
Control method of industrial robot and related device
CN122008260B