Kinematics calibration method and device for mine rescue robot
By using multimodal sensor data fusion and feature point matching technology, the kinematic calibration problem of mine rescue robots in high dust and low light environments was solved, achieving accurate kinematic parameter estimation and real-time rescue capabilities.
Patent Information
- Application Number
- CN202511510960.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-12-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional kinematic calibration methods for mine rescue robots are not accurate enough in high dust and low light environments, cannot accurately obtain pose information, and lack the ability to perceive and adapt to the dynamic environment in real time, which hinders rescue operations.
By employing multimodal sensor data fusion technology, including LiDAR, dynamic vision sensors, millimeter-wave radar, and IMU, and through adaptive data fusion algorithms and feature point matching, combined with the SHOT descriptor and PointNet++ network, kinematic parameter estimation is optimized to achieve accurate motion calibration of the robot.
It improves the kinematic calibration accuracy and reliability of mine rescue robots in complex environments, meets real-time rescue needs, and reduces computational complexity.
Smart Images

Figure CN121048610A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of kinematic calibration, and in particular to a kinematic calibration method and apparatus for a mine rescue robot. Background Technology
[0002] Robot kinematic calibration aims to ensure the robot matches the theoretical design model, improving its point and trajectory accuracy, or enabling it to move entirely according to programmed actions. Calibration is necessary after robot assembly. Mine rescue robots are crucial equipment for coping with extreme environments such as mine disasters; the accuracy of their kinematic calibration directly determines their positioning, navigation, and operational capabilities in complex tunnels. The internal environment of mines typically presents challenges such as high concentrations of dust, low light levels, complex terrain, and potentially harmful gases, posing severe challenges to the robot's perception system.
[0003] In such complex mining environments, traditional robot kinematic calibration methods have many limitations. Common calibration methods based on a single sensor, such as relying solely on lidar, are susceptible to scattering and attenuation of the laser signal in high-dust environments, leading to a significant decrease in measurement accuracy. While relying solely on vision sensors for calibration suffers from severe image quality degradation under low-light conditions, greatly increasing the difficulty of feature extraction and matching, and consequently failing to accurately obtain the robot's pose information. Furthermore, most of these traditional methods do not fully consider the impact of environmental factors on sensor performance and robot kinematic parameters, lacking the ability to perceive and adaptively adjust to dynamic environments in real time. Inaccurate robot kinematic calibration will hinder rescue operations. Summary of the Invention
[0004] In view of the aforementioned problems, this application is proposed to provide a kinematic calibration method and apparatus for a mine rescue robot that overcomes or at least partially solves the aforementioned problems, comprising: A kinematic calibration method for a mine rescue robot, characterized by comprising the following steps: Acquire multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data from lidar, event streams from dynamic vision sensors, detection information from millimeter-wave radar, and acceleration and angular velocity data from IMU; The multimodal sensor data is weighted and fused using an adaptive data fusion algorithm to obtain multimodal sensor fused data. Based on multimodal sensor fusion data, core feature points are extracted by combining the SHOT descriptor and PointNet++ network. The core feature points are then searched for nearest neighbors to obtain matching nearest point pairs. The core feature points and the nearest points are then input into a preset point cloud model to obtain feature point matching pairs. The target pose parameters of the mine rescue robot are determined, and a residual function is constructed based on the target pose parameters and the feature point matching pair to obtain the motion parameters of the mine rescue robot.
[0005] Furthermore, the event stream processing steps of the dynamic vision sensor include: The raw event stream from the dynamic vision sensor is acquired, and a spatiotemporal filtering algorithm is used to eliminate noisy events in the raw event stream to construct event frames. Where x and y represent the pixel coordinates of the event, p is the event polarity, and t is the timestamp of the event. Based on the event frame, corner features and spatiotemporal gradient information are obtained using the DBSCAN algorithm; The event visual odometry method based on optical flow field estimation obtains the event flow of the dynamic visual sensor representing the robot's relative pose change based on the corner features and the spatiotemporal gradient information.
[0006] Furthermore, the processing steps for the detection information from the millimeter-wave radar include: The millimeter-wave radar signal is acquired, and the millimeter-wave radar signal is converted into sparse point cloud information using the range Doppler algorithm. The target point cloud information is then determined from the sparse point cloud information using the CFAR detection algorithm. The detection information of the millimeter-wave radar is obtained by using a point cloud enhancement algorithm based on the target point cloud information through multi-frame accumulation and Gaussian filtering. The event stream from the dynamic vision sensor and the detection information from the millimeter-wave radar are input into the radar vision feature fusion network for feature fusion.
[0007] Furthermore, the step of weightedly fusing the multimodal sensor data using an adaptive data fusion algorithm to obtain multimodal sensor fused data includes: Environmental feature vectors are extracted based on the multimodal sensor data. Where D represents dust concentration, Indicates relative humidity, Indicates light intensity, Ambient temperature; The environmental feature vector is input into the fuzzy inference model, and the output range of the fuzzy inference model is... Environmental adaptability factors ; Based on the aforementioned environmental adaptability factors The multimodal sensor data is obtained by weighted fusion of the data with the adaptive data fusion algorithm and the Q-learning reinforcement learning algorithm.
[0008] Furthermore, the steps prior to extracting core feature points based on multimodal sensor fusion data and combining the SHOT descriptor with the PointNet++ network include: By using voxel filtering, a voxel grid of equal size is set for the point cloud data in the multimodal sensor fusion data, and representative points in each voxel grid are extracted. Isolated points within the radius range of each representative point are identified and deleted by radius filtering to obtain denoised point cloud data.
[0009] Furthermore, the step of determining the target pose parameters of the mine rescue robot includes: Based on the rigid motion characteristics of the mine rescue robot, the pose of the mine rescue robot is represented as an element T of a special Euclidean group SE (3), and initial pose data is generated through the element T; wherein, the element T is a homogeneous transformation matrix; Based on the correspondence between Lie groups and Lie algebras, a left perturbation model is used to parameterize the initial pose data, and pose perturbation is selected. and will Parameterized as elements in the Lie algebra se(3) , It is a 3×4 matrix; Using exponential mapping exp( )right Perform a Lie group transformation to generate the perturbed target pose parameters. ;in This represents group multiplication.
[0010] Furthermore, the step of constructing a residual function based on the target pose parameters and the feature point matching pairs to obtain the motion parameters of the mine rescue robot also includes: Multiple residual functions are constructed based on the target pose parameters and the feature point matching pairs; The Huber kernel function is used to handle outliers in each of the residual functions. The motion parameters of the mine rescue robot are obtained by weighted summation of all the residual terms.
[0011] A kinematic calibration device for a mine rescue robot, the device comprising: The data acquisition module is used to acquire multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data of lidar, event stream of dynamic vision sensor, detection information of millimeter-wave radar, and acceleration and angular velocity data of IMU; The data fusion module is used to perform weighted fusion of the multimodal sensor data using an adaptive data fusion algorithm to obtain multimodal sensor fused data; The data matching module is used to extract core feature points based on multimodal sensor fusion data, combined with SHOT descriptors and PointNet++ network, perform nearest neighbor search on the core feature points to obtain matching neighbor point pairs, and input the core feature points and the neighbor points into a preset point cloud model to obtain feature point matching pairs. The parameter determination module is used to determine the target pose parameters of the mine rescue robot, and construct a residual function based on the target pose parameters and the feature point matching pair to obtain the motion parameters of the mine rescue robot.
[0012] A computer electronic device includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When executed by the processor, the computer program implements the steps of the kinematic calibration method for a mine rescue robot as described above.
[0013] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the kinematic calibration method for a mine rescue robot as described above.
[0014] This application has the following advantages: In the embodiments of this application, addressing the problem in the prior art that it does not fully consider the impact of environmental factors on sensor performance and robot kinematic parameters, and lacks the ability to perceive and adaptively adjust to dynamic environments in real time, this application provides a solution for achieving accurate motion calibration of a robot by fusing multimodal data and using robust feature matching and pose estimation. Specifically, the solution involves: acquiring multimodal sensor data with the same timestamp; wherein the multimodal sensor data includes point cloud data from LiDAR, event streams from dynamic vision sensors, detection information from millimeter-wave radar, and acceleration and angular velocity data from IMU; weighting and fusing the multimodal sensor data using an adaptive data fusion algorithm to obtain multimodal sensor fused data; extracting core feature points based on the multimodal sensor fused data, combining the SHOT descriptor and PointNet++ network, performing nearest neighbor search on the core feature points to obtain matching nearest neighbor point pairs, and inputting the core feature points and the nearest neighbor points into a preset point cloud model to obtain feature point matching pairs; determining the target pose parameters of the mine rescue robot, and constructing a residual function based on the target pose parameters and the feature point matching pairs to obtain the motion parameters of the mine rescue robot. Through feature point extraction and matching processes, and by preprocessing the point cloud, the amount of data and the impact of outliers are reduced. By combining the SHOT descriptor with the PointNet++ network, the accuracy of feature point matching is greatly improved, the kinematic parameter estimation is optimized, the accuracy of parameter estimation is improved, the computational complexity is reduced, and the real-time requirements of mine rescue are met. Attached Figure Description
[0015] To more clearly illustrate the technical solution of this application, the drawings used in the description of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating the steps of a kinematic calibration method for a mine rescue robot according to an embodiment of this application; Figure 2 This is a structural block diagram of a kinematic calibration device for a mine rescue robot according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a computer electronic device provided in an embodiment of the present invention; 1. Computer electronic device; 2. External device; 3. Processing unit; 4. Bus; 5. Network adapter; 6. I / O interface; 7. Display; 8. Memory; 9. Random access memory; 10. Cache memory; 11. Storage system; 12. Program / utility; 13. Program module. Detailed Implementation
[0017] To make the objectives, features, and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0018] The inventors, through analysis of existing technologies, discovered that multimodal heterogeneous sensor fusion technology is gradually becoming an important direction for improving the environmental perception capabilities of robots. By integrating the advantages of multiple sensors, it is expected to achieve more accurate perception in complex environments. Simultaneously, with the rapid development of machine learning and artificial intelligence technologies, data-driven adaptive algorithms offer new ideas for improving the accuracy of robot kinematic calibration, enabling real-time adjustments to calibration strategies based on environmental changes. By constructing a multimodal heterogeneous sensor fusion system, combining spatiotemporally synchronized adaptive data fusion, robust feature point extraction and matching, kinematic parameter estimation based on Lie group optimization, and dynamic environmental perception and error compensation, the shortcomings of traditional methods can be overcome, improving the accuracy and reliability of kinematic calibration for mine rescue robots in complex environments and ensuring the efficient operation of mine rescue work.
[0019] Reference Figure 1This application illustrates a kinematic calibration method for a mine rescue robot according to an embodiment of the present application: S110. Acquire multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data of lidar, event stream of dynamic vision sensor, detection information of millimeter-wave radar, and acceleration and angular velocity data of IMU. S120. The multimodal sensor data is weighted and fused using an adaptive data fusion algorithm to obtain multimodal sensor fused data; S130. Based on the multimodal sensor fusion data, the core feature points are extracted by combining the SHOT descriptor and the PointNet++ network. The core feature points are searched for nearest neighbors to obtain matching nearest point pairs. The core feature points and the nearest points are then input into a preset point cloud model to obtain feature point matching pairs. S140. Determine the target pose parameters of the mine rescue robot, and construct a residual function based on the target pose parameters and the feature point matching pair to obtain the motion parameters of the mine rescue robot.
[0020] In the embodiments of this application, the feature point extraction and matching process reduces the amount of data and the impact of outliers by preprocessing the point cloud; the combination of SHOT descriptor and PointNet++ network greatly improves the accuracy of feature point matching, optimizes kinematic parameter estimation, improves parameter estimation accuracy, reduces computational complexity, and meets the real-time requirements of mine rescue.
[0021] The kinematic calibration method for a mine rescue robot in this exemplary embodiment will be further described below.
[0022] As described in step S110, multimodal sensor data with the same timestamp are acquired. The multimodal sensor data includes point cloud data from lidar, event streams from dynamic vision sensors, detection information from millimeter-wave radar, and acceleration and angular velocity data from IMU.
[0023] In one embodiment of the present invention, the specific processing procedure of the "event flow of the dynamic vision sensor" described in step S110 can be further explained in conjunction with the following description.
[0024] As described in the following steps, the raw event stream from the dynamic vision sensor is acquired, and a spatiotemporal filtering algorithm is used to eliminate noise events in the raw event stream to construct event frames. Where x and y represent the pixel coordinates where the event occurred, p is the event polarity (usually +1 indicates an increase in brightness, and -1 indicates a decrease in brightness), and t is the timestamp of the event.
[0025] It should be noted that dynamic vision sensors inevitably generate noise events during operation, which can interfere with the extraction of valid information. Eliminating noise events through spatiotemporal filtering algorithms can improve data purity, ensuring that subsequent processing is based on accurate event information. Constructing event frames involves structuring the raw event data, making it easier for subsequent algorithms to process and analyze.
[0026] As described in the following steps, corner features and spatiotemporal gradient information are obtained from the event frame using the DBSCAN algorithm.
[0027] It should be noted that in complex mine environments, directly processing all event data is inefficient and makes it difficult to extract key information. Utilizing the DBSCAN algorithm to detect corner features allows for the selection of representative key points from a large amount of event data. These key points contain important structural information about the environment. Calculating spatiotemporal gradients can further describe the changing characteristics of these key points, providing richer feature information for subsequent motion estimation.
[0028] As described in the following steps, the event visual odometry method based on optical flow field estimation obtains the event flow of the dynamic visual sensor that characterizes the robot's relative pose change based on the corner features and spatiotemporal gradient information.
[0029] It should be noted that the event visual odometry calculation method for optical flow field estimation in the above steps also needs to be combined with the concept of event time surface in order to achieve visual modeling in the spatiotemporal domain, solve for motion vectors of multiple corner features, adjust the estimated value of motion vectors in continuous iterative optimization, and finally achieve motion estimation with higher pixel-level accuracy to obtain event flow data of the robot's relative pose changes at continuous moments.
[0030] In one embodiment of the present invention, the specific processing procedure of the "detection information of millimeter-wave radar" described in step S110 can be further explained in conjunction with the following description.
[0031] As described in the following steps, a millimeter-wave radar signal is acquired, the millimeter-wave radar signal is converted into sparse point cloud information using a range Doppler algorithm, and the target point cloud information is determined from the sparse point cloud information using a CFAR detection algorithm.
[0032] It should be noted that millimeter-wave radar performs stably in the high-dust, low-light environment of mines, but the resulting raw point cloud is relatively sparse and susceptible to noise interference. The range-Doppler algorithm can extract range and velocity information from the raw point cloud, forming effective sparse point cloud information. The CFAR (Constant False Alarm Rate) detection algorithm, through adaptive thresholding, can accurately identify target points in complex noisy backgrounds, avoiding noise being misidentified as targets.
[0033] As described in the following steps, the detection information of the millimeter-wave radar is obtained based on the target point cloud information through a point cloud enhancement algorithm of multi-frame accumulation and Gaussian filtering.
[0034] It should be noted that the multi-frame accumulation algorithm combines and fuses point cloud data from multiple consecutive frames to supplement missing target details in a single frame and increase the density of the point cloud; Gaussian filtering smooths the enhanced point cloud and reduces noise interference.
[0035] Furthermore, the Kalman filter algorithm can be used to predict the current position and velocity based on the target's historical motion state, and then corrected by combining new observation data to achieve continuous tracking of the target's trajectory.
[0036] As described in the following steps, the event stream of the dynamic vision sensor and the detection information of the millimeter-wave radar are input into the radar vision feature fusion network for feature fusion.
[0037] It should be noted that the target distance, velocity and other features extracted by millimeter-wave radar are associated and matched with the image features acquired by the visual sensor. The attention mechanism will evaluate the reliability of the two modal features based on the current environment (such as dust concentration and lighting conditions), assign higher weights to the more reliable features, and finally output the fused comprehensive features.
[0038] As described in step S120, the multimodal sensor data is weighted and fused using an adaptive data fusion algorithm to obtain multimodal sensor fused data.
[0039] It should be noted that the adaptive data fusion algorithm is as follows: , Among them, the Let be the weight of the i-th multimodal sensor data at time t during the fusion process; Environmental adaptability factors; Let be the measurement variance of the i-th sensor at time t; n represents the total number of sensors participating in data fusion.
[0040] In one embodiment of the present invention, the specific steps of step S120, "weighting and fusing the multimodal sensor data using an adaptive data fusion algorithm to obtain multimodal sensor fused data," can be further explained in conjunction with the following description.
[0041] As described in the following steps, environmental feature vectors are extracted based on the multimodal sensor data. Where D represents dust concentration, Indicates relative humidity, Indicates light intensity, Ambient temperature.
[0042] It should be noted that by extracting environmental feature vectors, the impact of the environment on the sensor can be quantified, providing a basis for dynamically adjusting the sensor weights based on the environment. For example, when the dust concentration is high, the lidar signal is easily scattered. This change can be detected by extracting environmental feature vectors, and the lidar data weight can be reduced in subsequent steps.
[0043] As described in the following steps, the environmental feature vector is input into the fuzzy inference model, and the fuzzy inference model outputs a range of... Environmental adaptability factors .
[0044] It should be noted that the fuzzy inference model contains nearly 50 rules generated by combining expert experience and data training. It further combines the Mamdani inference rules to output environmental adaptation factors, thereby achieving adaptive adjustment of environmental adaptation factors.
[0045] As described in the following steps, based on the environmental adaptability factor The multimodal sensor data is obtained by weighted fusion of the data with the adaptive data fusion algorithm and the Q-learning reinforcement learning algorithm.
[0046] It should be noted that the environment in mines is diverse, and the fuzzy rules in the fuzzy inference model may not be fully adaptable to all scenarios. The reinforcement learning in the Q-learning algorithm, by continuously trying different rule adjustment strategies, enables the fuzzy inference model to evolve in various complex environments, continuously improving the accuracy and robustness of the adaptive weighted fusion algorithm, and ensuring that the multi-sensor data fusion results always maintain high precision.
[0047] As described in step S130, based on the multimodal sensor fusion data, core feature points are extracted using the SHOT descriptor and PointNet++ network. Nearest neighbor search is performed on these core feature points to obtain matching neighbor point pairs. The core feature points and the neighbor points are then input into a point cloud model to obtain the feature point matching pairs. .
[0048] In one embodiment of the present invention, the specific processing steps prior to step S130, "extracting core feature points based on multimodal sensor fusion data and combining the SHOT descriptor with the PointNet++ network," can be further explained in conjunction with the following description.
[0049] As described in the following steps, a voxel grid of equal size is set for the point cloud data in the multimodal sensor fusion data by voxel filtering, and representative points in each voxel grid are extracted.
[0050] It should be noted that the purpose of this step is to reduce the data density of point cloud data in multimodal sensor fusion data. Specifically, a three-dimensional voxel grid is set for the point cloud data, and all point cloud data is traversed, mapping each point to the corresponding voxel grid. For each grid, the point closest to the grid center (or the average point within the grid) is retained as the "representative point" of that grid, and other redundant points within the grid are deleted. This operation reduces the point cloud data density and reduces the amount of subsequent computation.
[0051] As described in the following steps, isolated points within the radius range of each representative point are determined by radius filtering and deleted to obtain denoised point cloud data.
[0052] It should be noted that, in the point cloud after voxel filtering, radius filtering sets a spherical neighborhood with a fixed radius for each point, counts the number of other points contained in each spherical neighborhood (i.e., "point count"), and determines points with fewer than a preset threshold as "isolated points" and deletes them, finally obtaining denoised and simplified point cloud data.
[0053] Voxel filtering and radius filtering are used to reduce the density and noise of point cloud data from multimodal sensor fusion data, providing more concise point cloud data for subsequent extraction of core feature points by the SHOT descriptor and PointNet++ network.
[0054] As described in step S140, the target pose parameters of the mine rescue robot are determined, and the target pose parameters and the feature points are matched accordingly. Construct a residual function to obtain the motion parameters of the mine rescue robot; In one embodiment of the present invention, the specific steps of "determining the target pose parameters of the mine rescue robot" in step S140 can be further explained in conjunction with the following description.
[0055] As described in the following steps, based on the rigid motion characteristics of the mine rescue robot, the pose of the mine rescue robot is represented as an element T of a special Euclidean group SE (3), and initial pose data is generated through the element T; wherein, the element T is a homogeneous transformation matrix.
[0056] As described in the following steps, based on the correspondence between Lie groups and Lie algebras, a left perturbation model is used to parameterize the initial pose data, and the pose perturbation is selected. and will Parameterized as elements in the Lie algebra se(3) , It is a 3×4 matrix.
[0057] As described in the following steps, the exponential mapping exp( )right Perform a Lie group transformation to generate the perturbed target pose parameters. .in This represents group multiplication.
[0058] In one embodiment of the present invention, step S140, "matching the target pose parameters and the feature points," can be further explained in conjunction with the following description. The specific steps for "constructing the residual function to obtain the motion parameters of the mine rescue robot" are as follows.
[0059] As described in the following steps, multiple residual functions are constructed based on the target pose parameters and the feature point matching pairs; outliers in each residual function are processed using the Huber kernel function; and all residual terms are weighted and summed to obtain the motion parameters of the mine rescue robot. The residual function is:
[0060] in, For the current feature point, For the feature points corresponding to the reference frame, For information matrix, For residual terms; The formula for weighted summation of all the residual terms is as follows:
[0061] in, This represents the Huber kernel function.
[0062] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0063] Reference Figure 2 This application illustrates a kinematic calibration device for a mine rescue robot according to an embodiment of the present application, which specifically includes the following modules: Specifically, it includes: The data acquisition module 210 is used to acquire multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data of lidar, event stream of dynamic vision sensor, detection information of millimeter-wave radar and acceleration and angular velocity data of IMU; The data fusion module 220 is used to perform weighted fusion of the multimodal sensor data through an adaptive data fusion algorithm to obtain multimodal sensor fused data; The data matching module 230 is used to extract core feature points based on multimodal sensor fusion data, combined with SHOT descriptor and PointNet++ network, perform nearest neighbor search on the core feature points to obtain matching neighbor point pairs, and input the core feature points and the neighbor points into a preset point cloud model to obtain feature point matching pairs. The parameter determination module 240 is used to determine the target pose parameters of the mine rescue robot, and construct a residual function based on the target pose parameters and the feature point matching pair to obtain the motion parameters of the mine rescue robot.
[0064] In one embodiment of the present invention, the data acquisition module 210 includes: The event frame construction submodule is used to acquire the raw event stream from the dynamic vision sensor, and to eliminate noise events in the raw event stream using a spatiotemporal filtering algorithm to construct event frames. Where x and y represent the pixel coordinates of the event, p is the event polarity, and t is the timestamp of the event. The corner feature and spatiotemporal gradient information calculation submodule is used to obtain corner features and spatiotemporal gradient information based on the event frame using the DBSCAN algorithm. The event flow calculation submodule is used for event visual odometry calculation based on optical flow field estimation. Based on the corner features and the spatiotemporal gradient information, it obtains the event flow of the dynamic visual sensor representing the robot's relative pose change.
[0065] In one embodiment of the present invention, the data acquisition module 210 includes: The target point cloud information calculation submodule is used to acquire millimeter-wave radar signals, convert the millimeter-wave radar signals into sparse point cloud information through the range Doppler algorithm, and use the CFAR detection algorithm to determine the target point cloud information in the sparse point cloud information. The detection information calculation submodule is used to obtain the detection information of the millimeter-wave radar based on the target point cloud information through a point cloud enhancement algorithm of multi-frame accumulation and Gaussian filtering; The feature fusion submodule is used to input the event stream of the dynamic vision sensor and the detection information of the millimeter-wave radar into the radar vision feature fusion network for feature fusion.
[0066] In one embodiment of the present invention, the data fusion module 220 includes: The environmental feature extraction submodule is used to extract environmental feature vectors based on the multimodal sensor data. Where D represents dust concentration, Indicates relative humidity, Indicates light intensity, Ambient temperature; The adaptive factor calculation submodule is used to input the environmental feature vector into the fuzzy inference model, and the fuzzy inference model outputs a range of... Environmental adaptability factors ; The fusion data calculation submodule is used to calculate based on the environmental adaptability factors. The multimodal sensor data is obtained by weighted fusion of the data with the adaptive data fusion algorithm and the Q-learning reinforcement learning algorithm.
[0067] In one embodiment of the present invention, the data matching module 230 includes: The representative point extraction submodule is used to set an equal-sized voxel grid for the point cloud data in the multimodal sensor fusion data through voxel filtering, and extract representative points in each voxel grid. The point cloud data denoising submodule is used to determine and delete isolated points within the radius range of each representative point through radius filtering, thereby obtaining denoised point cloud data.
[0068] In one embodiment of the present invention, the parameter determination module 240 includes: The pose initial data calculation submodule is used to represent the pose of the mine rescue robot as an element T of a special Euclidean group SE (3) based on the rigid motion characteristics of the mine rescue robot, and generate pose initial data through the element T; wherein, the element T is a homogeneous transformation matrix; The perturbation parameterization submodule is used to parameterize the initial pose data using a left perturbation model based on the correspondence between Lie groups and Lie algebras, selecting pose perturbations. and will Parameterized as elements in the Lie algebra se(3) , It is a 3×4 matrix; The Lie group transformation submodule is used to utilize the exponential mapping exp( )right Perform a Lie group transformation to generate the perturbed target pose parameters. ;in This represents group multiplication.
[0069] In one embodiment of the present invention, the parameter determination module 240 includes: The residual function construction submodule is used to construct multiple residual functions based on the target pose parameters and the feature point matching pairs; The outlier handling submodule is used to process outliers in each of the residual functions using the Huber kernel function; The weighted summation submodule is used to perform a weighted summation of all the residual terms to obtain the motion parameters of the mine rescue robot.
[0070] Reference Figure 3 The diagram illustrates a computer electronic device for implementing a method for kinematic calibration of a mine rescue robot according to the present invention, which may specifically include the following: The aforementioned computer electronic device 1 is manifested in the form of a general-purpose computing device. The components of the computer electronic device 1 may include, but are not limited to: one or more processors or processing units 3, memory 8, and a bus 4 connecting different system components (including memory 8 and processing unit 3).
[0071] Bus 4 represents one or more of several bus architectures, including memory buses or memory controllers, peripheral buses, graphics acceleration ports, processors, or local buses using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Audio / Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0072] Computer electronic device 1 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer electronic device 1, including volatile and non-volatile media, removable and non-removable media.
[0073] Memory 8 may include computer system readable media in the form of volatile memory, such as random access memory 9 and / or cache memory 10. Computer electronic device 1 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 11 may be used to read and write non-removable, non-volatile magnetic media (commonly referred to as a "hard disk drive"). Although Figure 3 As not shown, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 4 via one or more data media interfaces. The memory may include at least one program product having a set (e.g., at least one) of program modules 13 configured to perform the functions of the embodiments of this application.
[0074] A program / utility 12 having a set (at least one) of program modules 13 may be stored, for example, in memory. Such program modules 13 include—but are not limited to—an operating system, one or more application programs, other program modules 13, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 13 typically perform the functions and / or methods described in the embodiments of this application.
[0075] The computer electronic device 1 can also communicate with one or more external devices 2 (e.g., keyboard, pointing device, display 7, camera, etc.), and with one or more devices that enable an operator to interact with the computer electronic device 1, and / or with any device that enables the computer electronic device 1 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed through the I / O interface 6. Furthermore, the computer electronic device 1 can also communicate with one or more networks (e.g., local area network (LAN)), wide area network (WAN), and / or public networks (e.g., the Internet) through the network adapter 5. Figure 3 As shown, network adapter 5 communicates with other modules of computer electronic device 1 via bus 4. It should be understood that, although... Figure 3 Not shown, it may be combined with other hardware and / or software modules, including but not limited to: microcode, device drivers, redundant processing unit 3, external disk drive array, RAID system, tape drive and data backup storage system 11, etc.
[0076] The processing unit 3 executes various functional applications and data processing by running programs stored in memory 8, such as implementing a method for kinematic calibration of a mine rescue robot provided in the embodiments of this application.
[0077] That is, when the processing unit 3 executes the above program, it achieves the following: acquiring multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data of lidar, event stream of dynamic vision sensor, detection information of millimeter-wave radar, and acceleration and angular velocity data of IMU; weighting and fusing the multimodal sensor data through an adaptive data fusion algorithm to obtain multimodal sensor fused data; based on the multimodal sensor fused data, extracting core feature points by combining SHOT descriptor and PointNet++ network, performing nearest neighbor search on the core feature points to obtain matching neighbor point pairs, and inputting the core feature points and the neighbor points into a preset point cloud model to obtain feature point matching pairs; determining the target pose parameters of the mine rescue robot, and constructing a residual function based on the target pose parameters and the feature point matching pairs to obtain the motion parameters of the mine rescue robot.
[0078] In this application embodiment, the application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for kinematic calibration of a mine rescue robot as provided in all embodiments of the application.
[0079] That is, when the program is executed by the processor, it performs the following: acquiring multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data of lidar, event stream of dynamic vision sensor, detection information of millimeter-wave radar, and acceleration and angular velocity data of IMU; weighting and fusing the multimodal sensor data through an adaptive data fusion algorithm to obtain multimodal sensor fused data; based on the multimodal sensor fused data, extracting core feature points by combining SHOT descriptor and PointNet++ network, performing nearest neighbor search on the core feature points to obtain matching neighbor point pairs, and inputting the core feature points and the neighbor points into a preset point cloud model to obtain feature point matching pairs; determining the target pose parameters of the mine rescue robot, and constructing a residual function based on the target pose parameters and the feature point matching pairs to obtain the motion parameters of the mine rescue robot.
[0080] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0081] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0082] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the operator's computer, partially on the operator's computer, as a standalone software package, partially on the operator's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the operator's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider). The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably.
[0083] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0084] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0085] The above provides a detailed description of the kinematic calibration method and device for a mine rescue robot provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A kinematic calibration method for a mine rescue robot, characterized in that, Including the following steps: Acquire multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data from lidar, event streams from dynamic vision sensors, detection information from millimeter-wave radar, and acceleration and angular velocity data from IMU; The multimodal sensor data is weighted and fused using an adaptive data fusion algorithm to obtain multimodal sensor fused data. Based on multimodal sensor fusion data, core feature points are extracted by combining SHOT descriptors and PointNet++ network. The core feature points are then searched for nearest neighbors to obtain matching nearest point pairs. The core feature points and the nearest points are then input into a preset point cloud model to obtain feature point matching pairs. The target pose parameters of the mine rescue robot are determined, and a residual function is constructed based on the target pose parameters and the feature point matching pair to obtain the motion parameters of the mine rescue robot.
2. The method according to claim 1, characterized in that, The event stream processing steps of the dynamic vision sensor include: The raw event stream from the dynamic vision sensor is acquired, and a spatiotemporal filtering algorithm is used to eliminate noisy events in the raw event stream to construct event frames. Where x and y represent the pixel coordinates of the event, p is the event polarity, and t is the timestamp of the event. Based on the event frame, corner features and spatiotemporal gradient information are obtained using the DBSCAN algorithm; The event visual odometry method based on optical flow field estimation obtains the event flow of the dynamic visual sensor representing the robot's relative pose change based on the corner features and the spatiotemporal gradient information.
3. The method according to claim 1, characterized in that, The processing steps for the detection information from the millimeter-wave radar include: The millimeter-wave radar signal is acquired, and the millimeter-wave radar signal is converted into sparse point cloud information using the range Doppler algorithm. The target point cloud information is then determined from the sparse point cloud information using the CFAR detection algorithm. The detection information of the millimeter-wave radar is obtained by using a point cloud enhancement algorithm based on the target point cloud information through multi-frame accumulation and Gaussian filtering. The event stream from the dynamic vision sensor and the detection information from the millimeter-wave radar are input into the radar vision feature fusion network for feature fusion.
4. The method according to claim 1, characterized in that, The step of weighted fusing the multimodal sensor data using an adaptive data fusion algorithm to obtain multimodal sensor fused data includes: Environmental feature vectors are extracted based on the multimodal sensor data. Where D represents dust concentration, Indicates relative humidity, Indicates light intensity, Ambient temperature; The environmental feature vector is input into the fuzzy inference model, and the output range of the fuzzy inference model is... Environmental adaptability factors ; Based on the aforementioned environmental adaptability factors The multimodal sensor data is obtained by weighted fusion of the data with the adaptive data fusion algorithm and the Q-learning reinforcement learning algorithm.
5. The method according to claim 1, characterized in that, The steps preceding the extraction of core feature points based on multimodal sensor fusion data, combined with the SHOT descriptor and PointNet++ network, include: By using voxel filtering, a voxel grid of equal size is set for the point cloud data in the multimodal sensor fusion data, and representative points in each voxel grid are extracted. Isolated points within the radius range of each representative point are identified and deleted by radius filtering to obtain denoised point cloud data.
6. The method according to claim 1, characterized in that, The step of determining the target pose parameters of the mine rescue robot includes: Based on the rigid motion characteristics of the mine rescue robot, the pose of the mine rescue robot is represented as an element T of a special Euclidean group SE (3), and the initial pose data is generated through the element T; wherein, the element T is a homogeneous transformation matrix; Based on the correspondence between Lie groups and Lie algebras, a left perturbation model is used to parameterize the initial pose data, and pose perturbation is selected. and will Parameterized as elements in the Lie algebra se(3) , It is a 3×4 matrix; Using exponential mapping exp( )right Perform a Lie group transformation to generate the perturbed target pose parameters. ;in This represents group multiplication.
7. The method according to claim 6, characterized in that, The step of constructing a residual function based on the target pose parameters and the feature point matching pair to obtain the motion parameters of the mine rescue robot includes: Multiple residual functions are constructed based on the target pose parameters and the feature point matching pairs; The Huber kernel function is used to handle outliers in each of the residual functions. The motion parameters of the mine rescue robot are obtained by weighted summation of all the residual terms.
8. A kinematic calibration device for a mine rescue robot, characterized in that, The device includes: The data acquisition module is used to acquire multimodal sensor data with the same timestamp; wherein, the multimodal sensor data includes point cloud data of lidar, event stream of dynamic vision sensor, detection information of millimeter-wave radar, and acceleration and angular velocity data of IMU; The data fusion module is used to perform weighted fusion of the multimodal sensor data using an adaptive data fusion algorithm to obtain multimodal sensor fused data; The data matching module is used to extract core feature points based on multimodal sensor fusion data, combined with SHOT descriptors and PointNet++ network, perform nearest neighbor search on the core feature points to obtain matching neighbor point pairs, and input the core feature points and the neighbor points into a preset point cloud model to obtain feature point matching pairs. The parameter determination module is used to determine the target pose parameters of the mine rescue robot, and construct a residual function based on the target pose parameters and the feature point matching pair to obtain the motion parameters of the mine rescue robot.
9. A computer electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the kinematic calibration method for a mine rescue robot as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the kinematic calibration method for a mine rescue robot as described in any one of claims 1 to 7.