Robot multi-source data fusion sensing method, chip and electronic equipment
By constructing an adaptive fusion architecture and a multimodal sensor array, the robot achieves accurate synchronization and intelligent fusion of data from multiple types of sensors, solving the problem of perception deviation in complex dynamic environments, improving perception accuracy and environmental adaptability, and meeting the application requirements of high-precision and high-dynamic scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 杭州智芯科微电子科技有限公司
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-12
AI Technical Summary
Existing robot perception systems mostly use a single sensor or simple data overlay, which can easily lead to perception bias in complex and dynamic environments, resulting in problems such as motion lag and decision-making errors, which seriously restricts the application of robots in high-precision and high-dynamic scenarios.
By constructing an adaptive fusion architecture, accurate synchronization, quality assessment, and intelligent fusion of multi-type sensor data are achieved, outputting high-precision environmental data. A multi-modal sensor array is used for data preprocessing, a multi-dimensional data quality assessment model is constructed, confidence quantification and weight calibration are performed, and the optimal state data is output by combining local feature fusion and global state fusion.
It significantly improves the robot's perception accuracy and environmental adaptability, enhances real-time performance and robustness, enables stable output in complex scenarios, reduces development costs, and improves the robot's application capabilities in high-precision and high-dynamic scenarios.
Smart Images

Figure CN122008270A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot perception technology, specifically to a robot multi-source data fusion perception method, chip, and electronic device. Background Technology
[0002] With the rapid development of robotics technology, robots have been widely used in various fields such as industrial manufacturing, smart homes, outdoor exploration, and service delivery. During task execution, robots need to acquire environmental information through various sensors to achieve functions such as autonomous navigation, target recognition, and obstacle avoidance. The accuracy and real-time nature of environmental perception directly determine the robot's maneuverability and environmental adaptability.
[0003] Existing robot perception systems mostly acquire environmental data using single sensors or simple data overlay methods, which have significant limitations: visual sensors are easily affected by factors such as changes in lighting and occlusion, leading to distortion in target feature extraction; inertial measurement units (IMUs) suffer from cumulative errors, and attitude drift occurs when operating alone for extended periods; while lidar offers high ranging accuracy, feature point matching is difficult in complex textured environments, and data redundancy is high; force sensors can only provide feedback on contact status and cannot acquire global environmental information. The limitations of single-sensor data make robots prone to perceptual biases in complex dynamic environments, leading to problems such as motion stuttering and decision-making errors, severely restricting the application of robots in high-precision, high-dynamic scenarios. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a robot multi-source data fusion perception method. By constructing an adaptive fusion architecture, it achieves accurate synchronization, quality assessment, and intelligent fusion of data from multiple types of sensors, outputting high-precision environmental data to provide reliable support for robot motion planning and significantly improve the robot's flexibility and environmental adaptability. The robot multi-source data fusion perception method includes the following steps: S1. Collect raw data from the surrounding environment using a multimodal sensor array, and then preprocess the raw data to obtain standardized environmental data.
[0005] S2. Confidence level is obtained by quantifying the confidence of standardized environmental data through a constructed multi-dimensional data quality assessment model.
[0006] S3. Determine each confidence level G. i Is it greater than the corresponding standard confidence threshold g? i If yes, then the standardized environmental data is calibrated as valid data and its weight coefficient is marked; otherwise, the standardized environmental data is determined to be low-quality data, its weight is reduced (provisionally designated as valid data) or the channel data is temporarily blocked to avoid error amplification.
[0007] S4. Perform multi-source data fusion on the effective data to obtain the optimal state data.
[0008] S5. Perform consistency verification between the optimal state data output by fusion and historical data. Compare multiple consecutive frames of data using a sliding window mechanism to calculate the comprehensive deviation value.
[0009] S6. Determine whether the overall deviation value is within the deviation threshold range. If yes, confirm the result is valid and output it to the robot motion controller; otherwise, trigger the feedback adjustment mechanism.
[0010] Preferably, the sensor array includes a vision sensor, an inertial measurement unit, a lidar, and / or a force sensor.
[0011] Preferred methods include: raw data preprocessing, including time synchronization, spatial alignment, noise filtering, and data standardization.
[0012] Preferred method: Time synchronization adopts a hardware interrupt triggering mechanism to generate a unified microsecond-level timestamp, calibrate the sampling time of the multimodal sensor array to synchronize it, and eliminate timing deviations.
[0013] Preferred method: Spatial alignment includes: constructing a robot base coordinate system and obtaining the coordinate system where the original data is located, and then, based on the preset extrinsic parameter matrix of the multimodal sensor array, translating the coordinate systems of different sensors to unify them into the robot base coordinate system to achieve spatial position matching.
[0014] Preferred methods include: noise filtering using corresponding algorithms for different sensor characteristics; Gaussian filtering to remove noise from visual data; moving average filtering to suppress vibration interference from IMU data; and pass-through filtering to remove invalid point clouds at long distances from LiDAR data; and data standardization to map various types of data to a unified numerical range, providing consistent input for subsequent fusion calculations.
[0015] Preferred: The evaluation dimensions of the multi-dimensional data quality assessment model include: data integrity, signal-to-noise ratio, feature identification, and temporal stability.
[0016] The preferred multi-dimensional data quality assessment model process is as follows: First, weights are assigned to each dimension based on sensor characteristics and application scenarios. Then, the acquisition coefficient G for each dimension is calculated. i j , where i is the standardized environmental data number and j is the dimension number.
[0017] Preferred: Data integrity data , where n i N is the effective data volume of the standardized environmental data collection numbered i. i It is the theoretical sampled data volume of the standardized environmental data numbered i.
[0018] Preferred: Signal-to-noise ratio data integrity data , where n SNR-i It is the actual SNR of the standardized environmental data collection numbered i, N SNR-i It is the target SNR of the standardized environmental data numbered i.
[0019] Preferred: Feature recognition accuracy is calculated by combining sensor type through the proportion of feature points or deviation from the baseline, for visual sensors and LiDAR, feature recognition accuracy. , where n i0 ' is the number of valid feature points collected in the standardized environmental data collection, numbered i, and N is the number of valid feature points collected. i0 'This represents the total number of feature points extracted from the standardized environmental data numbered i. MU and force sensor: feature recognition rate.' In this formula, n i1 It is the actual data value collected in the standardized environmental data collection, numbered i, and n. i1 ' is the baseline data value for the standardized environmental data numbered i, N i1 'This is the maximum allowable deviation of the standardized environmental data numbered i.
[0020] Preferred: Timing stability , where s i S is the unit-time data fluctuation variance of the standardized environmental data numbered i. i It is the maximum allowable variance of the standardized environmental data numbered i.
[0021] Preferred: Final confidence level , where w j G is the weight value of dimension j, where J is the total number of dimensions, j=1, 2, ..., J; i j It is the collection coefficient of the j-th dimension of the standardized environmental data with number i.
[0022] Preferably: the confidence level value , where w j G is the weight value of dimension j, where J is the total number of dimensions, j=1, 2, ..., J; i j G is the collection coefficient of the j-th dimension of the standardized environmental data, numbered i. i j ' is the standard data of the j-th dimension of the standardized environmental data with number i.
[0023] Preferred: Standard confidence threshold g i The system adopts a "fixed baseline threshold + dynamic compensation" mechanism, which specifically includes: first, presetting a baseline threshold g according to the work scenario. iThe complexity coefficient is calculated based on environmental features. High-confidence fusion results with a confidence level ≥ 0.8 from the last 100 frames are selected and denoted as the reference frame set S. The mean value of each environmental state dimension in S is calculated as the calibration baseline value Xr. The fusion results from the last 100 frames are denoted as Xm, where m is the data frame number. The relative deviation of each frame from the calibration baseline value and the fusion deviation rate δt are calculated. If δt > 3%, the fusion result deviation is too large, and a preset baseline threshold g is set. i 'Raise the threshold by 0.05; if 1%≤δt≤3%: the fusion accuracy meets the standard, maintain the current threshold unchanged; if δt<1%: preset the baseline threshold g.' i Lower the threshold by 0.05; the calibrated standard confidence threshold should be limited to the interval [0.4, 0.9] as the standard confidence threshold g. i .
[0024] Preferred method for obtaining complexity coefficient includes: selecting standardized environmental data from the first 3 frames, focusing on the core dimensions of environmental representation, calculating the variance of the first 3 frames of data for each core dimension; weighted summation of the variances of all core dimensions to obtain the comprehensive fluctuation variance of the first 3 frames of data, and mapping the comprehensive fluctuation variance to the interval [-0.2, 0.3] to obtain the complexity coefficient.
[0025] Preferably, the optimal state data is obtained by using a method of local feature fusion and global state fusion.
[0026] Preferred: Local feature fusion layer: (1) Visual-LiDAR feature fusion; (2) IMU-Odometry feature fusion; (3) Force feature extraction.
[0027] Preferred method: Visual-LiDAR feature fusion includes: the visual sensor performs lightweight convolution and pooling operations on preprocessed image data to extract target contour and texture features, outputting a visual feature map with dimensions [H×W×C], where H is the image height, W is the width, and C is the number of feature channels; the LiDAR data, based on the spatial coordinates and density information of the point cloud, filters out the spatial feature points of obstacles, generates a three-dimensional point cloud feature set, and then maps it to the visual feature map through projection transformation. Figure 1 The two-dimensional plane is used to obtain the LiDAR feature map; finally, the two types of feature maps are fused by feature stitching and attention mechanism, and a dense environment feature map is generated after weighted fusion.
[0028] The preferred method involves IMU-odometry feature fusion, which includes: the IMU acquiring angular velocity and linear acceleration data of the robot, and obtaining a preliminary estimate of the robot's attitude angle through integration; and the odometry calculating the robot's displacement to obtain motion trajectory data. Prediction-update iterations are performed using state equations and observation equations to correct for IMU cumulative errors and odometry slippage errors, outputting preliminary estimates of the robot's attitude and motion trajectory.
[0029] Preferred: The state equation is Where k is the discretization time node number, X k-1 It is the state vector at time point k-1, containing attitude angle, displacement, and velocity; X k is the state vector at time point k, i.e., the state vector at the current time point; A is the state transition matrix; B is the control matrix; u is the input value of the IMU at time point k; w is the process noise at time point k.
[0030] Preferred: Observation equation is Z k The current time point represents the odometry observation; H represents the observation matrix; v k The observation noise is at time point k.
[0031] Preferred: Force feature extraction: targeting the contact pressure (F) collected by the force sensor x F γ The data of force direction (Fz) and force direction are filtered by sliding window to remove instantaneous interference, and then the data are mapped to the range [0, 1] by feature normalization.
[0032] Preferred: Global state fusion includes: integrating the results of each local fusion and defining a state vector X; calculating the contribution and the optimal state data X from the previous frame based on the confidence level. k-1 Predict the current frame state using the state transition matrix A. and the predicted covariance matrix The results of each local fusion are used as the observation value Z. k The observation matrix H is determined by combining the contribution, and the intermediate coefficients are calculated. The final update yields the optimal state data. .
[0033] Preferred: State vector X = [x, y, z, α, β, γ, O1, O2, ..., O n [E], where (x, y, z) are the robot's three-dimensional position coordinates, (α, β, γ) are the attitude angles, and O1~O n E represents the position and size information of each obstacle, and E represents the contact state characteristics.
[0034] Preferred: Current frame state A is the state transition matrix.
[0035] Preferred: Predicted covariance matrix Where A is the state transition matrix, P k-1 A is the optimal posterior covariance matrix of the (k-1)th frame. T Let Q be the transpose of the state transition matrix, and let Q be the process noise covariance matrix.
[0036] Preferred: Intermediate coefficient Where R is the observation noise covariance matrix, reflecting the measurement accuracy of the sensor, and H... T The contribution determines the transpose of the observation matrix.
[0037] Preferred: Optimal state data .
[0038] Preferred method for obtaining contribution includes obtaining a paired subset based on the current scenario and valid data, and then calculating the contribution based on the paired subset. Where l is the paired subset number, f is the valid data number, sl is the valid data number in the paired subset with number l, and G f γ is the confidence level corresponding to the valid data with ID f, F is the total number of valid data in the paired subset with ID l, f = 1, 2, ..., F; L is the total number of paired subsets, l = 1, 2, ..., L; l The scene adaptation coefficient corresponding to the pairing subset numbered l. It is the average confidence level corresponding to the paired subsets.
[0039] Preferred method: The calculation method for the comprehensive deviation value can be , where X k,f It is the fused data of the f-th dimension of the current frame, X w,k,f It is the arithmetic mean of the f-th dimension data of all frames within the sliding window; It is the historical data mean of the f-th dimension within the current frame sliding window; ε is the minimum value to prevent... When the denominator is 0, the dimensional weights w are then assigned. f 'Calculate the overall deviation value' .
[0040] The present invention also proposes a robot multi-source data fusion sensing chip, which integrates the above-mentioned multi-source data fusion sensing method and is used to execute the above-mentioned robot multi-source data fusion sensing method.
[0041] This embodiment proposes an electronic device, which includes the aforementioned multi-source data fusion sensing chip, power module, storage module, and sensor array. The sensor array is electrically connected to the sensor interface module of the chip. The storage module is used to store fusion algorithm parameters, historical data, and configuration files. The power module provides stable power to the chip and sensor array.
[0042] The technical effects and advantages of this invention are as follows: 1. Significantly improved perception accuracy: By constructing a redundant perception system through a multimodal sensor array and combining it with a hierarchical adaptive fusion architecture, the limitations of a single sensor are effectively compensated for. After fusion, the average error of environmental data is greatly reduced, and the accuracy of robot posture estimation is significantly improved. It can accurately identify obstacles, target objects and environmental changes in complex scenes.
[0043] 2. Enhanced real-time performance and robustness: The chip adopts a heterogeneous computing architecture, and hardware acceleration is achieved in core processes such as preprocessing and fusion computing. The total processing time of a single frame of data is greatly reduced, meeting the high dynamic motion requirements of robots. The dynamic quality assessment mechanism can adaptively shield low-confidence data, and can still output reliable environmental information stably in complex environments such as strong light, occlusion, and vibration. Its robustness is better than traditional fusion solutions.
[0044] 3. Strong versatility and scalability: The sensor interface module supports access from multiple protocols and types of devices. The fusion algorithm parameters can be dynamically configured to adapt to different types of robots, such as industrial, service, and mobile robots. There is no need to redesign hardware and algorithms for specific scenarios, which reduces development costs.
[0045] 4. High hardware feasibility: The chip adopts a lightweight design, balancing computing power and power consumption, making it suitable for embedded integration. The electronic device supports modular expansion, and sensors and storage resources can be flexibly configured according to needs, facilitating large-scale mass production and engineering implementation, and promoting the widespread application of robots in high-precision and complex scenarios. Attached Figure Description
[0046] Figure 1 This is a flowchart illustrating a multi-source data fusion perception method for robots proposed in this invention.
[0047] Figure 2 This is a flowchart illustrating the calculation method of the standard confidence threshold in a robot multi-source data fusion perception method proposed in this invention.
[0048] Figure 3 This is a flowchart illustrating the method for obtaining the complexity coefficient in a robot multi-source data fusion perception method proposed in this invention.
[0049] Figure 4 This is a flowchart illustrating the vision-lidar feature fusion method in a robot multi-source data fusion perception method proposed in this invention.
[0050] Figure 5 This is a flowchart illustrating the IMU-odometer feature fusion method in a robot multi-source data fusion sensing method proposed in this invention.
[0051] Figure 6 This is a flowchart illustrating the global state fusion method in a robot multi-source data fusion perception method proposed in this invention. Detailed Implementation
[0052] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the invention, and should not be construed as limiting the invention. Rather, embodiments of the invention include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.
[0053] Example 1 refer to Figure 1 This embodiment proposes a multi-source data fusion perception method for robots. By constructing an adaptive fusion architecture, it achieves accurate synchronization, quality assessment, and intelligent fusion of data from multiple types of sensors, outputting high-precision environmental data to provide reliable support for robot motion planning and significantly improve the robot's flexibility and environmental adaptability. The robot multi-source data fusion perception method includes the following steps: S1. Raw data is collected from the surrounding environment using a multimodal sensor array. This raw data is then preprocessed to obtain standardized environmental data. The multimodal sensor array can be mounted on a robot, or it can be used to connect to sensors installed in the environment; details are omitted here. The sensor array includes at least vision sensors (such as a vision system composed of multiple cameras, a monocular camera, and / or an RGB-D camera), an inertial measurement unit (IMU), lidar, and / or force sensors; specific configurations depend on the robot's intended use and are not detailed here. The collected raw data cannot be used directly and requires preprocessing, which may include time synchronization, spatial alignment, noise filtering, and data standardization. Specifically, time synchronization employs a hardware interrupt triggering mechanism to generate a unified microsecond-level timestamp, calibrating the sampling time of the multimodal sensor array to synchronize it and eliminate timing deviations. Spatial alignment includes: constructing a robot base coordinate system and obtaining the coordinate system of the original data; then, based on the preset extrinsic parameter matrix of the multimodal sensor array, translating the coordinate systems of different sensors to unify them to the robot base coordinate system, achieving spatial position matching. Noise filtering uses corresponding algorithms for different sensor characteristics; visual data can use Gaussian filtering to remove noise, IMU data can use moving average filtering to suppress vibration interference, and LiDAR data can use pass-through filtering to remove invalid point clouds at long distances. Data standardization maps various types of data to a unified numerical range, providing consistent input for subsequent fusion calculations. For example, an industrial camera acquires images of the workpiece and its surrounding environment (sampling frequency 30Hz), an MPU6050 acquires the robot's joint angular velocity and acceleration (sampling frequency 500Hz), a lidar scans the distribution of obstacles in the working area (sampling frequency 10Hz), and a force sensor detects the contact pressure between the end effector and the workpiece (sampling frequency 100Hz). The chip synchronously acquires data through the sensor interface module, completes time synchronization calibration and spatial coordinate alignment, uses Gaussian filtering and moving average filtering to remove various types of noise, and then normalizes it to the [0, 1] region.
[0054] S2. A multi-dimensional data quality assessment model is constructed to quantify the confidence level of standardized environmental data and obtain the confidence score. The assessment dimensions of the multi-dimensional data quality assessment model can include: data integrity, signal-to-noise ratio, feature recognition, and temporal stability. A multi-dimensional weighted scoring method is then used to quantify and calculate the confidence score. The specific process is as follows: First, weights are assigned to each dimension based on sensor characteristics and application scenarios (summing to 1). For example, in an indoor robot scenario, the visual sensor emphasizes feature recognition (weight 0.25), the IMU emphasizes temporal stability (weight 0.2), the LiDAR emphasizes data integrity (weight 0.35), and the force sensor emphasizes data integrity (weight 0.2). Then, the acquisition coefficient G for each dimension is calculated. ij (Score range 0~1), where i is the standardized environmental data number, such as image data obtained from a visual sensor, and j is the dimension number. Collection coefficient G i j This is the j-th dimension of the standardized environmental data, numbered i. (Data integrity data) , where n i N is the effective data volume of the standardized environmental data collection numbered i. i This refers to the theoretical sampled data volume of the standardized environmental data, designated as i, which can be obtained from the collected data information. (Signal-to-noise ratio data, data integrity data) (Upper limit 1), where n SNR-i It is the actual SNR of the standardized environmental data collection numbered i, N SNR-i The target SNR of the standardized environmental data is numbered i; the feature discrimination is calculated by combining the sensor type through the proportion of feature points or the deviation from the baseline, for visual sensors and LiDAR, feature discrimination. , where n i0 ' is the number of valid feature points collected in the standardized environmental data collection, numbered i, and N is the number of valid feature points collected. i0 'This refers to the total number of feature points extracted from standardized environmental data numbered i. Taking an industrial robot grasping scenario as an example: the vision sensor extracts a total of 200 feature points of the workpiece contour. After screening and removing 30 overlapping points and 10 background interference points, the number of effective feature points is 160; the feature recognition score is 160 / 200 = 0.8. IMU and force sensor: Feature recognition score In this formula, n i1 It is the actual data value collected in the standardized environmental data collection, numbered i, and n. i1 ' is the baseline data value for the standardized environmental data numbered i, N i1 'i' represents the maximum allowable deviation of standardized environmental data. Taking a robot navigation scenario (IMU angular velocity monitoring) as an example: actual data value = 1.21 rad / s, baseline data value = 1.2 rad / s, maximum allowable deviation = 0.02 rad / s; deviation = |1.21 - 1.2| = 0.01 rad / s; feature recognition score = 1 - 0.01 / 0.02 = 0.5. Temporal stability. , where s i S is the unit-time data fluctuation variance of the standardized environmental data numbered i. i This is the maximum permissible variance for the standardized environmental data numbered i. Final confidence value. , where w j G is the weight value of dimension j, where J is the total number of dimensions, j=1, 2, ..., J; i jThis refers to the collection coefficient of the j-th dimension of the standardized environmental data, numbered i. This method allows us to obtain the confidence level of each standardized environmental data point. The confidence level ranges from 0 to 1, with higher confidence levels closer to 1, thus enabling accurate evaluation of each standardized data point and facilitating subsequent calculations. The confidence level calculation method may also include confidence level values. , where w j G is the weight value of dimension j, where J is the total number of dimensions, j=1, 2, ..., J; i j G is the collection coefficient of the j-th dimension of the standardized environmental data, numbered i. i j 'This refers to the standard data for the j-th dimension of the standardized environmental data, designated as i. Its average or median can be calculated based on various historical collection coefficients, or a set standard value can be established, such as 0.8. Other data settings are also possible, but will not be elaborated upon here. This method can quickly weaken data with large deviations in various collection coefficients, thus avoiding data distortion caused by excessive deviations and providing an accurate data foundation for subsequent calculations, thereby improving the accuracy of later calculations.
[0055] S3. Determine each confidence level G. i Is it greater than the corresponding standard confidence threshold g? i If yes, the standardized environmental data is calibrated as valid data and its weight coefficient is marked; otherwise, the standardized environmental data is judged as low-quality data, its weight is reduced (provisionally considered valid data) or the channel data is temporarily blocked to avoid error amplification. For example, an industrial camera with a confidence level of 0.9 under normal lighting (above the threshold of 0.7), LiDAR data with a confidence level of 0.95, IMU data with a confidence level of 0.85, and a force sensor with a confidence level of 0.9 when in contact with a workpiece are all judged as valid data and their corresponding weight coefficients are marked. In scenarios with strong light obstruction, the feature recognition of a vision sensor decreases, and its confidence level falls below the threshold; the system automatically reduces its weight, details of which are not elaborated here. Standard confidence threshold g i The settings can be configured with fixed data based on experience, or a "fixed baseline threshold + dynamic compensation" mechanism can be used to ensure adaptability to different scenarios. (Reference) Figure 2 Specifically, this may include: first, presetting a baseline threshold g according to the work scenario. i For example, the complexity coefficient is 0.7 for high-precision scenarios (such as industrial assembly), 0.6 for ordinary scenarios (such as indoor navigation), and 0.5 for complex dynamic scenarios (such as outdoor exploration); the complexity coefficient is calculated through environmental features. (Reference) Figure 3Methods for obtaining the complexity coefficient may include: selecting standardized environmental data from the first 3 frames, focusing on the core dimensions of environmental representation; for visual / LiDAR systems: selecting the variance of key pixel grayscale values in the feature map and the coordinates of the point cloud cluster centers; for IMU systems: selecting angular velocity (ω... x / ω γ / ωz), linear acceleration (a x / a γ / az); For each core dimension, calculate the variance of the first 3 frames of data; sum the variances of all core dimensions with weights to obtain the comprehensive fluctuation variance of the first 3 frames of data. The weights of each dimension in the weighted sum can be set according to the importance of the sensor, such as a weight of 0.3 for visual features and 0.25 for IMU angular velocity, with all weights summing to 1. Map the comprehensive fluctuation variance to the interval [-0.2, 0.3] to obtain the complexity coefficient. The mapping method can use a linear mapping formula, which will not be elaborated here. For example, assume the core dimensions and variances of the first 3 frames of data: visual feature grayscale variance σ1 2 =0.03 (weight 0.3); IMU angular velocity variance σ² 2 =0.04 (weight 0.25); force pressure variance σ3 2 =0.02 (weight 0.2); LiDAR cluster center variance σ4 2 =0.05 (weight 0.25); given σmin=0.01, σmax=0.1. Calculation process: Comprehensive fluctuation variance σ = 0.03×0.3 + 0.04×0.25 + 0.02×0.2 + 0.05×0.25 = 0.009 + 0.01 + 0.004 + 0.0125 = 0.0355; Complexity coefficient = (0.0355-0.01) / (0.1-0.01)×0.5-0.2 = (0.0255 / 0.09)×0.5-0.2 ≈ 0.1417-0.2 = -0.0583 (within the range [-0.2, 0.3], no truncation required); Select the "high confidence fusion result with confidence ≥ 0.8" from the last 100 frames as the reference frame set S, calculate the mean of each environmental state dimension (such as position coordinates, attitude angle, obstacle position) in S, and use it as the calibration benchmark value Xr. The fusion results of nearly 100 frames are denoted as Xm, where m is the data frame number. The relative deviation of each frame from the calibration reference value is calculated using the formula: Single Frame Relative Deviation (Xr≠0) and fusion bias rate δt (mean of relative bias over 100 frames). If δt > 3%, the fusion result bias is too large, indicating that the current threshold is too low and low-confidence data is mixed in. The preset baseline threshold g is used. i 'Raise the threshold by 0.05; if 1%≤δt≤3%: the fusion accuracy meets the standard, maintain the current threshold unchanged; if δt<1%: the fusion deviation is acceptable but there is room for optimization, preset the baseline threshold g.' iLower the threshold by 0.05 (to retain more data and improve robustness); the standard confidence threshold after calibration should be limited to the range of [0.4, 0.9] to avoid data filtering failure due to excessively low or high standard confidence thresholds. For example, if the calibration baseline value Xr = 500mm (robot X-axis position) of the fusion results of the last 100 frames, and the sum of the relative deviations of the 100 frames = 2.8, then: the fusion deviation rate δt = 2.8 / 100 = 2.8% (in the range of 1%~3%), the current dynamic threshold = 0.66, and the threshold after calibration = 0.66 - 0.05 = 0.61. If the sum of the relative deviations of the last 100 frames is 3.5, then δ_total = 3.5% (>3%), and the threshold after calibration = 0.66 + 0.05 = 0.71. This invention obtains a standard confidence threshold that accurately matches the dynamic characteristics of the environment. Based on the variance of the fluctuations in the first three frames of data, it calculates the complexity coefficient, achieving adaptive adjustment where "the more complex the environment, the higher the threshold; the more stable the environment, the lower the threshold." This avoids the problem of traditional fixed thresholds mixing in low-confidence data in complex scenes (strong light, vibration, occlusion) or excessive filtering leading to insufficient data redundancy in stable scenes, ensuring a strong binding between the threshold and the real-time environmental state. It integrates the variance of fluctuations from multiple core dimensions of sensors such as vision, IMU, LiDAR, and force sensing, rather than data from a single sensor, avoiding misjudgments of complexity due to local anomalies in a particular sensor, making the basis for threshold adjustment more comprehensive and reliable. Every 100 frames, it uses high-reliability historical data for backtracking and comparison, dynamically calibrating the threshold based on the fusion deviation rate. If the fusion result deviation is too large (>3%), it indicates that the threshold is too low and contains invalid data, and it is adjusted upwards in a timely manner; if the deviation is too small (<1%), the threshold is appropriately lowered to retain more redundancy, forming a closed loop of "dynamic adjustment - deviation monitoring - calibration optimization," preventing the threshold from deviating from the optimal range during long-term use. When the environment is highly complex (e.g., with many dynamic obstacles or large fluctuations in sensor data), the threshold is significantly increased, retaining only high-confidence data for fusion to minimize the interference of low-quality data on the fusion results and ensure perception accuracy in complex scenarios. When the environment is stable (e.g., in static operations or with small fluctuations in sensor data), the threshold is appropriately reduced, allowing more data to participate in fusion. Data redundancy enhances the anti-interference capability of the fusion results and prevents perception interruptions caused by temporary anomalies in a single sensor.
[0056] S4. Perform multi-source data fusion on effective data to obtain optimal state data. Specifically, this can be achieved using a "local feature fusion - global state fusion" method. The local feature fusion layer focuses on targeted collaborative extraction and preliminary fusion of similar / same-dimensional feature data. This eliminates redundancy in single-sensor data, strengthens the expression of local environmental features, and provides high-quality feature input for global fusion. See below for details. Figure 4(1) Visual-LiDAR feature fusion: The visual sensor performs convolution and pooling operations on the preprocessed image data in a lightweight manner to extract target contours (such as the edge contours of workpieces in industrial scenes and the contour features of the human body in service scenes) and texture features (such as the surface texture of workpieces and the texture of ground materials), and outputs a visual feature map with dimensions [H×W×C] (H is the image height, W is the width, and C is the number of feature channels); the LiDAR data is based on the spatial coordinates and density information of the point cloud, and the spatial feature points of obstacles (such as walls, equipment, and dynamic obstacles) are selected to generate a three-dimensional point cloud feature set, which is then mapped to the visual feature through projection transformation. Figure 1 A two-dimensional plane is used to obtain the LiDAR feature map. Finally, feature stitching and an attention mechanism are used to fuse the two types of feature maps. The attention weight is determined by feature similarity calculation (formula: attention weight = cosine similarity between visual features and LiDAR features / sum of similarities of all feature pairs). The weighted fusion generates a dense environmental feature map, preserving visual texture details while maintaining LiDAR spatial positioning accuracy. Example: In an industrial robot operation scenario, a visual sensor extracts the workpiece contour and surface texture features (feature map size 256×256×64). The LiDAR extracts the workpiece's three-dimensional spatial coordinates and surrounding obstacle features through clustering, mapping them to a 256×256×64 LiDAR feature map. After cosine similarity calculation, the visual feature weight for the workpiece area is 0.6, and the LiDAR feature weight is 0.4. The fusion yields a dense feature map containing workpiece details and spatial location. Reference Figure 5 (2) IMU-odometry feature fusion: The IMU acquires the robot's angular velocity (ω) x ω γ , ω_z) and linear acceleration (a x a γ The robot's attitude angles (pitch, roll, yaw) are initially estimated using data from the a_z data and integrated to obtain preliminary values. Odometry (such as wheeled odometry or visual odometry) calculates the robot's displacements (Δx, Δy, Δz) to obtain trajectory data. The state equation is... Where k is the discretization time node number, X k-1 It is the state vector at time point k-1, containing attitude angle, displacement, and velocity; X k Here, is the state vector at time point k, i.e., the state vector at the current time point. A is the state transition matrix; B is the control matrix; u is the input value of the IMU at time point k; and w is the process noise at time point k, the specific method of which is existing technology and will not be elaborated here. The observation equation is... Z k The current time point represents the odometry observation; H represents the observation matrix; v kFor the observation noise at time point k, the IMU cumulative error and the odometry slippage error are corrected through prediction-update iteration, and the preliminary estimation results of robot attitude and motion trajectory are output. Example: In the navigation scenario of mobile robot, the IMU collects 500 sets of angular velocity and acceleration data per second, and the initial estimated values of attitude angles are obtained by integration (pitch angle 0.5°, roll angle 0.3°, heading angle 90.2°). The odometry collects displacement data (Δx=0.5m, Δy=0, Δz=0). After EKF fusion, the attitude angles are corrected to pitch angle 0.48°, roll angle 0.29°, and heading angle 90.0°. The trajectory error is reduced from ±3mm to ±1mm. (3) Force feature extraction: For the contact pressure (F) collected by the force sensor x F γ The force sensor collects contact state features (such as contact mark = 1, pressure level = 0.5, and force direction = Z) and force direction data. A sliding window filter is used to remove instantaneous interference, and then feature normalization maps the data to the [0, 1] range. Contact state features (such as whether there is contact, contact pressure level, and force direction) are extracted separately. This does not require fusion with other sensors; it serves only as supplementary features for close-range interaction scenarios, avoiding interference with global environmental features. Example: When an industrial robot grasps a workpiece, the force sensor collects a contact pressure of 50N, with the force direction along the negative Z-axis. After filtering and normalization, the contact state features are output (contact mark = 1, pressure level = 0.5, force direction = Z). - This provides a basis for subsequent grasping posture adjustments. Global State Fusion Layer: Based on the fusion results of each local layer, a global environment state vector is constructed. Then, an adaptive weighted fusion algorithm is used to calculate the contribution of each local result, eliminate feature conflicts, and output the optimal state data based on the maximum contribution. Reference Figure 6 The specific process is as follows: Global environment state vector construction: Integrate the results of each local fusion and define the state vector X=[x, y, z, α, β, γ, O1, O2, ..., O n [E], where (x, y, z) are the robot's three-dimensional position coordinates, (α, β, γ) are the attitude angles, and O1~O n E represents the position and size information of each obstacle, and E represents the contact state characteristics (output of the force sensing module), achieving a unified representation of the robot's own state and the state of its surrounding environment. The contribution is calculated based on the confidence scores corresponding to the standardized data obtained in step S2. The contribution can be obtained by methods such as obtaining paired subsets based on the current scene and valid data, and then calculating the contribution based on these paired subsets. Where l is the paired subset number, f is the valid data number, sl is the valid data number in the paired subset with number l, and G f γ is the confidence level corresponding to the valid data with ID f, F is the total number of valid data in the paired subset with ID l, f = 1, 2, ..., F; L is the total number of paired subsets, l = 1, 2, ..., L; lThe scene adaptation coefficient corresponding to the pairing subset numbered l. It is the average confidence level corresponding to the paired subsets. Based on the optimal state data X from the previous frame. k-1 Predict the current frame state using the state transition matrix A. A is the state transition matrix; the prediction covariance matrix is also included. Where A is the state transition matrix, P k-1 A is the optimal posterior covariance matrix of the (k-1)th frame. T Z is the transpose of the state transition matrix, and Q is the process noise covariance matrix, which is dynamically adjusted according to contribution; local modules with higher contributions have lower noise weights. The fusion results of each local module are used as the observation value Z. k The observation matrix H is determined by combining the contribution values (the values of each observation dimension in H represent the corresponding contribution values), and the intermediate coefficients are calculated. Where R is the observation noise covariance matrix, reflecting the measurement accuracy of the sensor, and H... T The contribution determines the transpose of the observation matrix. The final update yields the optimal state data. This eliminates feature conflicts and redundancies in the local fusion results and outputs the globally optimal environmental state. For example, in an industrial robot grasping scene (scene adaptation coefficients of 1.2, 1.0, and 0.8 respectively), step S2 obtains the confidence scores of each sensor: visual confidence score of 0.9, LiDAR confidence score of 0.95, IMU confidence score of 0.85, odometry confidence score of 0.8, and force perception confidence score of 0.9. Calculate the local mean values: C1=(0.9+0.95) / 2=0.925, C2=(0.85+0.8) / 2=0.825, C3=0.9; initial weights W1=0.925×1.2=1.11, W2=0.825×1.0=0.825, W3=0.9×0.8=0.72; contribution P1=1.11 / (1.11+0.825+0.72)≈0.41, P2≈0.30, P3≈0.29. During global fusion, the vision-LiDAR fusion result contributes the most, dominating the position representation of workpieces and obstacles in the environment; the IMU-odometry fusion result assists in correcting the robot's posture; force features supplement the contact state, and after Kalman filtering prediction and updating, the optimal state data is output. A layered strategy of "local feature fusion + global state fusion" is adopted, and visual-LiDAR environmental features, IMU-odometer attitude trajectory and force contact features are complementarily verified, which effectively reduces the noise interference of a single sensor. The errors of core state dimensions such as position and attitude are controlled within the acceptable range in engineering, and the fusion deviation rate is stable and extremely small.
[0057] S5. The optimal state data output by fusion is compared with historical data for consistency verification. A sliding window mechanism is used to compare multiple consecutive frames of data to calculate the comprehensive deviation value. The sliding window mechanism is existing technology and will not be elaborated here. The method for calculating the comprehensive deviation value can be... , where X k,f It is the fused data of the f-th dimension of the current frame (such as the x-axis position of the k-th frame), X w,k,f It is the arithmetic mean of the f-th dimension data of all frames within the sliding window. It is the historical data mean of the f-th dimension within the current frame sliding window, and ε is a local minimum to prevent... When the denominator is 0, the dimensional weights w are then assigned. f '(The sum of all weights is 1), calculate the overall deviation value.' , where w f ' is the weight of the assigned dimension for valid data numbered f. Its value is based on the "priority of the state dimension" (e.g., position x / y / z is more critical to navigation and has a higher weight), and is fine-tuned in combination with sensor confidence, but the core is the "dimension priority". The specific details of how it is obtained will not be elaborated here.
[0058] S6. Determine if the overall deviation value is within the deviation threshold range. If yes, confirm the result is valid and output to the robot motion controller; otherwise, trigger the feedback adjustment mechanism to readjust the data quality assessment threshold and fusion weight coefficient, optimize the fusion algorithm parameters, and ensure the accuracy of subsequent outputs. The deviation threshold range can be set through the assembly scenario, and can be determined separately or comprehensively. For example, in an industrial robot assembly scenario, the x / y / z position threshold is 5% (or 2mm, taking the maximum value); the attitude angle α / β / γ is 10% (or 0.1°); and the obstacle distance is 8%; the overall threshold = 0.4×5%+0.3×10%+0.3×8%=6.9%. This is just a simple example; other cases will not be elaborated here. The feedback adjustment mechanism can include: identifying the source of the positioning deviation (quickly pinpointing the problem); adjusting the data quality assessment threshold (filtering better data); and optimizing the fusion weight coefficient (balancing data contribution). Specific details will not be elaborated here.
[0059] Example 2 This embodiment also proposes a robot multi-source data fusion sensing chip. The chip integrates the above-mentioned multi-source data fusion sensing method and adopts a heterogeneous computing architecture, including a sensor interface module, a data preprocessing module, a quality assessment module, a hierarchical fusion computing module, a verification feedback module, and a control interface module. Each module is interconnected through an on-chip bus to achieve high-speed data transmission and parallel processing.
[0060] Sensor interface module: integrates multiple protocol interfaces such as I2C, SPI, UART, and CAN, supports plug-and-play for various types of devices such as vision sensors, IMU, LiDAR, and force sensors, realizes synchronous acquisition and transmission of raw data, and the interface rate can reach up to 80MHz to meet the data transmission requirements of high-frequency sensors.
[0061] Data preprocessing module: Built-in hardware acceleration unit, which implements preprocessing algorithms such as time synchronization and noise filtering in hardware, and supports real-time data processing with a sampling frequency of 500Hz.
[0062] Quality assessment module: Enables multi-dimensional confidence assessment, with a built-in configurable threshold register, supporting dynamic adjustment of assessment parameters according to different application scenarios.
[0063] Layered fusion computing module: integrates a lightweight CNN accelerator and a Kalman filter hardware engine. The CNN accelerator adopts channel pruning and parameter sharing mechanisms to reduce computing power consumption while ensuring feature extraction accuracy; it supports adaptive weight adjustment to achieve high-speed solution of global state fusion.
[0064] Verification feedback module: It has a built-in data caching unit and deviation detection circuit, which caches 10 consecutive frames of fusion results. It implements consistency verification through hardware logic and can quickly update fusion parameters when feedback adjustment is triggered.
[0065] Control interface module: Provides interfaces such as PWM and Ethernet to transmit the integrated environmental data to the robot motion controller and decision module, while receiving external control commands to realize the collaborative work between the chip and the robot system.
[0066] Example 3 This embodiment proposes an electronic device, which includes the aforementioned multi-source data fusion sensing chip, power module, storage module, robot motion controller, and sensor array. The sensor array is electrically connected to the sensor interface module of the chip. The storage module is used to store fusion algorithm parameters, historical data, and configuration files. The power module provides stable power to the chip and sensor array. The robot motion controller parses the effective optimal state data and generates running instructions based on the optimal state data.
[0067] This electronic device can be integrated into various robot bodies such as industrial robots, service robots, and mobile robots. It connects to the robot control system through a standardized interface, supports modular expansion, and can add or remove sensor types according to robot application scenarios to adapt to different operation requirements such as indoor and outdoor, high precision, and high dynamics.
[0068] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0069] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for multi-source data fusion sensing in robots, characterized in that, The robot multi-source data fusion perception method includes the following steps: S1. Collect raw data from the surrounding environment using a multimodal sensor array, and then preprocess the raw data to obtain standardized environmental data; S2. The confidence level G is obtained by quantifying the standardized environmental data using a constructed multi-dimensional data quality assessment model. i ; S3. Determine each confidence level G. i Is it greater than the corresponding standard confidence threshold g? i If yes, the standardized environmental data is calibrated as valid data; otherwise, the standardized environmental data is judged as low-quality data. S4. Perform multi-source data fusion on the effective data to obtain the optimal state data; S5. Perform consistency verification between the optimal state data output by fusion and historical data. Compare multiple consecutive frames of data through a sliding window mechanism to calculate the comprehensive deviation value. S6. Determine whether the overall deviation value is within the deviation threshold range. If it is, confirm that the result is valid; if not, trigger the feedback adjustment mechanism.
2. The robot multi-source data fusion sensing method according to claim 1, characterized in that, The sensor array includes: a vision sensor, an inertial measurement unit, a lidar, and / or a force sensor.
3. The robot multi-source data fusion sensing method according to claim 1, characterized in that, Preprocessing of raw data includes: time synchronization, spatial alignment, noise filtering, and data standardization.
4. The robot multi-source data fusion sensing method according to claim 1, characterized in that, The evaluation dimensions of the multi-dimensional data quality assessment model include: data integrity, signal-to-noise ratio, feature discriminability, and temporal stability.
5. The robot multi-source data fusion sensing method according to claim 1, characterized in that, The multi-dimensional data quality assessment model process is as follows: First, weights are assigned to each dimension based on sensor characteristics and application scenarios; then, the acquisition coefficients G for each dimension are calculated. i j , where i is the standardized environmental data number and j is the dimension number.
6. The robot multi-source data fusion sensing method according to claim 1, characterized in that, Standard confidence threshold g i The methods for obtaining the threshold include: first, pre-setting a baseline threshold g according to the work scenario. i The complexity coefficient is calculated based on environmental features; the fusion results with a confidence level ≥ 0.8 from the last 100 frames are selected and denoted as the reference frame set S; the mean value of each environmental state dimension in S is calculated as the calibration benchmark value Xr; the fusion results from the last 100 frames are denoted as Xm, where m is the data frame number; the relative deviation of each frame from the calibration benchmark value and the fusion deviation rate δt are calculated; if δt > 3%: the fusion result deviation is too large, and the preset benchmark threshold g is set. i 'Raise the threshold by 0.05; if 1%≤δt≤3%: the fusion accuracy meets the standard, maintain the current threshold unchanged; if δt<1%: preset the baseline threshold g.' i Lower the threshold by 0.05; the calibrated standard confidence threshold should be limited to the interval [0.4, 0.9] as the standard confidence threshold g. i .
7. A robot multi-source data fusion sensing method according to claim 6, characterized in that, The method for obtaining the complexity coefficient includes: selecting the first 3 frames of standardized environmental data and focusing on the core dimensions of environmental representation; calculating the variance of the first 3 frames of data for each core dimension; weighted summing of the variances of all core dimensions to obtain the comprehensive fluctuation variance of the first 3 frames of data; and mapping the comprehensive fluctuation variance to the interval [-0.2, 0.3] to obtain the complexity coefficient.
8. A robot multi-source data fusion sensing method according to claim 1, characterized in that, The optimal state data is specifically processed using a local feature fusion-global state fusion method; the local feature fusion layer consists of: (1) visual-LiDAR feature fusion; and (2) IMU-odometer feature fusion. (3) Force feature extraction; The global state fusion layer includes: global environment state vector construction: integrating the results of each local fusion and defining the state vector X; The contribution is calculated based on the confidence level, using the optimal state data X from the previous frame. k-1 Predict the current frame state using the state transition matrix A. and the predicted covariance matrix The results of each local fusion are used as the observation value Z. k The observation matrix H is determined by combining the contribution, and the intermediate coefficients are calculated. The final update yields the optimal state data. .
9. A chip, characterized in that, The chip is used to execute the robot multi-source data fusion perception method according to any one of claims 1-8.
10. An electronic device, characterized in that, The electronic device includes the chip of claim 9.