An intelligent driving multi-sensor fusion data processing system
Through the intelligent driving multi-sensor fusion data processing system with data preprocessing, feature-level and decision-making fusion modules, the problem of insufficient perception and decision-making accuracy in the prior art is solved, precise perception and reliable decision-making in complex traffic scenarios are achieved, and the adaptability and safety of the system are enhanced.
Patent Information
- Application Number
- CN202510580736.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing intelligent driving multi-sensor fusion data processing technology has problems such as insufficient accuracy and poor system expansion in the perception and decision-making links, and it is difficult to adapt to complex and changeable traffic scenarios, resulting in insufficient driving safety and system stability.
The data preprocessing unit is used to standardize, clean and synchronize the format of the sensor data. The feature-level fusion module extracts multi-sensor features through deep neural networks. The decision-level fusion module uses weighted voting and fuzzy logic inference to generate driving decisions, and combines the data post-processing unit to optimize decisions and feedback sensor parameters.
It improves the accurate perception and reliability of sensor data and the reliability of decision-making, enhances the adaptability and response speed of the system, can operate stably under different models and road conditions, reduces misjudgments and misjudgments, and improves the safety and efficiency of intelligent driving.
Smart Images

Figure CN120105350B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fusion data processing, and particularly to an intelligent driving multi-sensor fusion data processing system. Background Art
[0002] With the rapid development of intelligent driving technology, multi-sensor fusion data processing has become a core element in enhancing driving safety and intelligence. In today's complex and ever-changing traffic scenarios, vehicles need to accurately perceive surrounding environmental information and then make reliable decisions, which poses strict requirements on multi-sensor fusion systems.
[0003] Traditional intelligent driving multi-sensor fusion data processing solutions have many limitations. At the perception level, most early systems simply spliced data collected by different sensors without deeply exploring the internal relationships of the characteristics of each sensor. For example, lidar data was only roughly used to obtain the general outline of the target, and the visual information of the camera was not effectively combined with geometric features, resulting in poor target recognition accuracy. In complex intersections and areas with mixed traffic of people and vehicles, it often occurs that the target is misidentified or cannot be accurately positioned, and it is difficult to clearly distinguish the details of different types of vehicles and pedestrian postures, and key road conditions are easily missed, laying hidden dangers for driving decisions.
[0004] There are also deficiencies in the decision-making link. In the past, a fixed weight method was often used to fuse the preliminary decisions of each sensor without considering the performance fluctuations of the sensors in different scenarios. In sunny days, the weight of the visual information of the camera is too high. Once encountering bad weather such as heavy rain and thick fog, the camera is severely interfered by light and raindrops, but still participates in the decision-making according to the established high weight, resulting in frequent errors in the final driving instructions and causing unnecessary braking and steering misoperations. In the face of fuzzy road condition information, the traditional binary decision-making mode lacks a flexible response mechanism, and the yes-or-no judgment standard cannot fit a large number of boundary-fuzzy situations in actual driving, with a low error tolerance and difficult to ensure driving safety.
[0005] In addition, the scalability and adaptability of the system are even more short boards of traditional solutions. The old architecture is rigid. When accessing new sensors, the original design needs to be overthrown and the algorithm needs to be greatly reconstructed, with a long R & D cycle and high cost. And there is no effective linkage with the subsequent data processing link, and it is impossible to optimize the sensor parameters and fusion algorithm according to the vehicle driving history data and real-time road condition feedback, resulting in the system being difficult to adapt to different vehicle models and road conditions. In special scenarios such as mountain curves and urban congested sections, the response is laggy and the stability is poor, and it is difficult to meet the continuous advancement needs of intelligent driving.
[0006] In summary, the existing intelligent driving multi-sensor fusion data processing technology urgently needs to be improved to overcome the above problems in order to fit the increasingly complex and diverse application scenarios of intelligent driving and ensure driving safety and efficient operation of the system. Summary of the Invention
[0007] The main object of the present invention is to provide an intelligent driving multi-sensor fusion data processing system, which can effectively solve the problems mentioned in the background art.
[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0009] An intelligent driving multi-sensor fusion data processing system includes a data preprocessing unit, a data fusion unit, and a data postprocessing unit:
[0010] The data preprocessing unit includes a data format standardization module, a data cleaning and noise reduction module, and a data calibration and synchronization module, which are used to sequentially perform format unified conversion, noise and abnormal data elimination, and calibration and synchronization of spatial coordinates and time on the original data received from lidar, cameras, millimeter wave radars, and ultrasonic sensors.
[0011] The data fusion unit is provided with a feature-level fusion module and a decision-level fusion module. The feature-level fusion module is used to extract corresponding features from the data of each sensor, and adopts a deep neural network architecture to fuse different features into a fusion feature vector including target space, vision, motion, and close-range information. The decision-level fusion module uses a classifier to carry out fine target classification decisions, and at the same time independently conducts motion and close-range situation assessments based on the data of millimeter wave radars and ultrasonic sensors, and then uses a weighted voting and fuzzy logic reasoning strategy to fuse the preliminary decision results to generate driving decision instructions.
[0012] The data postprocessing unit includes a decision optimization and risk assessment module and a data storage and feedback module, which optimize the driving decision instructions in combination with the vehicle's real-time state and road environment information, estimate the risk probability, and are also responsible for storing key data and feedback information to optimize the sensor parameters and fusion algorithms, and push warning information as needed.
[0013] The multi-sensor fusion data processing steps of the system are as follows:
[0014] Data acquisition and transmission step: Each sensor is started synchronously, data is collected at a predetermined frequency and time stamps are attached, and the data is transmitted to the data preprocessing unit via the in-vehicle high-speed bus.
[0015] Data preprocessing step: The original data is sequentially processed by the data format standardization module, the data cleaning and noise reduction module, and the data calibration and synchronization module, and is stored in the buffer for fusion.
[0016] Data fusion step: At the feature-level fusion module, features of each sensor are extracted in parallel, multi-feature deep fusion is completed to generate a fusion feature vector and sent to the decision-level fusion module; after receiving it, the decision-level fusion module first classifies and evaluates, then fuses the preliminary decision results, and outputs accurate driving instructions to the data postprocessing unit.
[0017] Data post - processing steps. The data post - processing unit optimizes instructions, evaluates risks, stores data, feeds back information, and pushes warnings according to set rules.
[0018] Preferably, the data format standardization module converts lidar point cloud data into a structured array containing three - dimensional coordinates and reflection intensity information, unifies camera image data into a digital image matrix with a specific resolution and color - coding format, arranges millimeter - wave radar data into a data set containing target distance, speed, and angle parameters, processes ultrasonic sensor data into a simple distance numerical sequence, and attaches accurate timestamps to each data.
[0019] The data cleaning and denoising module, for lidar, uses statistical filtering combined with outlier removal algorithms to screen out noise and outlier data based on local point cloud density thresholds, distance thresholds, and point cloud distribution models; for camera images, it uses median filtering, Gaussian filtering, and image repair algorithms to remove salt - and - pepper noise, high - frequency noise, and image defects; millimeter - wave radar filters false echoes with a constant false alarm rate (CFAR) detection algorithm; the ultrasonic sensor calibrates the ranging value according to the temperature compensation model and eliminates measurement errors caused by temperature fluctuations.
[0020] The data calibration and synchronization module constructs a coordinate transformation matrix, maps each sensor's data to the vehicle coordinate system based on the pre - calibrated sensor installation position and angle parameters; uses a high - precision atomic clock or in - vehicle synchronous bus signal to stamp synchronous timestamps on the data, and eliminates time deviations caused by acquisition frequency differences and transmission delays through linear interpolation and timestamp re - alignment techniques.
[0021] Preferably, when the feature - level fusion module extracts the geometric features of lidar, it uses a point cloud segmentation algorithm to divide the point cloud clusters of target objects, calculates the dimensions of the target bounding box, and obtains the dimension parameters by finding the coordinate extreme values along the coordinate axes; calculates the volume by using the convex hull algorithm to decompose the convex hull into triangular pyramids and accumulate the volumes; calculates the surface flatness by selecting the neighboring points of the points and fitting a local plane with principal component analysis, and calculates the standard deviation of the distances from the points to the plane; calculates the centroid position according to the coordinate mean formula; the specific extraction method is as follows:
[0022] Point cloud segmentation: Use the Euclidean clustering algorithm for point cloud segmentation. For each point in the point cloud data set , calculate its Euclidean distance from other points ; set a distance threshold , if the distance between point and point is less than or equal to the threshold, then group them into the same cluster; traverse the entire point cloud data set to gradually form different clusters, and each cluster represents the point cloud cluster of a potential target object.
[0023] Target bounding box size calculation: For each segmented target point cloud cluster, find the maximum value of the point cloud coordinates along the x, y, and z coordinate axes respectively. , , and the minimum value , , ; The target bounding box size parameter is , and its size information can intuitively reflect the occupancy range of the target in three-dimensional space.
[0024] Volume calculation: Assume that the target object is approximately convex hull-shaped, use the QuickHull algorithm to construct the convex hull, and the convex hull vertex set is . Decompose the convex hull into multiple simple triangular pyramids, and calculate the volume of each triangular pyramid through the triangular pyramid volume formula , where is the area of the bottom triangle, is the height, and accumulate to get the total volume of the target object.
[0025] Surface flatness calculation: Select the nearest neighbor points of each point in the target point cloud cluster ( The value of is determined according to experience or experiment, such as ), use the principal component analysis (PCA) method to fit the local plane; calculate the standard deviation of the distances from these points to the fitted plane . The smaller the value, the higher the surface flatness; otherwise, the surface is rougher.
[0026] Centroid position calculation: The centroid coordinates are calculated by the following formula: , , , where is the number of points in the target point cloud cluster, , , are the coordinates of each point in the point cloud.
[0027] Preferably, the feature-level fusion module extracts camera visual features using a deep learning object detection model. After convolution and downsampling, a feature map is generated. Use anchor boxes to match the target prediction class probability, bounding box coordinate offset, and confidence score. Then, through non-maximum suppression, high-quality detection results are retained. For the retained regions, further convolution and fully connected processing are performed to output visual feature vectors composed of shape, texture, color features, and class probabilities. The specific extraction method is as follows:
[0028] The YOLOv5 model is selected for object detection and visual feature extraction; after the image is input into the YOLOv5 network, through several convolutional layers and downsampling layers, feature maps of different scales are generated. On the feature maps, the pre-set anchor boxes are used to match with the target objects, and each anchor box predicts the class probability (covering common categories such as cars, trucks, pedestrians, bicycles, etc.), the bounding box coordinate offset ( , , , ), and the object confidence score . Through the non-maximum suppression (NMS) algorithm, the overlapping and redundant detection boxes are removed, and the high-quality detection results are retained.
[0029] For the retained object detection box regions, visual features are further extracted. The corresponding region images are processed through convolutional layers and fully connected layers, and the output shape features (such as the aspect ratio of the object, the contour curve information), texture features (using the gray-level co-occurrence matrix GLCM to statistically analyze the gray-level changes of pixels in different directions and distances), and color features (extracting the mean and variance information of the RGB channels) are combined with the class probability information to form a visual feature vector for subsequent fusion.
[0030] Preferably, the feature-level fusion module extracts the motion features of the millimeter-wave radar, collects the distance, radial velocity, and angle information at fixed time intervals, obtains the target radial acceleration through differential operations, calculates the angular velocity using the angle information of two frames, and combines them into the form of a motion feature vector. The specific extraction method is as follows:
[0031] The millimeter-wave radar continuously outputs the distance , radial velocity , and angle information of the target relative to the radar, and collects the data sequence at a fixed time interval . The target radial acceleration is calculated through differential operations: , where is the radial velocity at the current moment, and is the radial velocity at the previous moment.
[0032] The target angular velocity is calculated through the angle information of two consecutive frames: , reflecting the change rate of the target motion direction; these motion features are combined into the vector form to describe the dynamic characteristics of the target.
[0033] Preferably, the feature-level fusion module extracts the short-distance obstacle features of the ultrasonic sensor, sets the distance threshold classification, and generates the feature values representing the short-distance target state according to the comparison between the measured distance and the threshold, in the form of binary or graded.
[0034] Preferably, when the feature-level fusion module fuses features, a fusion convolutional neural network is constructed. The feature branches of the lidar and the camera each undergo convolution, pooling, and fully connected operations. The convolution uses a convolution kernel of a specific size, a stride, and a padding method. The pooling is max pooling, and then the channels are concatenated to output a joint feature. For the millimeter-wave radar and ultrasonic features, a recurrent neural network branch is introduced, and the long short-term memory unit is used to capture dynamic changes. Finally, the outputs of each branch are concatenated into a fused feature vector. The specific implementation method is as follows:
[0035] Construction and training of the fusion convolutional neural network:
[0036] Construct a multi-branch convolutional neural network architecture, which is divided into a lidar geometric feature branch and a camera visual feature branch. The lidar geometric feature vector is used as input and connected to a sub-network containing multiple convolutional, pooling, and fully connected layers. The camera visual feature vector is similarly connected to another independent but structurally similar sub-network. In the convolutional layer, a convolution kernel with a stride of 1 and a padding method of SAME is used for feature extraction. The calculation formula is: , where is the feature map of the th layer, is the convolution kernel weight, and is the input feature of the th layer.
[0037] The pooling layer uses max pooling with a window size of and a stride of 2 to downsample the feature map, reduce the data volume, and retain key information. After being processed by their respective sub-networks, the feature maps output by the two branches are concatenated in the channel dimension and then connected to the subsequent fully connected layer. The number of neurons in the fully connected layer is adjusted according to the experiment and task complexity, such as setting it to 256 neurons. Finally, a joint feature representation vector is output, reflecting the target space and visual fusion features.
[0038] Processing of motion and close-range features by the recurrent neural network branch:
[0039] For the millimeter-wave radar motion feature vector and the ultrasonic close-range obstacle feature, a recurrent neural network branch composed of long short-term memory units is introduced. Inside the long short-term memory unit, the core calculation steps are as follows:
[0040] Forget gate: , where is the forget gate weight matrix, is the previous hidden state, is the current millimeter-wave radar motion feature vector, and is the current ultrasonic close-range obstacle feature vector. is the bias vector, is the sigmoid function.
[0041] Input gate: , which controls the degree to which the current input information is incorporated into the memory cell.
[0042] Update memory cell: , , to achieve the update of the memory cell.
[0043] Output gate: , , to determine the current output hidden state, capture the dynamic changes of the motion and close-range states over time. After being processed by multiple long short-term memory cells, the output is the motion and close-range fusion feature vector.
[0044] Generate the fusion feature vector: Concatenate the spatial-visual fusion feature vector output by the convolutional neural network branch and the motion and close-range fusion feature vector output by the recurrent neural network branch in dimension to form the final fusion feature vector. Its dimension is the sum of the output dimensions of the two branches. This vector completely encompasses the target's space, vision, motion, and close-range information and is fed into the decision-level fusion module.
[0045] Preferably, when the decision-level fusion module performs target classification and situation assessment, it uses a support vector machine to classify the target, constructs an SVM model to find the optimal classification hyperplane, sets the objective function and constraints, and uses the Lagrange multiplier method to solve the dual problem to obtain the optimal solution. Based on this, the decision function value of the newly input fusion feature vector is calculated to judge the class membership. Specifically, it includes:
[0046] Fine target classification based on SVM and MLP:
[0047] Use a support vector machine (SVM) to classify the target, taking the fusion feature vector obtained by feature-level fusion as the input. For binary classification problems (such as distinguishing vehicles from pedestrians), constructing an SVM model aims to find the optimal classification hyperplane. The objective function is: , and the constraint conditions are , , where is the normal vector of the hyperplane, is the bias, is the fusion feature vector, is the class label. Use the Lagrange multiplier method to solve the dual problem to obtain the optimal solution and . For the newly input fusion feature vector , by calculating the decision function value , judge the class membership based on its positive or negative value.
[0048] The multi-layer perceptron (MLP) is used for multi-classification scenario expansion. The network includes an input layer, multiple hidden layers (the number of neurons in the hidden layer can be set to a decreasing structure such as 128 and 64), and an output layer. The ReLU function is selected as the activation function, and the softmax function is used in the output layer to convert the output into the probability distribution of each category. For example, , where is the weight matrix of the output layer, is the output of the last hidden layer, is the bias of the output layer, and the category with the highest probability is the target classification result.
[0049] Situation assessment based on millimeter-wave radar and ultrasonic sensors:
[0050] According to the millimeter-wave radar data, set the speed threshold , and the acceleration threshold , . If the target radial velocity and the acceleration , it is determined that the target is approaching at high speed, and the danger level is recorded as high; if and , the danger level is recorded as medium; if and , the danger level is recorded as low.
[0051] Combined with the short-range obstacle characteristics of the ultrasonic sensor, when the characteristic value is 2 (extremely close), directly increase the overall danger level; if it is 1 (close) and the millimeter-wave radar also indicates that the target is approaching, also adjust the danger level to a higher state, and comprehensively form a situation assessment result for subsequent decision-making fusion.
[0052] Preferably, when the decision-level fusion module fuses to generate the final driving decision, when using weighted voting to fuse the preliminary decisions, weights are dynamically assigned to each sensor according to different scenarios, and the weighted sum of the comprehensive scores is used to determine the driving operation; when using fuzzy logic reasoning, fuzzy sets and membership functions are defined, a fuzzy rule base is formulated, and after fuzzy composition and centroid defuzzification, the driving instruction is output, specifically including:
[0053] Weighted voting method:
[0054] Weights are assigned to each sensor according to different scenarios. For example, in a scenario of strong direct sunlight, the reliability of the camera is reduced due to light interference, and the weight is set to ; the weight of the lidar is set to ; the millimeter-wave radar is less affected by light, and the weight is set to ; the weight of the ultrasonic sensor is set to . When the SVM classification target is a pedestrian and the confidence level is 0.8 (denoted as ), the danger level of the millimeter-wave radar situation assessment is medium (denoted as ), the ultrasonic sensor indicates an obstacle in the near distance (denoted as ), the weighted voting comprehensive score . According to the predefined threshold, if , output a deceleration command; if it is between 0.3 - 0.5, maintain the current vehicle speed and strengthen monitoring; if it is less than 0.3, drive normally.
[0055] Fuzzy logic reasoning:
[0056] Define fuzzy sets and membership functions. For the target distance, set fuzzy sets of "near", "medium", and "far", for the speed, set fuzzy sets of "slow", "medium", and "fast", and for the danger level, set fuzzy sets of "high", "medium", and "low". For example, for the target distance, use a triangular membership function. If the actual target distance is , the membership function of the "near" fuzzy set is: , where , are distance parameters, which are set according to the actual scenario.
[0057] Formulate a fuzzy rule base, such as "IF the lidar target distance is near AND the camera identifies a pedestrian AND the millimeter-wave radar detects a fast speed AND the ultrasonic wave indicates an obstacle in the near distance THEN the danger level is high". Calculate the membership values of the fuzzy sets based on the data of each sensor, obtain the fuzzy output result through fuzzy synthesis operations (such as the Mamdani inference method), and finally use the centroid method for defuzzification to convert the fuzzy result into a clear driving decision instruction. Assume that the membership distribution of the fuzzy output danger level is [0.2 (low), 0.3 (medium), 0.5 (high)], and the corresponding driving decision strength values are [0 (no operation), 0.5 (deceleration), 1 (emergency braking)]. The defuzzification calculation formula is: , calculate to get , and accordingly output the corresponding driving decision instruction, such as a deceleration or braking operation.
[0058] Preferably, the data post-processing unit includes a decision optimization and risk assessment module, a data storage and feedback module, where:
[0059] The decision optimization and risk assessment module is used to optimize and adjust the fused decision instructions by combining the vehicle speed, acceleration, steering angle, tire grip information during real-time driving and the road curvature, slope, lane width, and traffic rule information of the road environment. When a deceleration instruction is received during high-speed driving, it optimizes the braking deceleration curve based on the vehicle speed and vehicle dynamics model to prevent sudden braking and loss of control. When turning, it precisely fine-tunes the steering angle instruction in combination with the road curvature to ensure a smooth turn. At the same time, based on historical data and the current situation, it predicts the risk probabilities of potential collisions and lane departures and adjusts the driving strategy in advance.
[0060] The data storage and feedback module is responsible for storing the original sensor data, preprocessing results, fused feature vectors, decision-making process, and vehicle driving status information in a local database according to the time series for accident backtracking, algorithm performance evaluation, and system fault troubleshooting. It regularly statistically analyzes the stored data, mines the performance trends of sensors and the data fusion effect, and adjusts the sensor parameters and optimizes the fusion algorithm based on the obtained results to achieve the adaptive upgrade of the system.
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] 1. The feature-level fusion module deeply mines the characteristics of each sensor, comprehensively extracts the geometric features of lidar, the visual features of cameras, the motion features of millimeter-wave radars, and the close-range obstacle features of ultrasonic sensors, abandons the traditional simple fusion method, and uses a deep neural network architecture to make various features fully complementary and deeply intertwined. In complex traffic scenarios, such as busy intersections and highway entrances and exits where vehicles frequently change lanes, it can accurately outline the surrounding target contours and carefully capture their dynamic changes, greatly reducing the missed and misjudged situations caused by fuzzy and one-sided perception, laying a precise perception foundation for subsequent decision-making, and enabling the intelligent driving system to clearly understand the surrounding conditions.
[0063] 2. The decision-level fusion module first uses classifiers and data evaluations based on millimeter-wave radars and ultrasonic sensors to achieve multi-dimensional and precise analysis of road conditions. The weighted voting strategy flexibly adjusts the weights of each sensor according to the actual scenario, highlighting the advantages of millimeter-wave radars in bad weather and emphasizing lidar data in strong light environments to ensure that the decision-making conforms to the current road conditions and effectively avoids unreasonable instructions, greatly enhancing the accuracy and reliability of decision-making. The fuzzy logic inference system, relying on custom fuzzy sets and a rich rule base, can well handle the common fuzzy and uncertain situations in driving, output continuous and reasonable flexible decision instructions, avoid the limitation of decision-making falling into a black-and-white situation, and calmly respond to sudden or complex road conditions.
[0064] 3. From the perspective of the system architecture, the design of the data fusion unit features significant modularity and structurality. If new sensors are to be connected in the future, they can be efficiently integrated by referring to the existing feature extraction and fusion processes, which is conducive to rapid iterative upgrades, saving R & D costs and time. Moreover, it is closely linked with the data post-processing unit. The post-processing unit deeply analyzes the performance trends of sensors and the effectiveness of data fusion based on the vast amount of stored data, and provides accurate feedback accordingly to dynamically adjust sensor parameters and iterate fusion algorithms, greatly improving the response speed of the system under different vehicle models and various road conditions, significantly enhancing its adaptability, and enabling it to closely follow the changing requirements of intelligent driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a schematic diagram of the composition architecture process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The following will detail the implementation manners of the present invention in conjunction with the drawings and embodiments.
[0067] As Figure 1 shown, an intelligent driving multi-sensor fusion data processing system is specifically divided into three parts: a data preprocessing unit, a data fusion unit, and a data post-processing unit.
[0068] In practical applications, when an intelligent driving vehicle starts, lidar, cameras, millimeter-wave radars, and ultrasonic sensors are synchronously turned on for operation. The lidar quickly scans the surrounding environment of the vehicle at a frequency of 15Hz, emits laser beams and receives reflected light to form point cloud data containing rich three-dimensional spatial information. In a parking lot scenario, the lidar can clearly capture the contours and distances of parked vehicles around, as well as the position information of fixed obstacles such as columns and walls, providing basic data for vehicle starting and parking planning; the cameras capture high-definition images of the vehicle's surroundings at a frame rate of 30fps, covering front view, rear view, and panoramic views. The front view camera focuses on identifying lane lines, traffic signs, and vehicles and pedestrians ahead, just like the "eyes" of a driver, obtaining visual scene information in real time; the millimeter-wave radar continuously emits millimeter-wave signals and collects the distance, radial velocity, and angle information of the target relative to the radar at intervals of every 0.1s, being good at monitoring long-distance and high-speed moving targets. In a high-speed driving scenario, it can lock in vehicles approaching or moving away quickly several meters away in advance to assist in judging the safe following distance; the ultrasonic sensor detects the short-distance area around the vehicle at a high frequency of 20Hz, accurately positioning obstacles within a range of 0 - 2 meters. When the vehicle is moving slowly or parking, it closely monitors the distances between the four corners of the vehicle body and surrounding objects to prevent scratches. Each sensor attaches an accurate timestamp to the collected data and quickly transmits it to the data preprocessing unit via the in-vehicle high-speed CAN bus.
[0069] I. Data Preprocessing
[0070] 1. Standardization of Data Format
[0071] The data format standardization module quickly gets to work. The lidar point cloud data is parsed and processed, converted into a structured array form, which details the three-dimensional coordinates and reflection intensity information of each point cloud, facilitating direct call and operation by subsequent algorithms; the camera image data is uniformly adjusted to a digital image matrix with a resolution of 1920×1080 and an RGB888 color encoding format, meeting the input requirements of the deep learning model; the millimeter-wave radar organizes the scattered signal data collected into a regular data set, clearly listing the distance, speed, and angle values of each target; the ultrasonic sensor data is streamlined into a simple distance value sequence, and the timestamps carried by each data are calibrated to be synchronized with other sensors and stored in a dedicated buffer, ready to be cleaned and denoised.
[0072] 2. Data Cleaning and Denoising
[0073] For lidar data, a statistical filtering algorithm is launched, setting local point cloud density thresholds and distance thresholds to screen out isolated noise points and outlier data caused by abnormal reflections. For example, in a scene with strong direct sunlight, some laser beams are reflected too strongly to form false far points, which can be accurately removed by this algorithm; the camera image uses median filtering to remove salt-and-pepper noise, with a window size set to 3×3 to smooth image details, and then combines Gaussian filtering to suppress high-frequency noise. At the same time, an image repair algorithm repairs the bad points and defects caused by lens dirt and glare; the millimeter-wave radar uses a constant false alarm rate (CFAR) detection algorithm to filter out false echoes generated by multipath reflections and electromagnetic interference with a false alarm probability of 1%; the ultrasonic sensor calibrates the ranging value by substituting the real-time temperature value collected by the built-in temperature sensor into the temperature compensation formula to eliminate the measurement error caused by temperature fluctuations and improve the data purity. The processed data flows back to the buffer.
[0074] 3. Data Calibration and Synchronization
[0075] Based on the pre-precisely calibrated sensor installation positions and angle parameters, a coordinate transformation matrix is constructed. Assuming the lidar is installed in the center of the front of the vehicle roof, with an angle of 0 degrees and a height of 1 meter with respect to the vehicle's central axis, the point cloud data is mapped to the vehicle coordinate system based on this; the on-vehicle high-precision atomic clock signal is used to stamp synchronous timestamps on the data of each sensor. For the time deviation caused by the acquisition frequency difference and transmission delay, linear interpolation technology is used for adjustment. For example, if the lidar data lags behind for a short period of time due to transmission delay at a certain moment, the data for this period is filled in by linear interpolation to make the data of each sensor accurately synchronized in the time dimension, ensuring that the data collected at the same moment participates in the fusion process.
[0076] II. Data Fusion
[0077] 1. Feature-Level Fusion
[0078] Feature extraction:
[0079] For lidar geometric feature extraction, an Euclidean clustering algorithm is used for point cloud segmentation. For any point in the point cloud dataset, calculate its Euclidean distance from the surrounding points, set an appropriate distance threshold. If the distance between two points is less than the threshold, they are grouped into the same cluster. Traverse the dataset to form point cloud clusters of different target objects. When calculating the size of the target bounding box, find the coordinate extrema along the coordinate axes to accurately present the spatial occupancy range of targets such as vehicles; use the QuickHull algorithm to calculate the target volume, first construct a convex hull and then decompose it into triangular pyramids to accumulate the volume; select a certain number of nearest neighbor points for each point and use principal component analysis (PCA) to fit the local plane, and calculate the standard deviation of the distance from the points to the plane to evaluate the surface flatness; calculate the centroid position according to the coordinate mean formula to locate the target center of gravity.
[0080] For camera vision feature extraction, the YOLOv5 model is selected. The image is input into the network, and through multiple convolutional layers and downsampling layers, feature maps of different scales are generated. Use anchor boxes to match the targets, and predict the class probability (covering common classes), the coordinate offset of the bounding box, and the confidence score. Remove redundant detection boxes through non-maximum suppression. For the remaining regions, further process through convolutional and fully connected layers, and output shape features (analyze the aspect ratio and contour curve), texture features (statistical pixel gray level changes by the gray level co-occurrence matrix GLCM), and color features (extract the mean and variance of the RGB channels) to form a visual feature vector.
[0081] For millimeter-wave radar motion feature extraction, data is collected at fixed time intervals. Obtain the target radial acceleration through differential operations, calculate the angular velocity from the angle information of two consecutive frames, and combine them into a motion feature vector to describe the target dynamics.
[0082] For ultrasonic sensor close-range obstacle feature extraction, set thresholds such as extremely close, close, and relatively close, and generate feature values according to the measured distance comparison, which intuitively reflects the state of close-range targets.
[0083] Feature fusion:
[0084] Construct a fusion convolutional neural network (CNN). The lidar and camera feature branches each go through multiple convolutional, pooling, and fully connected operations. The convolutional layer selects convolutional kernels of appropriate sizes, and carefully designs the stride and padding methods to extract features; the pooling layer performs downsampling to reduce the data volume; the output feature maps of the two branches are concatenated in the channel dimension and connected to a fully connected layer with a specific number of neurons to output a joint feature representation vector.
[0085] Introduce a recurrent neural network (RNN) branch for millimeter-wave radar and ultrasonic features, and utilize long short-term memory units (LSTM). The forget gate controls memory retention; the input gate regulates information integration; the memory unit is updated to achieve dynamic storage; the output gate outputs the hidden state, and after multi-layer processing, the motion and close-range fusion feature vectors are output; finally, the outputs of the two branches are concatenated to form a fusion feature vector encompassing all-round information of the target, and sent to the decision-level fusion module.
[0086] (2) Decision-level fusion
[0087] Target classification and situation assessment:
[0088] Use a support vector machine (SVM) to classify targets, construct an SVM model to find the optimal classification hyperplane, set the objective function and constraints, and use the Lagrange multiplier method to solve the dual problem to obtain the optimal solution, and thus judge the class attribution of the newly input fusion feature vector; at the same time, based on the data of millimeter-wave radar and ultrasonic sensors, set speed thresholds, acceleration thresholds, and ultrasonic distance thresholds to evaluate the target motion and close-range situation.
[0089] Fusion to generate the final driving decision:
[0090] Weighted voting is used to fuse the preliminary decisions. The weight of the camera in the strong light direct shooting scenario is set to 0.2, lidar 0.3, millimeter-wave radar 0.4, and ultrasonic sensor 0.1. If the SVM classifies the target as a pedestrian with a high confidence, the millimeter-wave radar indicates that the target is approaching at a low speed, and the ultrasonic detects an obstacle at a close range, the weighted sum of the comprehensive scores determines the driving operation. Fuzzy logic reasoning defines fuzzy sets (such as "near", "medium", "far" fuzzy sets for target distance), membership functions, and formulates a fuzzy rule base (such as "IF the lidar target distance is near and the volume is large, the camera identifies it as a truck, the millimeter-wave radar detects a high speed, and the ultrasonic indicates an obstacle at a close range THEN the danger level is high"), and through fuzzy synthesis and the centroid method for defuzzification, driving instructions such as accelerating, decelerating, steering, and braking are output to precisely command the vehicle's actions.
[0091] III. Data post-processing
[0092] (1) Decision optimization and risk assessment
[0093] When the vehicle receives a deceleration command while driving at high speed, the decision-making optimization and risk assessment module combines the vehicle speed (such as 120 km / h) and the vehicle dynamics model to optimize the braking deceleration curve, avoiding sudden braking and loss of control; in the turning scenario, it precisely fine-tunes the steering angle command based on the road curvature (such as a curve radius of 50 meters) to ensure a smooth turn; at the same time, based on historical data (records of handling similar road conditions in the past) and the current situation (real-time sensor fusion data), it predicts the risk probabilities of potential collisions and lane departures, and adjusts the driving strategy in advance, such as giving an early warning, fine-tuning the vehicle speed, or maintaining a wider vehicle distance.
[0094] (2) Data storage and feedback
[0095] The data storage and feedback module stores key information such as raw sensor data, preprocessing results, fusion feature vectors, decision-making processes, and vehicle driving states in a local database according to the time series for accident backtracking, algorithm performance evaluation, and system fault troubleshooting; it regularly statistically analyzes the stored data to mine the performance trends of sensors and the data fusion effect. If it is found that the point cloud noise of the lidar increases during a certain period, it feeds back and adjusts its transmission power and scanning frequency; it optimizes the algorithm parameters according to the fusion effect to achieve the adaptive upgrade of the system and continuously improve the intelligent driving performance and safety.
[0096] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent driving multi-sensor fusion data processing system, characterized in that: Including data pre-processing unit, data fusion unit and data post-processing unit: The data preprocessing unit includes a data format standardization module, a data cleaning and noise reduction module, and a data calibration and synchronization module, which are used to perform unified format conversion, noise and abnormal data elimination, and calibration and synchronization of spatial coordinates and time on the raw data received from the laser radar, camera, millimeter wave radar, and ultrasonic sensor in sequence; The data fusion unit is provided with a feature-level fusion module and a decision-level fusion module. The feature-level fusion module is used to extract corresponding features from each sensor data, and adopts a deep neural network architecture to fuse different features into a fusion feature vector that includes target space, vision, motion and close-range information. The decision-level fusion module uses a classifier to carry out fine classification decisions of targets, and independently performs motion and close-range situation assessment based on millimeter-wave radar and ultrasonic sensor data, and then uses weighted voting and fuzzy logic reasoning strategies to fuse preliminary decision results to generate driving decision instructions; specifically including: Fusion convolutional neural network construction and training: Construct a multi-branch convolutional neural network architecture, which is divided into a lidar geometric feature branch and a camera visual feature branch. The lidar geometric feature vector is used as input and fed into a sub-network containing multiple convolutional, pooling, and fully-connected layers; the camera visual feature vector is also fed into another independent but structurally similar sub-network; in the convolutional layer, a 3×3 convolutional kernel with a stride of 1 and a padding method of SAME is used for feature extraction, and the calculation formula is: where F l is the feature map of the l-th layer, K ij is the convolutional kernel weight, and I l is the input feature of the l-th layer; The pooling layer uses maximum pooling with a window size of 2×2 and a step size of 2 to downsample the feature map, reduce the amount of data and retain key information. After being processed by their respective sub-networks, the feature maps output by the two branches are spliced in the channel dimension and then connected to the subsequent fully connected layer. The number of neurons in the fully connected layer is adjusted according to the complexity of the experiment and task, and finally the joint feature representation vector is output to reflect the target space and visual fusion features. The recurrent neural network branch processes motion and close-range features: For the millimeter-wave radar motion feature vectors [r, v r , a r , ω] and the ultrasonic short-distance obstacle features, a recurrent neural network branch composed of long short-term memory units is introduced. Inside the long short-term memory units, the core calculation steps are as follows: Forgotten gate: f t = σ(W f · [h t-1 , m t , u t + b f ), where W f is the forgotten gate weight matrix, h t-1 is the previous hidden state, m t is the current millimeter-wave radar motion feature vector, u t is the current ultrasonic short-range obstacle feature vector, b f is the bias vector, and σ is the sigmoid function; Input gate: i t = σ(W i · [h t-1 , m t , u t + b i ), which controls the degree to which the current input information is integrated into the memory unit; Update memory unit: Implement memory unit update; Output gate: o t = σ(W o · [h t-1 , m t , u t + b o ), h t = o t · tanh(c t ), determine the current output hidden state, capture the dynamic changes of motion and close - range state over time, and after being processed by multiple long - short - term memory units, output the motion and close - range fusion feature vector.
2. The intelligent driving multi-sensor fusion data processing system according to claim 1, wherein: The data format standardization module converts the laser radar point cloud data into a structured array containing three-dimensional coordinates and reflection intensity information, unifies the camera image data into a digital image matrix with a specific resolution and color coding format, organizes the millimeter wave radar data into a data set containing target distance, speed, and angle parameters, and processes the ultrasonic sensor data into a simple distance value sequence, and each data is accompanied by a precise timestamp; The data cleaning and noise reduction module uses statistical filtering combined with an outlier removal algorithm for laser radar to filter out noise and outlier data based on the local point cloud density threshold, distance threshold and point cloud distribution model; Median filtering, Gaussian filtering and image restoration algorithms are used on camera images to remove salt and pepper noise, high-frequency noise and image defects; millimeter-wave radar uses a constant false alarm rate detection algorithm to filter false echoes; ultrasonic sensors calibrate the distance value based on the temperature compensation model to eliminate measurement errors caused by temperature fluctuations; The data calibration and synchronization module constructs a coordinate transformation matrix, and maps the data of each sensor to the vehicle coordinate system according to the pre-calibrated sensor installation position and angle parameters; uses a high-precision atomic clock or an on-board synchronous bus signal to synchronize the data with a timestamp, and eliminates the time deviation caused by acquisition frequency differences and transmission delays through linear interpolation and timestamp realignment technology.
3. An intelligent driving multi-sensor fusion data processing system according to claim 1, characterized in that: When the feature-level fusion module extracts the lidar geometric features, it uses a point cloud segmentation algorithm to divide the target object point cloud clusters, calculates the size of the target bounding box, and obtains the size parameters by finding the coordinate extreme values along the coordinate axes; calculates the volume by using the convex hull algorithm to decompose the convex hull into triangular pyramids and accumulate the volumes; selects the nearest neighbor points of the points to fit the local plane with principal component analysis to calculate the standard deviation of the distances from the points to the plane for the surface flatness; calculates the centroid position according to the coordinate mean formula; the specific extraction method is as follows: Point cloud segmentation: The Euclidean clustering algorithm is used for point cloud segmentation. For each point P in the point cloud dataset i (x i , y i , z i ), calculate the Euclidean distance between it and other points Set the distance threshold T d . If the distance d i between point P j and point P ij < T d , then group them into the same cluster; traverse the entire point cloud dataset to gradually form different clusters, and each cluster represents the point cloud cluster of a potential target object; Target Bounding Box Size Calculation: For each segmented target point cloud cluster, find the maximum values x max , y max , z max and the minimum values x min , y min , z min along the x, y, and z coordinate axes respectively; the target bounding box size parameters are [x min , x max , y min , y max , z min , z max , and its size information can intuitively reflect the occupancy range of the target in three-dimensional space; Volume calculation: Assume that the target object is approximately convex hull-shaped, and use the QuickHull algorithm to construct the convex hull. The set of convex hull vertices is V = {v1, v2, …, v n}, decompose the convex hull into multiple simple triangular pyramids, and calculate the volume of each triangular pyramid through the triangular pyramid volume formula , where S base is the area of the bottom triangle and h is the height, and accumulate to obtain the total volume of the target object; Surface flatness calculation: Select the k-nearest neighbors of each point within the target point cloud cluster. The value of k is determined based on experience or experiments. Use the principal component analysis method to fit the local plane; calculate the standard deviation σ of the distances from these points to the fitted plane. d , σ d The smaller the value of σ, the higher the surface flatness; conversely, the rougher the surface. Calculation of the centroid position: Centroid coordinates It is calculated by the following formula: where N is the number of points in the target point cloud cluster, and x i , y i , z i are the coordinates of each point in the point cloud.
4. An intelligent driving multi-sensor fusion data processing system according to claim 3, characterized in that: When the feature-level fusion module extracts the camera visual features, it uses a deep learning object detection model. Through convolution and downsampling, it generates feature maps, uses anchor boxes to match the target to predict the class probability, the bounding box coordinate offset, and the confidence score. Then, through non-maximum suppression, it retains the high-quality detection results. For the retained regions, it further performs convolution and fully connected processing to output shape, texture, color features, and class probability to form a visual feature vector. The specific extraction method is as follows: The YOLOv5 model is selected for object detection and visual feature extraction; after the image is input into the YOLOv5 network, through several convolutional layers and downsampling layers, feature maps of different scales are generated. On the feature maps, the pre-set anchor boxes are used to match with the target objects, and each anchor box predicts the class probability P c , the bounding box coordinate offsets (Δx, Δy, Δw, Δh), and the object confidence score S conf , and the overlapping and redundant detection boxes are removed through the non-maximum suppression algorithm to retain the high-quality detection results; For the retained target detection box regions, visual features are further extracted. The corresponding region images are processed through convolutional layers and fully connected layers to output shape features, texture features, color features, which together with the class probability information form a visual feature vector for subsequent fusion.
5. An intelligent driving multi-sensor fusion data processing system according to claim 4, characterized in that: When the feature-level fusion module extracts the millimeter-wave radar motion features, it collects distance, radial velocity, and angle information at a predetermined time interval, obtains the target radial acceleration through differential operations, calculates the angular velocity using the angle information of two frames, and combines them into the form of a motion feature vector. The specific extraction method is as follows: The millimeter-wave radar continuously outputs the distance r, radial velocity v of the target relative to the radar r and the angle θ information, and collects data sequences at a fixed time interval Δt. The target radial acceleration a r is calculated through differential operation: where v r (t) is the radial velocity at the current moment, and v r (t - Δt) is the radial velocity at the previous moment; The target angular velocity ω is calculated from the angle information of two consecutive frames: reflecting the rate of change of the target's motion direction; these motion features are combined into a vector form [r, v r , a r , ω], which is used to describe the dynamic characteristics of the target.
6. An intelligent driving multi-sensor fusion data processing system according to claim 5, characterized in that: When the feature-level fusion module extracts the short-distance obstacle features of the ultrasonic sensor, it sets a distance threshold for grading, and generates a feature value representing the state of the short-distance target by comparing the measured distance with the threshold, which is in binary or graded form.
7. An intelligent driving multi-sensor fusion data processing system according to claim 6, characterized in that: The feature-level fusion module further includes generating a fused feature vector: concatenating the spatial-visual fused feature vector output by the convolutional neural network branch and the motion and short-distance fused feature vector output by the recurrent neural network branch in dimension to form a final fused feature vector, whose dimension is the sum of the output dimensions of the two branches. This vector completely encompasses the target's spatial, visual, motion, and short-distance information and is sent to the decision-level fusion module.
8. An intelligent driving multi-sensor fusion data processing system according to claim 1, characterized in that: When the decision-level fusion module performs target classification and situation assessment, it uses a support vector machine to classify the target. It constructs a support vector machine model to find the optimal classification hyperplane, sets the objective function and constraints, and uses the Lagrange multiplier method to solve the dual problem to obtain the optimal solution. Based on this, it calculates the decision function value of the newly input fused feature vector to determine the class membership.
9. The intelligent driving multi-sensor fusion data processing system according to claim 1, characterized in that: When the decision-level fusion module fuses to generate the final driving decision, it uses weighted voting to fuse the preliminary decisions and dynamically assigns weights to each sensor according to different scenarios, and determines the driving operation by weighted summation of the comprehensive scores; when using fuzzy logic reasoning, it defines fuzzy sets and membership functions, formulates a fuzzy rule base, and outputs driving instructions through fuzzy composition and centroid defuzzification.
10. An intelligent driving multi-sensor fusion data processing system according to claim 1, characterized in that: The data post-processing unit includes a decision optimization and risk assessment module and a data storage and feedback module, where: The decision optimization and risk assessment module is used to optimize and adjust the fused decision instructions by combining the vehicle speed, acceleration, steering angle, tire grip information during real-time driving, and the road curvature, slope, lane width, and traffic rule information of the road environment; The data storage and feedback module is responsible for storing the original sensor data, preprocessing results, fused feature vectors, decision-making process, and vehicle driving status information into the local database in time series for accident backtracking, algorithm performance evaluation, and system fault troubleshooting; regularly statistically analyze the stored data, mine the sensor performance trends and data fusion effects, and feedback and adjust the sensor parameters and optimize the fusion algorithm according to the obtained results to achieve the adaptive upgrade of the system.
Citation Information
Patent Citations
Urban rail train running limit foreign matter sensing method, system and device and medium
CN116573017A
Safe driving method and system based on multi-source data and storage medium
CN118429935A