Intelligent sensing system for multi-modal data fusion

Through an intelligent perception system with multimodal data fusion, combined with millimeter-wave radar and camera data, using deep learning models and volumetric Kalman filters and other technologies, the problems of inaccurate detection and tracking stability of traditional single sensors in intelligent traffic are solved, and high-precision and high-stability target perception and tracking are achieved.

CN120105331AInactive Publication Date: 2025-06-06HEBEI INST OF MACHINERY ELECTRICITY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510173024.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional single sensor applications are inaccurate in the fields of intelligent transportation and other fields due to environmental constraints. Vision sensors are prone to obstruction or missed detection in complex traffic scenarios, affecting tracking stability.

Method used

The intelligent perception system using a multimodal data fusion is adopted to collect and pre-process millimeter-wave radar data and camera image data through the data acquisition and pre-processing module, and physical and visual features are extracted from radar and camera data through the multimodal data fusion module, and fused with deep learning models. The intelligent perception module uses volume Kalman filter and Hungarian algorithm for state estimation and target correlation, and the object detection and behavior prediction module uses YOLOv5 and DeepSORT algorithms for object detection and trajectory prediction.

Benefits of technology

It realizes high-precision and high-stability target perception and tracking, improves the accuracy and robustness of perception, ensures continuous and accurate tracking of goals and trajectory prediction, and improves the intelligence level and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105331A_ABST
    Figure CN120105331A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent sensing system for multi-modal data fusion, and relates to the technical field of intelligent traffic. The multi-modal data fusion intelligent sensing system comprises a data acquisition and preprocessing module, a multi-modal data fusion module, an intelligent sensing module, a target detection and behavior prediction module and a system verification and optimization module, and high-precision and high-stability target sensing and tracking are realized by integrating millimeter wave radar and camera data. The multi-modal data fusion module effectively combines physical features and visual features, and improves the accuracy and robustness of perception. The intelligent sensing module adopts a volume Kalman filter and a Hungary algorithm, state estimation and target association are optimized, and continuous and accurate tracking of the target is ensured. The target detection and behavior prediction module utilizes an advanced YOLOv5 model and a DeepSORT algorithm to realize multi-target accurate detection and trajectory prediction, and the intelligent level of the system is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation technology, and in particular to an intelligent perception system for multi-modal data fusion. Background Art

[0002] Under the current technological background, target perception and tracking face many challenges in areas such as intelligent transportation. Traditional single sensor applications often result in inaccurate detection due to environmental factors, and visual sensors are susceptible to occlusion or missed detection in complex traffic scenarios, affecting tracking stability. To solve this problem, the industry has begun to explore multimodal data fusion technology in the hope of improving the accuracy and robustness of perception by integrating data from different sensors. However, how to effectively combine physical features with visual features, and how to achieve efficient fusion of multimodal data, are still the difficulties of current research. Summary of the invention

[0003] In view of the shortcomings of the prior art, the present invention provides an intelligent perception system with multimodal data fusion, which solves the problems that traditional single sensor applications often lead to inaccurate detection due to environmental factors, and that visual sensors are easily affected by occlusion or missed detection in complex traffic scenes, affecting tracking stability.

[0004] To achieve the above objectives, the present invention is implemented through the following technical solutions: an intelligent perception system for multimodal data fusion, comprising:

[0005] The data acquisition and preprocessing module is used to collect millimeter-wave radar data and camera image data, and align the time and space of radar and camera data after data preprocessing;

[0006] Multimodal data fusion module, which is used to extract the physical features of the target from the radar data, extract the visual features of the target from the camera image, build a deep learning model, and fuse the features of the radar and camera for collaborative perception;

[0007] The intelligent perception module uses a volumetric Kalman filter to combine radar and camera observation data to improve the state estimation model and achieve accurate target tracking and state estimation. It improves the estimation accuracy and stability of the filter through adaptive volume point selection. It uses the Hungarian algorithm to associate radar and camera observation data to ensure continuous tracking and trajectory prediction of the target. It also builds an interactive multi-model system to dynamically select and switch models based on the target's motion mode and state.

[0008] The target detection and behavior prediction module is based on the YOLOv5 model and combines the multimodal data of radar and camera to accurately detect targets. It also combines the DeepSORT algorithm to achieve continuous tracking and trajectory prediction of multiple targets. Based on the historical trajectory and current status, it builds a target behavior prediction model to predict the target's future movement trajectory and possible behavior pattern.

[0009] The system verification and optimization module is used to build a multi-scenario dataset training and verification system, and evaluate and optimize algorithms and models through performance indicators.

[0010] Preferably, the work contents of the data acquisition and preprocessing module specifically include:

[0011] A1. Millimeter-wave radar data acquisition: Receive basic information of target distance, speed, and azimuth sent by the millimeter-wave radar, and use adaptive Kalman filtering to smooth the data to reduce noise interference. The filter parameters are set to α = 0.1 and β = 0.01;

[0012] A2. Camera image data acquisition: receiving image data captured by the camera, applying Gaussian blur filtering to reduce image noise, and setting the blur radius r = 3;

[0013] A3. Data alignment and synchronization, including time alignment and time alignment:

[0014] A3.1. Time alignment: synchronize the radar and camera data according to the timestamp, and set the time synchronization error threshold Δt = 10ms;

[0015] A3.2. Spatial alignment: Use the calibration parameters to transform and align the coordinate systems of the radar and camera, and set the spatial alignment error threshold Δd = 5 cm.

[0016] Preferably, the working contents of the multimodal data fusion module specifically include:

[0017] B1. Multi-scale feature extraction:

[0018] B1.1. Radar features: extract basic information of target distance d, velocity v, and azimuth angle θ;

[0019] B1.2, Camera features: Extract the shape, color, and texture visual features of the target;

[0020] B2. Intelligent fusion strategy:

[0021] Build a deep learning model, set the input layer size to [radar feature dimension + camera feature dimension], and the output layer to target category and location information;

[0022] Using cross entropy loss function Loss = -∑y i·log(p i ), where y i is the true label, p i To predict the probability, the model parameters are optimized through the back-propagation algorithm.

[0023] Preferably, the working contents of the intelligent sensing and tracking module specifically include:

[0024] C1. Cubic Kalman filter optimization:

[0025] C1.1. Improve the state estimation model, combine the observation data of radar and camera, set the state vector and observation vector, as well as the process noise covariance matrix Q and the observation noise covariance matrix R;

[0026] C1.2, Adaptive volume point selection, adjust the number and distribution range of volume points according to the dynamic changes of the target state to improve the estimation accuracy and stability of the filter;

[0027] C2. Multimodal data association and adaptive algorithm:

[0028] C2.1, using the Hungarian algorithm for data association, combining radar and camera observation data to achieve continuous tracking and trajectory prediction of the target;

[0029] C2.2. Design an adaptive parameter adjustment mechanism to dynamically adjust the filter parameters and strategies according to system performance and environmental changes to improve the adaptability and robustness of the system.

[0030] Preferably, the improved state estimation model parameter setting includes: setting the state vector x = [d, v, θ, a] where a is acceleration; the observation vector z = [d obs , v obs ], where d obs and v obs are the distance and speed observed by the radar respectively; set the process noise covariance matrix Q and the observation noise covariance matrix R: Q = diag (σd 2 ,σv2,σθ2,σa 2 ), R = diag(σd obs 2 ,σv obs 2 );

[0031] The adaptive volume point selection parameter setting includes: setting the number of volume points N=2n according to the dynamic change of the target state, where n is the dimension of the state vector, and adaptively adjusting the distribution range of the volume points;

[0032] The Hungarian algorithm has an association threshold of τ=0.5 for data association, which means that observations with an association degree higher than 0.5 are associated with the target;

[0033] The parameters and strategies of the filter are dynamically adjusted to adjust the filter parameters α and β according to the target tracking accuracy, or to adjust the exposure time and gain of the camera image according to the weather conditions.

[0034] Preferably, the working contents of the target detection and behavior prediction module specifically include:

[0035] D1. Application of deep learning algorithms:

[0036] D1.1. Based on the YOLOv5 model, set the input image size to 640x640 and the batch size to batch size =16, learning rate lr = 0.01, number of training rounds epochs = 50;

[0037] D1.2. Combine the DeepSORT algorithm and set the matching threshold IoU threshold =0.5, maximum number of lost frames max lost =30, continuously track the target and predict its trajectory;

[0038] D2. Construction of interactive multi-model system:

[0039] D2.1. Model selection and switching mechanism:

[0040] According to the target's motion mode and state, an interactive multi-model system is constructed; the model switching probability threshold p is set switch =0.3, when the target state changes beyond the threshold, the model is switched;

[0041] D2.2, Target behavior prediction: Combine historical trajectory and current state to build a target behavior prediction model, use the Markov chain-based prediction model, set the transition probability matrix P, and predict the future state based on the current state of the target.

[0042] Preferably, the work contents of the system verification and optimization module specifically include:

[0043] E1. Dataset construction and training:

[0044] Build a dataset containing highway, severe weather, and complex traffic flow scenarios for system training and verification; improve the generalization ability of the model by adjusting the dataset size and scenario distribution;

[0045] Use large-scale data sets to train deep learning models, use early stopping strategies to avoid overfitting, and optimize model performance by adjusting validation set ratio parameters;

[0046] E2. Performance evaluation and optimization iteration:

[0047] Use precision, recall, F1 score, and MOTA indicators to comprehensively evaluate the system's target detection and tracking performance; find out the optimization direction by comparing the performance of different algorithms and models;

[0048] Based on the evaluation results and actual application needs, the algorithms and models are continuously optimized and iterated. By adjusting the hyperparameters of the deep learning model and optimizing the target detection and tracking algorithm strategies, the adaptability and reliability of the system are improved. At the same time, the system's functions and performance are continuously improved in combination with actual application scenarios and needs.

[0049] Preferably, the parameter setting of the dataset construction and training includes: setting the dataset size Dataset size = 10000, where each scene contains at least 100 objects;

[0050] When training a deep learning model, set the validation set ratio val ratio =0.2;

[0051] Set the precision rate to Precision, the recall rate to Recall, and the F1 score to F1 = 2×(Precision×Recall) / (Precision+Recall).

[0052] The present invention provides an intelligent perception system for multimodal data fusion. Compared with the prior art, it has the following beneficial effects:

[0053] 1. The multimodal data fusion intelligent perception system achieves high-precision and high-stability target perception and tracking by integrating millimeter-wave radar and camera data. The multimodal data fusion module effectively combines physical features with visual features to improve the accuracy and robustness of perception. The intelligent perception module uses the volumetric Kalman filter and Hungarian algorithm to optimize state estimation and target association, ensuring continuous and accurate tracking of targets. The target detection and behavior prediction module uses the advanced YOLOv5 model and DeepSORT algorithm to achieve accurate detection and trajectory prediction of multiple targets, further improving the intelligence level of the system. Finally, the system verification and optimization module ensures the performance of algorithms and models in practical applications, providing strong support for the continuous optimization of the system.

[0054] 2. The multi-modal data fusion intelligent perception system effectively reduces the noise interference of millimeter wave radar and camera data through the application of adaptive Kalman filtering and Gaussian fuzzy filtering, and improves the accuracy and reliability of data. At the same time, strict time alignment and space alignment strategies ensure the high synchronization and consistency of radar and camera data, laying a solid foundation for subsequent data fusion and intelligent perception. These improvements not only improve the overall performance of the system, but also provide strong support for applications in fields such as intelligent transportation.

[0055] 3. The multimodal data fusion intelligent perception system, the multimodal data fusion module extracts the multi-scale features of radar and camera, and combines them with deep learning models for intelligent fusion, which significantly improves the accuracy and robustness of target detection. This module can make full use of the complementary advantages of different sensors and effectively respond to complex environmental challenges. At the same time, the cross entropy loss function is used to optimize the model parameters, which further improves the recognition performance of the system and provides more reliable technical support for applications in fields such as intelligent transportation.

[0056] 4. The intelligent perception system of multimodal data fusion significantly improves the accuracy and stability of target tracking by optimizing the volumetric Kalman filter and adopting multimodal data association and adaptive algorithms. This module can make full use of the observation data of radar and camera to achieve accurate state estimation and trajectory prediction. At the same time, the adaptive volume point selection and parameter adjustment mechanism enhances the system's adaptability to dynamic environments and target state changes, and improves the system's robustness and practicality. These improvements provide more accurate and reliable technical support for applications in the fields of intelligent transportation, autonomous driving, etc.

[0057] 5. The intelligent perception system of multimodal data fusion significantly improves the real-time performance of target detection and the stability of tracking by integrating YOLOv5 and DeepSORT algorithms. At the same time, the constructed interactive multi-model system can flexibly switch models according to the target motion state, improving the accuracy of state estimation and behavior prediction. The behavior prediction model based on Markov chain further enhances the system's ability to predict future states. These improvements provide more accurate and efficient technical means for target monitoring and prediction in complex scenarios such as intelligent traffic management and autonomous driving, which helps to improve the safety and reliability of the system.

[0058] 6. The intelligent perception system of multimodal data fusion has effectively improved the generalization ability of the model by building a rich and diverse data set. Using large-scale data sets for training and verification, combined with the early stopping strategy, overfitting is avoided and the stability and accuracy of the model are ensured. At the same time, multiple performance indicators are used to comprehensively evaluate the system, providing a clear direction for optimization. By continuously optimizing iterative algorithms and models, combined with actual application scenarios and needs, this module has significantly improved the adaptability and reliability of the system, providing more robust technical support for practical applications in fields such as intelligent transportation. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a system module block diagram of the present invention;

[0060] Figure 2 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0062] See also Figure 1-Figure 2 , the present invention provides six technical solutions:

[0063] The first implementation mode: an intelligent perception system for multimodal data fusion, comprising:

[0064] The data acquisition and preprocessing module is used to collect millimeter-wave radar data and camera image data, and align the time and space of radar and camera data after data preprocessing;

[0065] Multimodal data fusion module, which is used to extract the physical features of the target from the radar data, extract the visual features of the target from the camera image, build a deep learning model, and fuse the features of the radar and camera for collaborative perception;

[0066] The intelligent perception module uses a volumetric Kalman filter to combine radar and camera observation data to improve the state estimation model and achieve accurate target tracking and state estimation. It improves the estimation accuracy and stability of the filter through adaptive volume point selection. It uses the Hungarian algorithm to associate radar and camera observation data to ensure continuous tracking and trajectory prediction of the target. It also builds an interactive multi-model system to dynamically select and switch models based on the target's motion mode and state.

[0067] The target detection and behavior prediction module is based on the YOLOv5 model and combines the multimodal data of radar and camera to accurately detect targets. It also combines the DeepSORT algorithm to achieve continuous tracking and trajectory prediction of multiple targets. Based on the historical trajectory and current status, it builds a target behavior prediction model to predict the target's future movement trajectory and possible behavior pattern.

[0068] The system verification and optimization module is used to build a multi-scenario dataset training and verification system, and evaluate and optimize algorithms and models through performance indicators.

[0069] By integrating millimeter-wave radar and camera data, high-precision and high-stability target perception and tracking are achieved. The multimodal data fusion module effectively combines physical and visual features to improve the accuracy and robustness of perception. The intelligent perception module uses the volumetric Kalman filter and Hungarian algorithm to optimize state estimation and target association, ensuring continuous and accurate tracking of targets. The target detection and behavior prediction module uses the advanced YOLOv5 model and DeepSORT algorithm to achieve accurate detection and trajectory prediction of multiple targets, further improving the intelligence level of the system. Finally, the system verification and optimization module ensures the performance of algorithms and models in practical applications, providing strong support for the continuous optimization of the system.

[0070] The second implementation mode is mainly different from the first implementation mode in that the working contents of the data acquisition and preprocessing module specifically include:

[0071] A1. Millimeter-wave radar data acquisition: Receive basic information of target distance, speed, and azimuth sent by the millimeter-wave radar, and use adaptive Kalman filtering to smooth the data to reduce noise interference. The filter parameters are set to α = 0.1 and β = 0.01;

[0072] A2. Camera image data acquisition: receiving image data captured by the camera, applying Gaussian blur filtering to reduce image noise, and setting the blur radius r = 3;

[0073] A3. Data alignment and synchronization, including time alignment and time alignment:

[0074] A3.1. Time alignment: synchronize the radar and camera data according to the timestamp, and set the time synchronization error threshold Δt = 10ms;

[0075] A3.2. Spatial alignment: Use the calibration parameters to transform and align the coordinate systems of the radar and camera, and set the spatial alignment error threshold Δd = 5 cm.

[0076] Through the application of adaptive Kalman filtering and Gaussian blur filtering, the noise interference of millimeter wave radar and camera data is effectively reduced, and the accuracy and reliability of the data are improved. At the same time, strict time alignment and space alignment strategies ensure the high synchronization and consistency of radar and camera data, laying a solid foundation for subsequent data fusion and intelligent perception. These improvements not only improve the overall performance of the system, but also provide strong support for applications in fields such as intelligent transportation.

[0077] The third implementation mode is mainly different from the first implementation mode in that the working contents of the multimodal data fusion module specifically include:

[0078] B1. Multi-scale feature extraction:

[0079] B1.1. Radar features: extract basic information of target distance d, velocity v, and azimuth angle θ;

[0080] B1.2, Camera features: Extract the shape of the target (such as aspect ratio ratio ), color (such as the H value hue in the HSV color space), texture (such as the contrast of the gray-level co-occurrence matrix GLCM) and other visual features;

[0081] B2. Intelligent fusion strategy:

[0082] Build a deep learning model (such as convolutional neural network CNN), set the input layer size to [radar feature dimension + camera feature dimension], and the output layer is the target category and location information.

[0083] Using cross entropy loss function Loss = -∑y i ·log(p i ), where y i is the true label, p i To predict the probability, the model parameters are optimized through the back-propagation algorithm.

[0084] The multimodal data fusion module extracts multi-scale features from radar and camera and combines them with deep learning models for intelligent fusion, significantly improving the accuracy and robustness of target detection. This module can fully utilize the complementary advantages of different sensors to effectively cope with complex environmental challenges. At the same time, the cross entropy loss function is used to optimize model parameters, further improving the recognition performance of the system and providing more reliable technical support for applications in fields such as intelligent transportation.

[0085] The fourth implementation mode is mainly different from the first implementation mode in that the working contents of the intelligent sensing and tracking module specifically include:

[0086] C1. Cubic Kalman filter optimization:

[0087] C1.1. Improve the state estimation model, combine the observation data of radar and camera, set the state vector and observation vector, as well as the process noise covariance matrix Q and the observation noise covariance matrix R; set the state vector x = [d, v, θ, a] where a is acceleration; set the observation vector z = [d obs , v obs ], where d obs and v obs are the distance and speed observed by the radar respectively; set the process noise covariance matrix Q and the observation noise covariance matrix R: Q = diag (σd 2 ,σv2,σθ2,σa 2 ), R = diag(σd obs2 ,σv obs 2 );

[0088] C1.2, Adaptive volume point selection, according to the dynamic changes of the target state, adjust the number and distribution range of the volume points to improve the estimation accuracy and stability of the filter, set the number of volume points N = 2n, where n is the dimension of the state vector, and adaptively adjust the distribution range of the volume points;

[0089] C2. Multimodal data association and adaptive algorithm:

[0090] C2.1, using the Hungarian algorithm for data association, with an association threshold of τ = 0.5, indicating that observations with a correlation degree higher than 0.5 are associated with the target, and combined with the observation data of the radar and camera to achieve continuous tracking and trajectory prediction of the target;

[0091] C2.2. Design an adaptive parameter adjustment mechanism to dynamically adjust the filter parameters and strategies according to system performance and environmental changes to improve the adaptability and robustness of the system. The dynamic adjustment of the filter parameters and strategies is to adjust the filter parameters α and β according to the target tracking accuracy, or adjust the exposure time and gain of the camera image according to weather conditions.

[0092] By optimizing the volumetric Kalman filter and adopting multimodal data association and adaptive algorithms, the accuracy and stability of target tracking are significantly improved. This module can make full use of the observation data of radar and camera to achieve accurate state estimation and trajectory prediction. At the same time, the adaptive volume point selection and parameter adjustment mechanism enhances the system's adaptability to dynamic environments and target state changes, and improves the system's robustness and practicality. These improvements provide more accurate and reliable technical support for applications in the fields of intelligent transportation, autonomous driving, etc.

[0093] The fifth implementation mode is mainly different from the first implementation mode in that the working contents of the target detection and behavior prediction module specifically include:

[0094] D1. Application of deep learning algorithms:

[0095] D1.1. Based on the YOLOv5 model, set the input image size to 640x640 and the batch size to batch size =16, learning rate lr = 0.01, number of training rounds epochs = 50;

[0096] D1.2. Combine the DeepSORT algorithm and set the matching threshold IoU threshold =0.5, maximum number of lost frames max lost =30, continuously track the target and predict its trajectory;

[0097] D2. Construction of interactive multi-model system:

[0098] D2.1. Model selection and switching mechanism:

[0099] According to the target's motion mode and state, an interactive multi-model system is constructed, such as a constant velocity model (CV), a constant acceleration model (CA), etc.; a model switching probability threshold p is set switch =0.3, when the target state changes beyond this threshold, the model is switched;

[0100] D2.2, Target behavior prediction: Combine historical trajectory and current state to build a target behavior prediction model, use the Markov chain-based prediction model, set the transition probability matrix P, and predict the future state based on the current state of the target.

[0101] By integrating YOLOv5 and DeepSORT algorithms, the real-time performance of target detection and the stability of tracking are significantly improved. At the same time, the constructed interactive multi-model system can flexibly switch models according to the target motion state, improving the accuracy of state estimation and behavior prediction. The behavior prediction model based on Markov chain further enhances the system's ability to predict future states. These improvements provide more accurate and efficient technical means for target monitoring and prediction in complex scenarios such as intelligent traffic management and autonomous driving, which helps to improve the safety and reliability of the system.

[0102] The sixth implementation mode is mainly different from the first implementation mode in that the work contents of the system verification and optimization module specifically include:

[0103] E1. Dataset construction and training:

[0104] Build a dataset containing highways, severe weather, and complex traffic flow scenarios for system training and verification; set the dataset size size = 10000, where each scene contains at least 100 objects; the generalization ability of the model is improved by adjusting the dataset size and scene distribution;

[0105] Use large-scale data sets to train deep learning models and set the validation set ratio val ratio =0.2, using the early stopping strategy to avoid overfitting, and optimizing the model performance by adjusting the validation set ratio parameter;

[0106] E2. Performance evaluation and optimization iteration:

[0107] The system's target detection and tracking performance is comprehensively evaluated using precision, recall, F1 score F1 = 2 × (Precision × Recall) / (Precision + Recall), and MOTA indicators; by comparing the performance of different algorithms and models, the optimization direction is found;

[0108] Based on the evaluation results and actual application needs, the algorithms and models are continuously optimized and iterated. By adjusting the hyperparameters of the deep learning model, optimizing the target detection and tracking algorithm and other strategies, the adaptability and reliability of the system are improved. At the same time, the system's functions and performance are continuously improved in combination with actual application scenarios and needs.

[0109] By building a rich and diverse data set, the generalization ability of the model is effectively improved. Using large-scale data sets for training and verification, combined with the early stopping strategy, overfitting is avoided and the stability and accuracy of the model are ensured. At the same time, multiple performance indicators are used to comprehensively evaluate the system, providing a clear direction for optimization. By continuously optimizing iterative algorithms and models, combined with actual application scenarios and needs, this module significantly improves the adaptability and reliability of the system, and provides more robust technical support for practical applications in fields such as intelligent transportation.

[0110] Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0111] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0112] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent perception system for multimodal data fusion, characterized in that: include: The data acquisition and preprocessing module is used to collect millimeter-wave radar data and camera image data, and align the time and space of radar and camera data after data preprocessing; Multimodal data fusion module, which is used to extract the physical features of the target from the radar data, extract the visual features of the target from the camera image, build a deep learning model, and fuse the features of the radar and camera for collaborative perception; The intelligent perception module uses a volumetric Kalman filter to combine radar and camera observation data to improve the state estimation model and achieve accurate target tracking and state estimation. It improves the estimation accuracy and stability of the filter through adaptive volume point selection. It uses the Hungarian algorithm to associate radar and camera observation data to ensure continuous tracking and trajectory prediction of the target. According to the target's motion mode and state, an interactive multi-model system is constructed to dynamically select and switch models; The target detection and behavior prediction module is based on the YOLOv5 model and combines the multimodal data of radar and camera to accurately detect targets; Combined with the DeepSORT algorithm, it can achieve continuous tracking and trajectory prediction of multiple targets; based on historical trajectories and current states, it builds a target behavior prediction model to predict the target's future motion trajectory and possible behavior patterns; The system verification and optimization module is used to build a multi-scenario dataset training and verification system, and evaluate and optimize algorithms and models through performance indicators.

2. The multimodal data fusion intelligent perception system according to claim 1, characterized in that: The work content of the data acquisition and preprocessing module specifically includes: A1. Millimeter-wave radar data acquisition: Receive basic information of target distance, speed, and azimuth sent by the millimeter-wave radar, and use adaptive Kalman filtering to smooth the data to reduce noise interference. The filter parameters are set to α = 0.1 and β = 0.01; A2. Camera image data acquisition: receiving image data captured by the camera, applying Gaussian blur filtering to reduce image noise, and setting the blur radius r = 3; A3. Data alignment and synchronization, including time alignment and time alignment: A3.

1. Time alignment: synchronize the radar and camera data according to the timestamp, and set the time synchronization error threshold Δt = 10ms; A3.

2. Spatial alignment: Use the calibration parameters to transform and align the coordinate systems of the radar and camera, and set the spatial alignment error threshold Δd = 5 cm.

3. The multimodal data fusion intelligent perception system according to claim 1, characterized in that: The working contents of the multimodal data fusion module specifically include: B1. Multi-scale feature extraction: B1.

1. Radar features: extract basic information of target distance d, velocity v, and azimuth angle θ; B1.2, Camera features: Extract the shape, color, and texture visual features of the target; B2. Intelligent fusion strategy: Build a deep learning model, set the input layer size to [radar feature dimension + camera feature dimension], and the output layer to target category and location information; Using cross entropy loss function Loss = -∑y i ·log(p i ), where y i is the true label, p i To predict the probability, the model parameters are optimized through the back-propagation algorithm.

4. The multimodal data fusion intelligent perception system according to claim 1, characterized in that: The working contents of the intelligent perception and tracking module specifically include: C1. Cubic Kalman filter optimization: C1.

1. Improve the state estimation model, combine the observation data of radar and camera, set the state vector and observation vector, as well as the process noise covariance matrix Q and the observation noise covariance matrix R; C1.2, Adaptive volume point selection, adjust the number and distribution range of volume points according to the dynamic changes of the target state to improve the estimation accuracy and stability of the filter; C2. Multimodal data association and adaptive algorithm: C2.1, using the Hungarian algorithm for data association, combining radar and camera observation data to achieve continuous tracking and trajectory prediction of the target; C2.

2. Design an adaptive parameter adjustment mechanism to dynamically adjust the filter parameters and strategies according to system performance and environmental changes to improve the adaptability and robustness of the system.

5. The multi-modal data fusion intelligent perception system according to claim 4, characterized in that: The improved state estimation model parameter setting includes: setting the state vector x=[d, v, θ, a], where a is acceleration; the observation vector z=[d obs , v obs ], where d obs and v obs are the distance and speed observed by the radar respectively; set the process noise covariance matrix Q and the observation noise covariance matrix R: Q = diag (σd 2 ,σv2,σθ2,σa 2 ), R = diag(σd obs 2 ,σv obs 2 ); The adaptive volume point selection parameter setting includes: setting the number of volume points N=2n according to the dynamic change of the target state, where n is the dimension of the state vector, and adaptively adjusting the distribution range of the volume points; The Hungarian algorithm has an association threshold of τ=0.5 for data association, which means that observations with an association degree higher than 0.5 are associated with the target; The parameters and strategies of the filter are dynamically adjusted to adjust the filter parameters α and β according to the target tracking accuracy, or to adjust the exposure time and gain of the camera image according to the weather conditions.

6. The multimodal data fusion intelligent perception system according to claim 1, characterized in that: The work content of the target detection and behavior prediction module specifically includes: D1. Application of deep learning algorithms: D1.

1. Based on the YOLOv5 model, set the input image size to 640x640 and the batch size to batch size =16, learning rate lr = 0.01, number of training rounds epochs = 50; D1.

2. Combine the DeepSORT algorithm and set the matching threshold IoU threshold =0.5, maximum number of lost frames max lost =30, continuously track the target and predict its trajectory; D2. Construction of interactive multi-model system: D2.

1. Model selection and switching mechanism: According to the target's motion mode and state, an interactive multi-model system is constructed; the model switching probability threshold p is set switch =0.3, when the target state changes beyond the threshold, the model is switched; D2.2, Target behavior prediction: Combine historical trajectory and current state to build a target behavior prediction model, use the Markov chain-based prediction model, set the transition probability matrix P, and predict the future state based on the current state of the target.

7. The multimodal data fusion intelligent perception system according to claim 1, characterized in that: The work content of the system verification and optimization module specifically includes: E1. Dataset construction and training: Build a dataset containing highway, severe weather, and complex traffic flow scenarios for system training and verification; improve the generalization ability of the model by adjusting the dataset size and scenario distribution; Use large-scale data sets to train deep learning models, use early stopping strategies to avoid overfitting, and optimize model performance by adjusting validation set ratio parameters; E2. Performance evaluation and optimization iteration: Use precision, recall, F1 score, and MOTA indicators to comprehensively evaluate the system's target detection and tracking performance; find out the optimization direction by comparing the performance of different algorithms and models; Based on the evaluation results and actual application needs, the algorithms and models are continuously optimized and iterated. By adjusting the hyperparameters of the deep learning model and optimizing the target detection and tracking algorithm strategies, the adaptability and reliability of the system are improved. At the same time, the system's functions and performance are continuously improved in combination with actual application scenarios and needs.

8. The multimodal data fusion intelligent perception system according to claim 7, characterized in that: The parameter settings for dataset construction and training include: setting the dataset size size = 10000, where each scene contains at least 100 objects; When training a deep learning model, set the validation set ratio val ratio =0.2; Set the precision rate to Precision, the recall rate to Recall, and the F1 score to F1 = 2×(Precision×Recall) / (Precision+Recall).

Citation Information

Cited By

  • Multi-variety vegetable harvester dynamic identification and feeding control system based on image processing

    CN120359903A

  • Vehicle-mounted monitoring system based on deep learning

    CN120385999A

  • Unmanned aerial vehicle target tracking and intelligent route planning method and system based on YOLO and DSM

    CN120740587A

  • ARM-based embedded remote monitoring camera

    CN120897124A

  • Mineral dynamic identification and accurate sampling method and system based on deep learning

    CN121564628A