A traffic control system based on multimodal data fusion

Through a multimodal data fusion system, using the synchronous acquisition and neural network analysis of image sensors and millimeter-wave radar sensors, the accuracy problem of the traffic control system under low light and weather changes is solved, the optimization of traffic flow and the timely identification of abnormal behavior are achieved, and the accuracy and stability of traffic management are improved.

CN120612819BActive Publication Date: 2025-10-03HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511100141.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-10-03
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing intelligent traffic control systems have poor accuracy in low-light environments and weather changes, image sensor data is easily distorted, and they fail to effectively integrate the differences between image and radar data, leading to misjudgments in traffic management and poor traffic flow control.

Method used

A multimodal data fusion system is used to synchronously collect data through image sensors and millimeter-wave radar sensors, perform adaptive preprocessing, construct a multimodal original data set, use a neural network model to analyze modal differences, generate strategy fusion coefficients, and dynamically adjust traffic control strategies.

Benefits of technology

It improves the accuracy of traffic participant behavior identification, reduces traffic congestion, optimizes traffic flow, dynamically responds to traffic anomalies, improves reaction speed and decision-making accuracy, and can more accurately identify different types of traffic participants, especially in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612819B_ABST
    Figure CN120612819B_ABST
Patent Text Reader

Abstract

The present invention provides a traffic control system based on multimodal data fusion, which belongs to the field of traffic control technology. It includes a multi-source data acquisition and adaptive preprocessing module, a modal difference recognition module, a modal difference modeling and analysis module, and a strategy fusion module. The multi-source data acquisition and adaptive preprocessing module is used to construct a multimodal original data set, the modal difference recognition module is used to construct a difference index sequence, the modal difference modeling and analysis module is used to predict the position difference coefficient, speed difference coefficient, and behavior state difference coefficient of the kth type of traffic target, and the strategy fusion module is used to construct a strategy fusion coefficient. This system first analyzes the position, speed, and behavior state differences of traffic participants, and finally improves the prediction effect of traffic conditions through strategic fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic control, and in particular to a traffic control system based on multimodal data fusion. Background Art

[0002] With the acceleration of urbanization, the increase in traffic volume and the types of traffic participants, traditional traffic control systems face huge challenges. Therefore, to adapt to the development of modernization, a series of automatic control systems will be adopted. Although the current intelligent traffic control system has begun to use multimodal data (images, radar, etc.) to assist decision-making, it still has many shortcomings.

[0003] First, in existing technologies, image sensors have poor accuracy in low-light environments and are easily affected by weather factors, resulting in data distortion. In traffic management, the modal differences between image and radar data are often optimized using traditional manual threshold setting methods, which cannot dynamically adapt to changing traffic conditions and environments.

[0004] Based on the image sensors' capture of traffic, existing traffic control systems fail to fully consider the correlation between position differences, speed differences, and behavioral state differences when making real-time adjustments to traffic conditions. Simply integrating these differences may lead to misjudgment of traffic participants' behavior, thereby affecting signal timing and traffic flow control.

[0005] Therefore, a traffic control system based on multimodal data fusion is proposed to solve this problem. Summary of the Invention

[0006] In order to remedy the above deficiencies, the present invention provides a traffic control system based on multimodal data fusion that overcomes the above technical problems or at least partially solves the above problems.

[0007] The present invention is achieved in that:

[0008] The present invention provides a traffic control system based on multimodal data fusion, comprising a multi-source data acquisition and adaptive preprocessing module, a modal difference recognition module, a modal difference modeling and analysis module, and a strategy fusion module;

[0009] The multi-source data acquisition and adaptive preprocessing module is used to deploy image sensors and millimeter-wave radar sensors at multiple locations in key urban traffic areas to synchronously collect image information and radar motion information of motor vehicles, pedestrians, and non-motor vehicles. It also performs adaptive processing of image frame rate and radar scan rate based on ambient brightness below 300 Lux and weather conditions such as rain, snow, or fog to construct a multimodal raw data set.

[0010] The modal difference recognition module is used to compare and analyze the features recognized by the image sensor and the features recognized by the millimeter wave radar sensor based on the multimodal original data set, and to construct a modal difference indicator sequence;

[0011] The modal difference modeling and analysis module is used to build a neural network model, input the modal difference index sequence into the neural network model for analysis, output the prediction results of traffic abnormal behavior, and predict the k-th traffic target position difference coefficient. , speed difference coefficient and behavioral state variance coefficient ;

[0012] Strategy fusion module, used to calculate the position difference coefficient of the kth type of traffic target , speed difference coefficient and behavioral state variance coefficient The strategy fusion coefficient FDC is constructed and optimized to generate the fourth strategy.

[0013] In a preferred solution, the multi-source data acquisition and adaptive pre-processing module includes a division unit, an image acquisition unit, a radar acquisition unit and a pre-processing unit;

[0014] The division unit divides the urban traffic road into a plurality of roads, and collects the three-dimensional coordinates x, y of the j-th intersection in the i-th road through a two-dimensional coordinate system;

[0015] The image acquisition unit is configured to deploy image sensors and millimeter-wave radar sensors at multiple locations at the j-th intersection on the i-th road, to acquire image modalities, including image information of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road, and to obtain the number of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road and their positions by recognizing graphic texture features;

[0016] The radar acquisition unit is configured to deploy millimeter-wave radar sensors at multiple locations at the j-th intersection on the i-th road, to acquire radar modes, including motion echo signals of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road, and to calculate the relative velocity vectors of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road based on the distance change and Doppler shift of the echoes;

[0017] Construct a multimodal raw dataset based on image and radar modalities;

[0018] The preprocessing unit is used to perform time synchronization and spatial registration processing on the number of motor vehicles, the number of pedestrians, and the number of non-motor vehicles with the relative speeds of motor vehicles, pedestrians, and non-motor vehicles, and to establish a target-level correspondence relationship based on the position matching relationship between motor vehicles, pedestrians, and non-motor vehicles in the image frame and the radar scan. At the same time, a brightness sensor is deployed to collect ambient brightness. When the ambient brightness is lower than 300 Lux, it indicates that the ambient brightness is unqualified, so that there is a risk of insufficient image brightness during the acquisition process of the image sensor and the millimeter-wave radar sensor. In the case of rain, snow, or fog, the image acquisition frame rate and the radar scanning frequency are dynamically adjusted to eliminate redundant or low-confidence data, realize the cleaning, completion, and standardization of multimodal data, and generate a highly consistent and time-space aligned fusion input data stream.

[0019] In a preferred solution, the modality difference recognition module includes a data extraction unit, a difference recognition unit and a difference sequence unit;

[0020] The data extraction unit is used to extract the number of vehicles, the number of pedestrians, and the number of non-motor vehicles at the j-th intersection on the i-th road and the relative speed characteristics of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road based on the multimodal original data set;

[0021] The difference recognition unit is used to generate difference measurement values ​​of each dimension by calculating the difference between the image and radar feature values ​​based on the extracted features, including position error, speed error and behavior judgment difference;

[0022] The difference sequence unit is used to construct a difference index sequence based on the number of vehicles, the number of pedestrians, and the number and speed of non-motor vehicles at the j-th intersection in the i-th road. The sequence contains the difference measurements of the positions, speeds, and behaviors of motor vehicles, pedestrians, and non-motor vehicles within a continuous time window, reflecting the multi-dimensional feature inconsistency of each target.

[0023] In a preferred solution, the modal difference modeling and analysis module includes a modeling unit, a capturing unit and a first evaluation unit;

[0024] The modeling unit is used to build a model using a convolutional neural network, and train and test the convolutional neural network model with a multimodal original data set, and use the trained convolutional neural network model as a motor vehicle, pedestrian and non-motor vehicle modal difference test and evaluation model, and use the intermediate layer output of the equipment operation status model as a motor vehicle, pedestrian and non-motor vehicle modal difference feature vector to identify the motor vehicle, pedestrian and non-motor vehicle modal difference information, and train and test the motor vehicle, pedestrian and non-motor vehicle modal difference feature test and evaluation model through the acquired motor vehicle, pedestrian and non-motor vehicle modal difference information, and use the trained motor vehicle, pedestrian and non-motor vehicle modal difference feature test and evaluation model as data operation prediction to predict the k-th type of traffic target position difference coefficient. , speed difference coefficient and behavioral state variance coefficient .

[0025] In a preferred solution, the k-th traffic target position difference coefficient The specific way to obtain it is:

[0026] S1, extracting the number of vehicles, the number of pedestrians, and the number and positions of non-motor vehicles at the j-th intersection in the i-th road in the multimodal original dataset through the capture unit;

[0027] S2. Based on feature data extracted from the image modality and radar modality, an image texture feature recognition algorithm is used to obtain a two-dimensional spatial position set in the image modality. At the same time, target reflection point information is extracted based on the radar echo signal to obtain a two-dimensional spatial position set in the radar modality.

[0028] S3. Use the nearest neighbor matching method to match the number of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road identified by the image and the radar, and calculate the position error of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road. ;

[0029] S4. Based on the correlation between the number L of vehicles, pedestrians and non-motor vehicles at the j-th intersection in the i-th road and the position error, the mean position difference of the k-th traffic target is obtained by the calculation formula ;

[0030] S5. Calculate the relative difference between the number of similar targets in the image and radar modes for the kth type of traffic targets ;

[0031] S6, based on the mean difference of the k-th traffic target position The k-th type of traffic target relative difference between the image and radar mode for the number of similar targets is associated with the k-th type of traffic target, and the k-th type of traffic target position difference coefficient is calculated by the formula .

[0032] In a preferred solution, the first evaluation unit is used to preset a position difference threshold Q and compare the position difference threshold Q with the position difference coefficient of the k-th type of traffic target. Perform a comparison and generate a first evaluation instruction, including:

[0033] when When Q > Q, it indicates that the image modality and the radar modality are abnormally different in the target position. The first strategy is generated, including reducing the current intersection streetlight duration by 10%-30% to shorten the detention period in the abnormal sensing area and reducing the maximum traffic flow of the relevant lane by 15%;

[0034] when When ≤Q, it means that the difference between the image mode and the radar mode in the target position is normal, and the current signal timing parameters are maintained to maintain the current traffic control rhythm.

[0035] In a preferred embodiment, the speed difference coefficient Specific methods of obtaining

[0036] S7. For the j-th intersection on the i-th road, extract the motion trajectories of motor vehicles, pedestrians, and non-motor vehicles using an image sensor and a millimeter-wave radar sensor, respectively, and calculate their velocity vectors per unit time;

[0037] The image sensor calculates the pixel speed of the target based on the position change between multiple frames of images and estimates the physical speed in combination with camera monitoring;

[0038] S8. Using the nearest neighbor matching method, match the data of the image modality and the radar modality under the same target;

[0039] S9. For the successfully matched targets, extract the k-th traffic target speed vector in the image mode and radar mode respectively. and , and calculate the speed error of the kth traffic target in the image mode and radar mode ;

[0040] The speed vector of the kth traffic target in the image modality is The acquisition method is: by analyzing the pictures continuously collected by the image sensor, identifying the coordinate changes of the traffic targets in the pictures, estimating the speed, and thus obtaining the speed vector ;

[0041] S10. Calculate the average speed difference of traffic targets based on the same target group at the jth intersection on the i-th road 、 and ;

[0042] S11. Weighted summary of the speed difference means of motor vehicles, pedestrians and non-motor vehicles to construct the speed difference coefficient ;

[0043] The second evaluation unit is used to preset a speed difference threshold W and compare the speed difference threshold W with the speed difference coefficient Performing a comparison and generating a second evaluation instruction, including:

[0044] when When the speed difference between the image and radar modalities is greater than W, a second strategy is generated, including increasing the radar sensor sampling frequency at the intersection by 11%-21% to enhance robustness, introducing a dynamic buffer window of 0.5 seconds to 1.2 seconds for traffic participants to avoid short-term misjudgments, and dynamically compensating the traffic light duration at the intersection by ±10% to address traffic fluctuations caused by speed recognition.

[0045] when When ≤W, it means that the speed difference between the image mode and the radar mode for traffic participants is normal, and the current signal timing parameters are maintained to maintain the current traffic control rhythm.

[0046] In a preferred embodiment, the behavioral state difference coefficient The specific way to obtain is:

[0047] S12, identification of traffic participant behavior states based on image and radar modalities, including stationary, moving, turning, and crossing;

[0048] S13. Using the images acquired simultaneously by the image modality and the radar modality, combined with texture features in the images, identify the same traffic target formation, and match the behavior states of the same traffic target in the two modalities to obtain a successfully matched traffic target.

[0049] S14. For successfully matched traffic targets, continue to identify the behavioral state differences of the same traffic target based on the image mode and radar mode, and count the number of target groups with different states. and the number of groups participating in matching traffic targets , the behavioral state difference coefficient is obtained by calculating ;

[0050] The third evaluation unit is used to preset the state difference threshold R and compare the state difference threshold R with the state difference coefficient Perform comparison and generate a third evaluation instruction, including:

[0051] when When the value is greater than R, it indicates that the image modality and radar modality have anomalies in their recognition of the behavior of traffic participants. This generates a third strategy, which includes increasing the image acquisition frame rate of the image sensor in the abnormal target area by 12%-23% and reducing the behavior recognition weight of the radar sensor in the area by 9%-24%, thereby reducing recognition interference and shortening the green light passage time by 7%-13%.

[0052] when When ≤R, it means that the image modality and radar modality are recognizing the behavior status of traffic participants normally, and the current traffic control parameters continue to be maintained and monitoring continues.

[0053] In a preferred solution, the strategy fusion module includes an association unit and an optimization unit;

[0054] The associated unit is used to convert the k-th type of traffic target position difference coefficient , speed difference coefficient and behavioral state variance coefficient The strategy fusion coefficient FDC is obtained by calculation.

[0055] In a preferred solution, the optimization unit is used to preset a strategy fusion threshold Y, and compare the strategy fusion threshold Y with the strategy fusion coefficient FDC to generate a fourth evaluation instruction, including:

[0056] When FDC>Y, it indicates that the current multimodal difference fusion is abnormal, and the fourth strategy is generated. This includes reducing the fusion weight of modal identification sources with significant differences by 12%-25% in the control strategy decision-making, while increasing the fusion weight of modalities with small modal differences by 10%-18%, adjusting the green light duration adjustment range by ±15%, adjusting the pedestrian waiting time by ±5%, and adjusting the non-motor vehicle priority passing efficiency by ±3%;

[0057] When FDC≤Y, it means that the current multimodal difference fusion is normal, and the current traffic control parameters are maintained to ensure normal traffic operation.

[0058] The present invention provides a traffic control system based on multimodal data fusion, which has the following beneficial effects:

[0059] 1. Through real-time analysis and fusion of multimodal data, traffic signals can be precisely controlled to reduce traffic congestion, optimize traffic flow, and improve road efficiency. The system analyzes the differences in the position, speed, and behavior of traffic participants detected by image sensors and radar sensors, and based on these differences, it can predict the behavior of traffic participants and identify abnormal behaviors in advance, such as rapid lane changes and sudden braking, so as to make timely signal adjustments and prevent traffic accidents.

[0060] 2. Based on the analysis of modal difference indicators by the neural network model, the system can dynamically adjust the control strategy, identify and respond to abnormal traffic changes, such as unexpected traffic jams and emergencies, and improve reaction speed and decision-making accuracy. The comparative analysis of image and radar data effectively reduces the error of a single sensor and improves the accuracy of target recognition. Especially in complex traffic environments, it can more accurately identify different types of traffic participants (motor vehicles, pedestrians, non-motor vehicles). Through the strategy fusion module, different coefficients (such as the k-th type of traffic target position difference coefficient, speed difference coefficient, etc.) are optimized. The system can dynamically adjust the green light duration, pedestrian waiting time and non-motor vehicle priority according to real-time traffic conditions, effectively reducing traffic delays and alleviating traffic pressure. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1 It is a system block diagram of the present invention. DETAILED DESCRIPTION

[0063] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0064] Example 1, with reference to Figure 1 ,The present invention provides a technical solution: a traffic control system based on multimodal data fusion, comprising a multi-source data acquisition and adaptive preprocessing module, a modal difference recognition module, a modal difference modeling and analysis module, and a strategy fusion module;

[0065] The multi-source data acquisition and adaptive preprocessing module is used to deploy image sensors and millimeter-wave radar sensors at multiple locations in key urban traffic areas to synchronously collect image information and radar motion information of motor vehicles, pedestrians, and non-motor vehicles. It also performs adaptive processing of image frame rate and radar scan rate based on ambient brightness below 300 Lux and weather conditions such as rain, snow, or fog to construct a multimodal raw data set.

[0066] The modal difference recognition module is used to compare and analyze the features recognized by the image sensor and the features recognized by the millimeter wave radar sensor based on the multimodal original data set, and to construct a modal difference indicator sequence;

[0067] The modal difference modeling and analysis module is used to build a neural network model, input the modal difference index sequence into the neural network model for analysis, output the prediction results of traffic abnormal behavior, and predict the k-th traffic target position difference coefficient. , speed difference coefficient and behavioral state variance coefficient ;

[0068] Strategy fusion module, used to calculate the position difference coefficient of the kth type of traffic target , speed difference coefficient and behavioral state variance coefficient The strategy fusion coefficient FDC is constructed and optimized to generate the fourth strategy.

[0069] In this embodiment, through real-time analysis and fusion of multimodal data, traffic signals can be accurately controlled, traffic congestion can be reduced, traffic flow can be optimized, and road traffic efficiency can be improved. By analyzing the modal differences between image sensors and radar sensors, the system can predict the behavior of traffic participants and identify abnormal behaviors in advance, such as rapid lane changes and sudden braking, so as to make timely signal adjustments and prevent traffic accidents. The system can also automatically adjust the image frame rate and radar scanning rate to cope with different weather conditions and ambient brightness, ensuring that accurate traffic data can still be efficiently obtained under low visibility conditions such as rainy days, nighttime or fog and haze, thereby ensuring the stability and reliability of the system.

[0070] Based on the analysis of modal difference indicators by the neural network model, the system can dynamically adjust the control strategy, identify and respond to abnormal traffic changes, such as unexpected traffic jams and emergencies, and improve reaction speed and decision-making accuracy. The comparative analysis of image and radar data effectively reduces the error of a single sensor and improves the accuracy of target recognition. Especially in complex traffic environments, it can more accurately identify different types of traffic participants (motor vehicles, pedestrians, non-motor vehicles). Through the strategy fusion module, different coefficients (such as the k-th type of traffic target position difference coefficient, speed difference coefficient, etc.) are optimized. The system can dynamically adjust the green light duration, pedestrian waiting time and non-motor vehicle priority according to real-time traffic conditions, effectively reducing traffic delays and alleviating traffic pressure.

[0071] Example 2: This example is an explanation of Example 1. Please refer to Figure 1 ,Specifically, the multi-source data acquisition and adaptive pre-processing module includes a ,division unit, an image acquisition unit, a radar acquisition unit and a ,pre-processing unit;

[0072] The division unit divides the urban traffic road into a plurality of roads, and collects the three-dimensional coordinates x, y of the j-th intersection in the i-th road through a two-dimensional coordinate system;

[0073] The image acquisition unit is configured to deploy image sensors and millimeter-wave radar sensors at multiple locations at the j-th intersection on the i-th road, to acquire image modalities, including image information of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road, and to obtain the number of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road and their positions by recognizing graphic texture features;

[0074] The radar acquisition unit is configured to deploy millimeter-wave radar sensors at multiple locations at the j-th intersection on the i-th road, to acquire radar modes, including motion echo signals of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road, and to calculate the relative velocity vectors of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road based on the distance change and Doppler shift of the echoes;

[0075] Construct a multimodal raw dataset based on image and radar modalities;

[0076] The preprocessing unit is used to perform time synchronization and spatial registration processing on the number of vehicles, the number of pedestrians, and the number of non-motor vehicles with the relative speeds of motor vehicles, pedestrians, and non-motor vehicles, and to construct a target-level correspondence relationship based on the position matching relationship between motor vehicles, pedestrians, and non-motor vehicles in the image frame and the radar scan. At the same time, a brightness sensor is deployed to collect ambient brightness. When the ambient brightness is lower than 300 Lux, it indicates that the ambient brightness is unqualified, so that there is a risk of insufficient image brightness during the acquisition process of the image sensor and the millimeter-wave radar sensor. In the case of rain, snow, or fog, the image acquisition frame rate and the radar scanning frequency are dynamically adjusted to eliminate redundant or low-confidence data, realize the cleaning, completion, and standardization of multimodal data, and generate a highly consistent and time-space aligned fusion input data stream.

[0077] In scenes where the ambient brightness is lower than 300 Lux and the weather is rainy, snowy, or foggy, the image preprocessing uses the following formula to complement the image with the radar:

[0078] Enhance radar signatures using image information;

[0079] ;

[0080] Where, The query matrix representing radar feature generation represents the query request sent by each position in the radar feature map to the image feature map. The key matrix generated for the image features is equivalent to a queryable directory or index provided by the image feature map, which represents the label of the semantic information contained in each spatial position of the image; The value matrix representing the image feature generation represents the actual and rich semantic content contained in the image feature map at each spatial position, where is the dimension of the key vector. The scaling operation is to prevent the dot product knot from being too large and causing the gradient to disappear, thereby stabilizing the training process. It means that the similarity scores are normalized and converted into weight coefficients. The sum of the weights of each radar position corresponding to all image positions is 1. The attention map representing the final output is a weighted sum obtained by applying the normalized weights to the value matrix of the image. The result is a new feature map of the same size as the radar feature map, but the feature vector at each position incorporates the most relevant semantic information from the image.

[0081] Enhanced radar signature to correct image features;

[0082] ;

[0083] Where, Represents the query matrix generated for the original image features, which represents the query request sent by each position in the image feature map to the radar feature map. Represents the key matrix generated by the enhanced radar features. The enhanced radar features are used because they already contain semantic information, which helps to more accurately match the image query. It represents the value matrix generated by the enhanced radar feature, which represents the precise geometric and velocity information provided by the radar after semantic enhancement. The attention map representing the final output corrects and enriches the image features degraded by raw materials such as weather by injecting precise physical information of the radar into the image features according to the correlation weights.

[0084] In this embodiment, by deploying image sensors and millimeter-wave radar sensors at key intersections on urban traffic roads, it is possible to simultaneously obtain information on traffic participants in image and radar modes, thereby achieving all-round, full-factor perception of vehicles, pedestrians, and non-motor vehicles. The image mode provides dense spatial texture features, and the radar mode provides accurate motion state information. Combining the advantages of the two modes can effectively improve perception accuracy and data richness. By collecting ambient lighting, weather conditions, and communication bandwidth status, dynamic adjustment of the image acquisition frame rate and radar scanning frequency can be achieved, ensuring stable and high-quality acquisition of key traffic information even in complex environments.

[0085] By establishing a position matching relationship between vehicles, pedestrians and non-motor vehicles in images and radars, a one-to-one correspondence at the target level is constructed, providing a unified and standard input basis for subsequent modal difference analysis and strategy fusion, eliminating redundant or low-confidence data, completing data standardization processing, and ensuring that the subsequent recognition and analysis stages are based on consistent and reliable multi-modal input data, thereby improving the accuracy of system perception and decision-making from the source.

[0086] Example 3, this example is the explanation in Example 1, please refer to Figure 1 ,Specifically, the modality difference recognition module includes a data extraction unit, a difference ,recognition unit and a difference sequence unit;

[0087] The data extraction unit is used to extract the number of vehicles, the number of pedestrians, and the number of non-motor vehicles at the j-th intersection on the i-th road and the relative speed characteristics of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road based on the multimodal original data set;

[0088] The difference recognition unit is used to generate difference measurement values ​​of each dimension by calculating the difference between the image and radar feature values ​​based on the extracted features, including position error, speed error and behavior judgment difference;

[0089] The difference sequence unit is used to construct a difference index sequence based on the number of vehicles, the number of pedestrians, and the number and speed of non-motor vehicles at the j-th intersection in the i-th road. The sequence contains the difference measurements of the positions, speeds, and behaviors of motor vehicles, pedestrians, and non-motor vehicles within a continuous time window, reflecting the multi-dimensional feature inconsistency of each target.

[0090] In this embodiment, the data extraction unit can simultaneously extract the number and speed characteristics of traffic participants in the image mode and radar mode, providing a complete feature basis for identifying potential differences between multiple modalities. The difference recognition unit performs refined comparative analysis of the image and radar in terms of target position, speed, behavioral status, etc., which can effectively identify perception deviations caused by sensor errors, environmental changes or occlusions, and enhance the system's ability to identify abnormal data or perception conflicts. The difference sequence unit calculates the feature differences within a continuous time window to form a difference indicator sequence that can be used for trajectory trend, behavioral stability and modal consistency analysis, which helps to identify potential problems such as sudden anomalies and long-term perception offsets.

[0091] Before data fusion, multi-modal consistency assessment and difference quantification are performed first, which can effectively avoid misjudgment, redundancy or fuzzy perception caused by direct fusion, and improve the input quality of subsequent fusion analysis and strategy decision-making modules. By introducing difference dimension analysis and indicator sequence tracking, the system can dynamically respond to recognition deviations and interference between different modalities, ensuring that the traffic control system still has high accuracy and stability in high-density, multi-target and multi-obstacle urban traffic environments.

[0092] Example 4: This example is an explanation of Example 1. Please refer to Figure 1 ,Specifically, the modal difference modeling and analysis module includes a ,modeling unit, a capturing unit and a first evaluation unit;

[0093] The modeling unit is used to build a model using a convolutional neural network, and train and test the convolutional neural network model with a multimodal original data set, and use the trained convolutional neural network model as a motor vehicle, pedestrian and non-motor vehicle modal difference test and evaluation model, and use the intermediate layer output of the equipment operation status model as a motor vehicle, pedestrian and non-motor vehicle modal difference feature vector to identify the motor vehicle, pedestrian and non-motor vehicle modal difference information, and train and test the motor vehicle, pedestrian and non-motor vehicle modal difference feature test and evaluation model through the acquired motor vehicle, pedestrian and non-motor vehicle modal difference information, and use the trained motor vehicle, pedestrian and non-motor vehicle modal difference feature test and evaluation model as data operation prediction to predict the k-th type of traffic target position difference coefficient. , speed difference coefficient and behavioral state variance coefficient .

[0094] In this embodiment, a convolutional neural network is used to automatically learn the difference information between image modalities and radar modalities on vehicles, pedestrians, and non-motor vehicle targets, which can efficiently extract and represent the complex nonlinear difference relationship between multiple modalities, enhance the accuracy and robustness of difference recognition, and use the feature output of the middle layer of the convolutional neural network model during training as the representation vector of modal differences to avoid the underlying noise and redundant information in the original input, effectively improving the precision and generalization ability of difference analysis.

[0095] The modal difference feature test and evaluation model obtained through training has the ability to adapt to new data patterns in different time periods and traffic scenarios, and supports the long-term online optimization and adaptive adjustment of the system. This module can not only identify the modal differences of traffic participants, but also accurately capture problems such as recognition distortion, occlusion effects, and modal drift in the perception system, providing a reliable basis for the system's subsequent strategic decision-making.

[0096] Example 5, this example is the explanation in Example 1, please refer to Figure 1 Specifically, the k-th traffic target position difference coefficient The specific way to obtain it is:

[0097] S1, extracting the number of vehicles, the number of pedestrians, and the number and positions of non-motor vehicles at the j-th intersection in the i-th road in the multimodal original dataset through the capture unit;

[0098] S2. Based on feature data extracted from the image modality and radar modality, an image texture feature recognition algorithm is used to obtain a two-dimensional spatial position set in the image modality. At the same time, target reflection point information is extracted based on the radar echo signal to obtain a two-dimensional spatial position set in the radar modality.

[0099] The set of two-dimensional spatial locations of the image modality;

[0100] ;

[0101] The two-dimensional spatial position set of radar modes;

[0102] ;

[0103] Where, Indicates the traffic target type, Represents a motor vehicle, Indicates pedestrians, Indicates non-motor vehicle; They are respectively represented as the coordinate points of the motor vehicle on the x-axis and y-axis in the image modality, They are represented as the coordinate points of the pedestrian on the x-axis and y-axis in the image modality, They are respectively represented as the coordinate points of the non-motor vehicle on the x-axis and y-axis in the image modality, They are respectively represented as the coordinate points of the motor vehicle on the x-axis and y-axis in the radar mode, Represented as the coordinate points of the pedestrian on the x-axis and y-axis in the radar mode, They are respectively represented as the coordinate points of the non-motorized vehicle on the x-axis and y-axis in the radar mode;

[0104] S3. Use the nearest neighbor matching method to match the number of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road identified by the image and the radar, and calculate the position error of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road. ;

[0105] ;

[0106] S4. Based on the correlation between the number L of traffic targets at the j-th intersection in the i-th road and the position error, the mean position difference of the k-th traffic target is obtained by the calculation formula ;

[0107] ;

[0108] Where, is the total number of traffic targets; It is represented by the number of traffic targets, k represents the type of traffic target, where k=1 means the above The corresponding is motor vehicle, k=2, that is, the above Corresponding to pedestrians, k=3 is the above The corresponding ones are non-motorized vehicles, and K represents the total number of traffic target types;

[0109] S5. Calculate the relative difference between the number of similar targets in the image and radar modes for the kth type of traffic targets;

[0110] ;

[0111] Where, Represents the number of targets detected by the image sensor, Represents the number of targets detected by the millimeter-wave radar sensor, It represents the maximum number of targets detected by the image sensor and the millimeter-wave radar sensor;

[0112] S6, based on the mean difference of the k-th traffic target position The k-th type of traffic target relative difference between the image and radar mode for the number of similar targets is associated with the k-th type of traffic target, and the k-th type of traffic target position difference coefficient is calculated by the formula ;

[0113] ;

[0114] Where, and Represents the weight coefficient, satisfying .

[0115] Preset 、 ;

[0116] By analyzing a large amount of historical traffic perception data, statistically analyzing the error variation range of image modalities and radar modalities in different scenarios, and evaluating the actual impact of position error and target recognition quantity differences on control results (such as traffic efficiency and false positive rate) under various traffic conditions, the value of the weight coefficient can be calculated based on statistical historical data.

[0117] The following is the k-th type of traffic target position difference coefficient The sample table is shown below;

[0118] Transportation participation goals…1…2…3…1…2…3;

[0119] Position error (cm)……1.2……1.5……0.8……1.1……1.3……2.0;

[0120] Millimeter wave radar

[0121] Number of targets detected by the sensor (items) ... 19 ... 34 ... 26 ... 23 ... 37 ... 42;

[0122] Image sensor detection

[0123] Number of targets detected

[0124] Quantity (pieces)…18…33…25…21…35…40;

[0125] Number of similar targets

[0126] kth type of traffic

[0127] Target relative difference

[0128] Difference (%)……0.05……0.03……0.04……0.09……0.06……0.05;

[0129] Category K traffic items

[0130] Marker position difference system

[0131] Number...0.74...0.91...0.50...0.70...0.80...1.22.

[0132] Example 6, this example is the explanation in Example 1, please refer to Figure 1 Specifically, the first evaluation unit is used to preset a position difference threshold Q;

[0133] Based on a large amount of historical traffic data, we conducted a statistical analysis of the positioning errors of image sensors and millimeter-wave radar in different traffic scenarios (such as morning and evening rush hours, rain and fog). We extracted the mean and standard deviation of the position error distribution of various traffic participants (motor vehicles, pedestrians, and non-motor vehicles) under normal perception conditions. Through statistical analysis of this data, we obtained a position difference threshold Q of 0.6.

[0134] The position difference threshold Q and the k-th traffic target position difference coefficient are combined Perform a comparison and generate a first evaluation instruction, including:

[0135] when When Q > Q, it indicates that the image modality and the radar modality are abnormally different in the target position. The first strategy is generated, including reducing the current intersection streetlight duration by 10%-30% to shorten the detention period in the abnormal sensing area and reducing the maximum traffic flow of the relevant lane by 15%;

[0136] when When ≤Q, it means that the difference between the image mode and the radar mode in the target position is normal, and the current signal timing parameters are maintained to maintain the current traffic control rhythm.

[0137] Based on the traffic participation target, continue to construct the position difference threshold Q and the k-th traffic target position difference coefficient A comparison example table is shown below;

[0138] Transportation participation goals…1…2…3…1…2…3;

[0139] Category K traffic items

[0140] Marker position difference system

[0141] Number...0.74...0.91...0.50...0.70...0.80...1.22;

[0142] Position difference threshold Q……0.6……0.6……0.6……0.6……0.6……0.6……0.6;

[0143] Evaluation result...abnormal, execute the first strategy...normal...abnormal, execute the first strategy...abnormal, execute the first strategy...abnormal, execute the first strategy...abnormal, execute the first strategy...abnormal, execute the first strategy.

[0144] In this implementation, the target position is jointly matched by the image sensor and the millimeter-wave radar. The nearest neighbor matching method and Euclidean distance calculation are used to accurately measure the spatial error in target recognition between the image modality and the radar modality. The cross-modal mean difference in the position of the k-th traffic target is effectively constructed. The joint calculation method of the mean difference in the position of the k-th traffic target and the relative difference in the number of the k-th traffic target is introduced to improve the overall sensitivity of the k-th traffic target position difference coefficient to the inconsistency in traffic target recognition, providing a reliable fusion metric for downstream evaluation and control strategies.

[0145] By setting an adjustable preset position difference threshold Q, real-time judgment of the perception consistency between image modality and radar modality is achieved, enabling the system to have intelligent self-assessment and self-judgment capabilities, thereby improving the adaptability and robustness of the perception system. When a modal difference abnormality is detected (the position difference coefficient of the kth type of traffic target is greater than Q), the system automatically executes response strategies such as shortening the intersection travel time and reducing the traffic flow, controlling the target's residence period in the abnormal area, thereby quickly intervening in traffic risks caused by potential perception deviations. When the difference coefficient is less than or equal to the threshold Q, the system maintains the current signal timing parameters and traffic control rhythm, does not interfere with the normal intersection traffic process, takes into account both safety and efficiency, and avoids resource waste and interference fluctuations.

[0146] Example 7, this example is the explanation in Example 1, please refer to Figure 1 Specifically, the speed difference coefficient Specific methods of obtaining

[0147] S7. For the j-th intersection on the i-th road, extract the motion trajectories of motor vehicles, pedestrians, and non-motor vehicles using an image sensor and a millimeter-wave radar sensor, respectively, and calculate their velocity vectors per unit time;

[0148] The image sensor calculates the pixel speed of the target based on the position change between multiple frames of images and estimates the physical speed in combination with camera monitoring;

[0149] S8. Using a nearest neighbor matching method, obtain the speed of the same group of traffic targets in the image mode and the radar mode, based on the speed data obtained in step S7, and match the obtained speed data to obtain successfully matched traffic targets;

[0150] S9. After successfully matching the traffic target, extract the k-th traffic target speed vector in the image mode and radar mode respectively. and , and calculate the speed error of the kth traffic target in the image mode and radar mode , ;

[0151] Where, Indicates target categories, including motor vehicles, pedestrians, and non-motor vehicles;

[0152] The speed vector of the kth traffic target in the image modality is The acquisition method is: by analyzing the pictures continuously collected by the image sensor, identifying the coordinate changes of the traffic targets in the pictures, estimating the speed, and thus obtaining the speed vector ;

[0153] The velocity vector of the kth traffic target in the radar mode , which can be directly obtained through radar sensors;

[0154] S10. Calculate the average speed difference of traffic targets based on the same target group at the jth intersection on the i-th road 、 and ;

[0155] ;

[0156] ;

[0157] ;

[0158] Where, Expressed as the total number of motor vehicles, is the total number of pedestrians, Expressed as the total number of non-motor vehicles, Represents a motor vehicle, Indicates pedestrians, Indicates non-motor vehicles, Expressed as the motor vehicle speed error, Pedestrian speed error, Expressed as the speed error of non-motor vehicles;

[0159] S11. Weighted summary of the speed difference means of motor vehicles, pedestrians and non-motor vehicles to construct the speed difference coefficient ;

[0160] ;

[0161] Where, 、 and is the weight coefficient, satisfying ;

[0162] Preset 、 and ;

[0163] Through long-term traffic data statistics, the actual proportion of traffic participants at the j-th intersection on the i-th road is analyzed. The proportions of motor vehicles, pedestrians and non-motor vehicles in the traffic flow are different. The weight coefficient should correspond to the flow proportion of each type of target. According to the collected data, the 、 and ;

[0164] Based on the traffic participation target, continue to build the speed difference coefficient The sample table is shown below;

[0165] Transportation participation goals…1…2…3…1…2…3;

[0166] In image mode

[0167] Target speed vector of the kth traffic category (km / h)……5.2……4.5……3.7……7.1……4.3……3.9;

[0168] Speed ​​of radar mode

[0169] Vector (km / h)……4.8……4.6……3.2……6.9……4.1……3.8;

[0170] Average speed difference (km / h)……0.4……0.1……0.5……0.2……0.2……0.1;

[0171] Speed ​​difference coefficient…0.32…0.09…0.45…0.14…0.15…0.07.

[0172] The second evaluation unit is configured to preset a speed difference threshold W;

[0173] Long-term monitoring and statistical analysis of speed recognition data from multimodal sensors (image sensors and millimeter-wave radar) in real-world traffic environments. By collecting speed variance coefficient data from a large number of intersections at different time periods and under different traffic flow conditions, and analyzing the data, we determined the speed variance threshold within the normal fluctuation range, ultimately achieving a speed variance threshold W of 0.2.

[0174] The speed difference threshold W and the speed difference coefficient Performing a comparison and generating a second evaluation instruction, including:

[0175] when When the speed difference between the image and radar modalities is greater than W, a second strategy is generated, including increasing the radar sensor sampling frequency at the intersection by 11%-21% to enhance robustness, introducing a dynamic buffer window of 0.5 seconds to 1.2 seconds for traffic participants to avoid short-term misjudgments, and dynamically compensating the traffic light duration at the intersection by ±10% to address traffic fluctuations caused by speed recognition.

[0176] when When ≤W, it means that the speed difference between the image mode and the radar mode for traffic participants is normal, and the current signal timing parameters are maintained to maintain the current traffic control rhythm.

[0177] Based on the traffic participation target, continue to construct the speed difference threshold W and speed difference coefficient A comparison example table is shown below;

[0178] Transportation participation goals…1…2…3…1…2…3;

[0179] Speed ​​difference coefficient……0.32……0.09……0.45……0.14……0.15……0.07;

[0180] Speed ​​difference threshold W……0.2……0.2……0.2……0.2……0.2……0.2……0.2;

[0181] Evaluation result...abnormal, execute the second strategy...normal...abnormal, execute the second strategy...normal...normal...normal.

[0182] In this embodiment, the speed vectors of similar traffic targets are extracted and matched by combining image sensors and millimeter-wave radars, achieving consistent calculation of speed features under multi-source data, enhancing the system's ability to understand the dynamic attributes of traffic participants, and introducing a complementary mechanism of image modality pixel-level speed estimation and radar echo speed calculation to address the recognition blind spots of a single modality under specific weather or lighting conditions, achieving robust enhancement of speed extraction accuracy in complex environments. Through multi-target and multi-modal speed difference calculations, a weighted fusion speed difference coefficient is constructed, providing standardized input for subsequent difference evaluation and strategy adjustment, and improving the versatility and scalability of the system structure.

[0183] The second evaluation unit is used to preset the speed difference threshold W, enabling real-time assessment of speed recognition anomalies for various traffic participants (motor vehicles, pedestrians, and non-motor vehicles), thereby proactively discovering potential perception deviations or failure risks. Through a difference metric-driven evaluation and feedback mechanism, the system can quickly respond to sudden changes in traffic flow or modal perception deviations, autonomously adjust perception frequency and control strategies, and enhance the stability and intelligence level of the overall system.

[0184] Example 8, this example is the explanation in Example 1, please refer to Figure 1 Specifically, the behavioral state difference coefficient The specific way to obtain is:

[0185] S12, identification of traffic participant behavior states based on image and radar modalities, including stationary, moving, turning, and crossing;

[0186] S13. Using the images acquired simultaneously by the image modality and the radar modality, combined with texture features in the images, identify the same traffic target formation, and match the behavior states of the same traffic target in the two modalities to obtain a successfully matched traffic target.

[0187] S14. For successfully matched traffic targets, continue to identify the behavioral state differences of the same traffic target based on the image mode and radar mode, and count the number of target groups with different states. and the number of groups participating in matching traffic targets , the behavioral state difference coefficient is obtained by calculating ;

[0188] ;

[0189] behavioral state variance coefficient It is an indicator that measures the consistency of the recognition results of the same traffic target behavior state by the image modality and the radar modality. It reflects the deviation of the two modalities in behavior recognition by counting the ratio of the number of targets with different behavior states recognized by the two modalities to the total number of matched targets.

[0190] For example, in a certain period of time, assuming that the number of traffic target groups that are successfully matched is 100, among which 18 groups of targets have different behavior state recognition results in image mode and radar mode. At this time, the behavior state difference coefficient is The calculation of is;

[0191] ,Comparing 0.18 with the threshold of 0.15, it is found that the recognition of ,surface light traffic targets is abnormal;

[0192] Based on the traffic participation goal, we continue to construct a sample table of behavior state difference coefficient (SDC), as shown below;

[0193] Transportation participation goals…1…2…3…1…2…3;

[0194] Number of targets with different statuses…2…6…4…5…3…1;

[0195] Number of matched traffic participant groups…18…32…20…22…18…15;

[0196] Coefficient of difference in behavioral states…0.11…0.19…0.20…0.23…0.11…0.07.

[0197] The third evaluation unit is configured to preset a state difference threshold R;

[0198] Based on a comprehensive analysis of historical data and real-time monitoring results on the performance of image and radar modalities in behavioral state recognition, we first collected data from a large number of actual traffic flow scenarios, statistically analyzed the consistency indicators of the behavioral states recognized by the two modalities under normal traffic conditions, and calculated the distribution range and mean of their behavioral state difference coefficients. Then, we selected a reasonable state difference threshold R based on the system's tolerance for recognition anomalies and actual traffic control requirements.

[0199] The state difference threshold R and the behavior state difference coefficient Perform comparison and generate a third evaluation instruction, including:

[0200] when When the value is greater than R, it indicates that the image modality and radar modality have anomalies in their recognition of the behavior of traffic participants. This generates a third strategy, which includes increasing the image acquisition frame rate of the image sensor in the abnormal target area by 12%-23% and reducing the behavior recognition weight of the radar sensor in the area by 9%-24%, thereby reducing recognition interference and shortening the green light passage time by 7%-13%.

[0201] when When ≤R, it means that the image modality and radar modality are recognizing the behavior status of traffic participants normally, and the current traffic control parameters continue to be maintained and monitoring continues.

[0202] Based on the traffic participation target, we continue to construct the state difference threshold R and the state difference coefficient A comparison example table is shown below;

[0203] Transportation participation goals…1…2…3…1…2…3;

[0204] Coefficient of variation of behavioral states…0.11…0.19…0.20…0.23…0.11…0.07;

[0205] State difference threshold R……0.15……0.15……0.15……0.15……0.15……0.15……0.15;

[0206] Evaluation result...normal...abnormal, execute the third strategy...abnormal, execute the third strategy...abnormal, execute the third strategy...normal...normal.

[0207] In this embodiment, the behavioral states of traffic participants (such as stationary, moving, turning, and crossing) are collaboratively identified through image modality and radar modality, and a cross-modal state matching mechanism is constructed to effectively improve the system's accurate understanding of changes in traffic target behavior. By comparing the behavioral states of the same traffic target in image and radar modalities, the number of inconsistent behavioral states and the total number of matches are calculated to form a standardized behavioral state difference coefficient, which helps to unify judgments and subsequent dynamic adjustments. The system can automatically identify behavioral perception anomalies based on the state difference threshold R, maintain high-precision state recognition in multi-target, high-traffic environments, and improve the stability and robustness of the perception system.

[0208] The system's difference recognition and dynamic adjustment mechanism has the capabilities of online judgment, graded response and automatic feedback, effectively improving the intelligence level of the multimodal data fusion system in complex urban traffic scenarios.

[0209] Example 9, this example is the explanation in Example 1, please refer to Figure 1 ,Specifically, the strategy fusion module includes an association unit and an ,optimization unit;

[0210] The associated unit is used to convert the k-th type of traffic target position difference coefficient , speed difference coefficient and behavioral state variance coefficient After dimensionless processing, the strategy fusion coefficient FDC is calculated by the following formula;

[0211] ;

[0212] Where, 、 and are weight coefficients respectively, satisfying .

[0213] Preset 、 and ;

[0214] Based on historical data and analysis of historical data, position differences directly affect the accurate judgment of the spatial distribution of traffic targets and are the basis for adjusting traffic signal timing. Therefore, based on past historical data and experience, position differences will be given a higher weight. Speed ​​differences reflect the dynamic changes of traffic flow and have a greater impact on traffic smoothness and safety, so they are given a medium weight. Although behavioral state differences are important, they are usually used as auxiliary judgment indicators and are given a relatively low weight. Therefore, it is concluded that 、 and ;

[0215] Based on the traffic participation goal, we continue to construct a sample table of the strategy fusion coefficient FDC, as shown below;

[0216] Transportation participation goals…1…2…3…1…2…3;

[0217] Category K traffic items

[0218] Coefficient of difference of target position……0.74……0.91……0.50……0.70……0.80……1.22;

[0219] Speed ​​difference coefficient……0.32……0.09……0.45……0.14……0.15……0.07;

[0220] Coefficient of variation of behavioral states…0.11…0.19…0.20…0.23…0.11…0.07;

[0221] Strategy fusion coefficient…0.49…0.52…0.43…0.44…0.47…0.65.

[0222] In this embodiment, by correlating the position difference coefficient, speed difference coefficient, and behavior state difference coefficient of the kth type of traffic target, the strategy fusion module can effectively integrate the difference information generated by multi-source sensors in the target recognition process, solve the problem of single modal error affecting the overall judgment, and use a dimensionless processing method to eliminate the scale influence between the three difference dimensions, so that various difference indicators have a unified comparison basis, improve the accuracy and applicability of the fusion coefficient FDC, and facilitate the fusion of cross-modal data in the same evaluation system. By setting adjustable weight coefficients (such as 、 and ), the system can flexibly adjust the importance of the three difference factors according to the actual road scene, time period or weather conditions, so that the strategy fusion results are closer to the actual perception needs and enhance the model's adaptability.

[0223] The fusion coefficient FDC serves as a unified indicator that dynamically reflects the degree of collaborative performance of the current multimodal perception system. A higher FDC indicates a greater modal difference, which can be used to drive subsequent optimization control measures. As the core judgment indicator in the fourth evaluation instruction, FDC can uniformly guide the system to adopt strategies such as adjusting the sensor sampling rate, changing the timing of traffic lights, or switching behavior recognition strategies to achieve a logical closed loop of system response.

[0224] Example 10: This example is an explanation of Example 1. Please refer to Figure 1 ,Specifically, the optimization unit is used to preset a strategy fusion threshold Y;

[0225] First, using a large amount of collected multimodal raw data (including image and radar modalities), the statistical distribution of the strategy fusion coefficient (FDC) was calculated for multiple time periods and intersections. The typical value range and fluctuation characteristics of FDC under normal traffic conditions were analyzed. Second, the FDC values ​​corresponding to actual abnormal traffic events (such as sensor failures, abnormal traffic congestion, and signal timing failures) were annotated to clarify the FDC threshold range under abnormal conditions. Then, using statistical methods (such as confidence interval analysis and cluster analysis), a strategy fusion threshold Y of 0.45 was determined to effectively distinguish between normal and abnormal conditions.

[0226] The strategy fusion threshold Y is compared with the strategy fusion coefficient FDC to generate a fourth evaluation instruction, including:

[0227] When FDC>Y, it indicates that the current multimodal difference fusion is abnormal, and the fourth strategy is generated. This includes reducing the fusion weight of modal identification sources with significant differences by 12%-25% in the control strategy decision-making, while increasing the fusion weight of modalities with small modal differences by 10%-18%, adjusting the green light duration adjustment range by ±15%, adjusting the pedestrian waiting time by ±5%, and adjusting the non-motor vehicle priority passing efficiency by ±3%;

[0228] When FDC≤Y, it means that the current multimodal difference fusion is normal, and the current traffic control parameters are maintained to ensure normal traffic operation.

[0229] Based on the traffic participation goal, we continue to build a table comparing the strategy fusion threshold Y and the strategy fusion coefficient FDC, as shown below;

[0230] Transportation participation goals…1…2…3…1…2…3;

[0231] Strategy fusion coefficient…0.49…0.52…0.43…0.44…0.47…0.65;

[0232] Strategy fusion threshold…0.45…0.45…0.45…0.45…0.45…0.45;

[0233] Evaluation result...abnormal, execute the fourth strategy...abnormal, execute the fourth strategy...normal...normal...abnormal, execute the fourth strategy...abnormal, execute the fourth strategy.

[0234] In this embodiment, by setting a strategy fusion threshold Y and comparing it with the fusion coefficient FDC, the system can accurately determine the overall difference between the current image modality and the radar modality. When an anomaly occurs (FDC>Y), the adaptive fourth strategy is triggered in a timely manner to effectively avoid the miscontrol problem caused by fusion failure. For identification sources with significant modal differences, the system automatically reduces their fusion weight in control decisions (by 12%-25%), while increasing the contribution weight of modalities with small differences (by 10%-18%), achieving intelligent redistribution of the contributions of different modalities and enhancing the accuracy and robustness of fusion decisions. When a modal fusion anomaly is detected, the system can synchronously adjust traffic control parameters, such as dynamically adjusting the green light duration by ±15%, to balance traffic pressure in different directions and optimize overall traffic efficiency.

[0235] The fourth strategy supports fine-grained adjustments: flexible adjustments of ±5% are made to pedestrian waiting time to reduce their waiting anxiety or congestion risks; and ±3% fine-tuning is implemented for the priority passage efficiency of non-motor vehicles to ensure smooth, safe and orderly traffic.

[0236] When FDC≤Y, the system maintains the stable operation of the existing control strategy, avoiding traffic fluctuations or interference caused by invalid adjustments, and improving the robustness and stability of the intelligent transportation system.

[0237] The threshold is set to facilitate comparison. The size of the threshold depends on the amount of sample data and the number of bases set by technicians in this field for each set of sample data; as long as it does not affect the proportional relationship between the parameter and the quantized value.

[0238] The above formulas are obtained by collecting a large amount of data and performing software simulation, and a formula close to the actual value is selected. The coefficients in the formula are set by those skilled in the art according to actual conditions. The above is only a preferred specific implementation method of the present invention, but the protection scope of the present invention is not limited to this. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, can make equivalent replacements or changes based on the technical solution and inventive concept of the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A traffic control system based on multimodal data fusion, characterized in that: It includes multi-source data acquisition and adaptive preprocessing module, modal difference identification module, modal difference modeling and analysis module and strategy fusion module; The multi-source data acquisition and adaptive preprocessing module is used to deploy image sensors and millimeter-wave radar sensors at multiple locations in key urban traffic areas to synchronously collect image information and radar motion information of motor vehicles, pedestrians, and non-motor vehicles to construct a multimodal raw data set; The modal difference recognition module is used to compare and analyze the features recognized by the image sensor and the features recognized by the millimeter wave radar sensor based on the multimodal original data set, and to construct a modal difference indicator sequence; The modal difference modeling and analysis module is used to build a neural network model, input the modal difference index sequence into the neural network model for analysis, output the prediction results of traffic abnormal behavior, and predict the k-th traffic target position difference coefficient. , speed difference coefficient and behavioral state variance coefficient ; Strategy fusion module, used to calculate the position difference coefficient of the kth type of traffic target , speed difference coefficient and behavioral state variance coefficient The strategy fusion coefficient FDC is constructed and optimized to generate the fourth strategy.

2. A traffic control system based on multimodal data fusion according to claim 1, characterized in that: The multi-source data acquisition and adaptive pre-processing module includes a division unit, an image acquisition unit, a radar acquisition unit and a pre-processing unit; The division unit divides the urban traffic road into a plurality of roads, and collects the three-dimensional coordinates x, y of the j-th intersection in the i-th road through a two-dimensional coordinate system; The image acquisition unit is configured to deploy image sensors and millimeter-wave radar sensors at multiple locations at the j-th intersection on the i-th road, to acquire image modalities, including image information of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road, and to obtain the number of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road and their positions by recognizing graphic texture features; The radar acquisition unit is configured to deploy millimeter-wave radar sensors at multiple locations at the j-th intersection on the i-th road, to acquire radar modes, including motion echo signals of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road, and to calculate the relative velocity vectors of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road based on the distance change and Doppler shift of the echoes; Construct a multimodal raw dataset based on image and radar modalities; The preprocessing unit is used to perform time synchronization and spatial registration processing on the number of motor vehicles, the number of pedestrians, and the number of non-motor vehicles with the relative speeds of motor vehicles, pedestrians, and non-motor vehicles, and to establish a target-level correspondence relationship based on the position matching relationship between motor vehicles, pedestrians, and non-motor vehicles in the image frame and the radar scan. At the same time, a brightness sensor is deployed to collect ambient brightness. When the ambient brightness is lower than 300 Lux, it indicates that the ambient brightness is unqualified, so that there is a risk of insufficient image brightness during the acquisition process of the image sensor and the millimeter-wave radar sensor. In the case of rain, snow, or fog, the image acquisition frame rate and the radar scanning frequency are dynamically adjusted to eliminate redundant or low-confidence data, realize the cleaning, completion, and standardization of multimodal data, and generate a highly consistent and time-space aligned fusion input data stream.

3. A traffic control system based on multimodal data fusion according to claim 2, characterized in that: The modality difference recognition module includes a data extraction unit, a difference recognition unit and a difference sequence unit; The data extraction unit is used to extract the number of vehicles, the number of pedestrians, and the number of non-motor vehicles at the j-th intersection on the i-th road and the relative speed characteristics of motor vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road based on the multimodal original data set; The difference recognition unit is used to generate difference measurement values ​​of each dimension by calculating the difference between the image and radar feature values ​​based on the extracted features, including position error, speed error and behavior judgment difference; The difference sequence unit is used to construct a difference index sequence based on the number of vehicles, the number of pedestrians, and the number and speed of non-motor vehicles at the j-th intersection in the i-th road. The sequence contains the difference measurements of the positions, speeds, and behaviors of motor vehicles, pedestrians, and non-motor vehicles within a continuous time window, reflecting the multi-dimensional feature inconsistency of each target.

4. A traffic control system based on multimodal data fusion according to claim 3, characterized in that: The modal difference modeling and analysis module includes a modeling unit, a capture unit, a first evaluation unit, a second evaluation unit, and a third evaluation unit; The modeling unit is used to build a model using a convolutional neural network, and train and test the convolutional neural network model with a multimodal original data set, and use the trained convolutional neural network model as a motor vehicle, pedestrian and non-motor vehicle modal difference test and evaluation model, and use the intermediate layer output of the equipment operation status model as a motor vehicle, pedestrian and non-motor vehicle modal difference feature vector to identify the motor vehicle, pedestrian and non-motor vehicle modal difference information, and train and test the motor vehicle, pedestrian and non-motor vehicle modal difference feature test and evaluation model through the acquired motor vehicle, pedestrian and non-motor vehicle modal difference information, and use the trained motor vehicle, pedestrian and non-motor vehicle modal difference feature test and evaluation model as data operation prediction to predict the k-th type of traffic target position difference coefficient. , speed difference coefficient and behavioral state variance coefficient .

5. The traffic control system based on multimodal data fusion according to claim 4, characterized in that: The k-th type of traffic target position difference coefficient The specific way to obtain it is: S1, extracting the number of vehicles, the number of pedestrians, and the number and positions of non-motor vehicles at the j-th intersection in the i-th road in the multimodal original dataset through the capture unit; S2. Based on feature data extracted from the image modality and radar modality, an image texture feature recognition algorithm is used to obtain a two-dimensional spatial position set in the image modality. At the same time, target reflection point information is extracted based on the radar echo signal to obtain a two-dimensional spatial position set in the radar modality. S3. Use the nearest neighbor matching method to match the number of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road identified by the image and the radar, and calculate the position error of vehicles, pedestrians, and non-motor vehicles at the j-th intersection on the i-th road. ; S4. Based on the correlation between the number L of vehicles, pedestrians and non-motor vehicles at the j-th intersection in the i-th road and the position error, the mean position difference of the k-th traffic target is obtained by the calculation formula ; S5. Calculate the relative difference between the number of similar targets in the image and radar modes for the kth type of traffic targets ; S6, based on the mean difference of the k-th traffic target position The k-th type of traffic target relative difference between the image and radar mode for the number of similar targets is associated with the k-th type of traffic target, and the k-th type of traffic target position difference coefficient is obtained by calculation .

6. The traffic control system based on multimodal data fusion according to claim 5, characterized in that: The first evaluation unit is used to preset a position difference threshold Q and compare the position difference threshold Q with the k-th type of traffic target position difference coefficient Perform a comparison and generate a first evaluation instruction, including: when When Q > Q, it indicates that the image modality and the radar modality are significantly different in the target position, and the first strategy is generated, including reducing the current intersection streetlight duration by 10%-30% and reducing the maximum traffic flow of the relevant lane by 15%; when When ≤Q, it means that the difference between the image mode and the radar mode in the target position is normal, and the current signal timing parameters are maintained to maintain the current traffic control rhythm.

7. The traffic control system based on multimodal data fusion according to claim 6, characterized in that: The speed difference coefficient Specific methods of obtaining S7. For the j-th intersection on the i-th road, extract the motion trajectories of motor vehicles, pedestrians, and non-motor vehicles using an image sensor and a millimeter-wave radar sensor, respectively, and calculate their velocity vectors per unit time; The image sensor calculates the pixel speed of the target based on the position change between multiple frames of images and estimates the physical speed in combination with camera monitoring; S8. Using the nearest neighbor matching method, match the data of the image modality and the radar modality under the same target; S9. For the successfully matched targets, extract the k-th traffic target speed vector in the image mode and radar mode respectively. and , and calculate the speed error of the kth traffic target in the image mode and radar mode , ; Where, Indicates target categories, including motor vehicles, pedestrians, and non-motor vehicles; The speed vector of the kth traffic target in the image modality is The acquisition method is: by analyzing the pictures continuously collected by the image sensor, identifying the coordinate changes of the traffic targets in the pictures, estimating the speed, and thus obtaining the speed vector ; The velocity vector of the kth traffic target in the radar mode , directly obtained through radar sensors; S10. Calculate the average speed difference of traffic targets based on the same target group at the jth intersection on the i-th road 、 and ; ; ; ; Where, Expressed as the total number of motor vehicles, is the total number of pedestrians, Expressed as the total number of non-motor vehicles, Represents a motor vehicle, Indicates pedestrians, Indicates non-motor vehicles, Expressed as the motor vehicle speed error, Pedestrian speed error, Expressed as the speed error of non-motor vehicles; S11. Weighted summary of the speed difference means of motor vehicles, pedestrians and non-motor vehicles to construct the speed difference coefficient ; ; Where, 、 and is the weight coefficient, satisfying ; The second evaluation unit is used to preset a speed difference threshold W and compare the speed difference threshold W with the speed difference coefficient Performing a comparison and generating a second evaluation instruction, including: when When ≥W, it indicates that the speed difference between the image modality and the radar modality of the traffic participants is abnormal, and the second strategy is generated, including increasing the radar sensor sampling frequency of the intersection by 11%-21%, introducing a dynamic buffer window of 0.5 seconds to 1.2 seconds for traffic participants, and performing a dynamic compensation of ±10% on the duration of the traffic light at the intersection; when When ≤W, it means that the speed difference between the image mode and the radar mode for traffic participants is normal, and the current signal timing parameters are maintained to maintain the current traffic control rhythm.

8. The traffic control system based on multimodal data fusion according to claim 7, characterized in that: The behavioral state difference coefficient The specific way to obtain is: S12, identification of traffic participant behavior states based on image and radar modalities, including stationary, moving, turning, and crossing; S13. Using the images acquired simultaneously by the image modality and the radar modality, combined with texture features in the images, identify the same traffic target formation, and match the behavior states of the same traffic target in the two modalities to obtain a successfully matched traffic target. S14. For successfully matched traffic targets, continue to identify the behavioral state differences of the same traffic target based on the image mode and radar mode, and count the number of target groups with different states. and the number of groups participating in matching traffic targets , the behavioral state difference coefficient is obtained by calculating ; The third evaluation unit is used to preset the state difference threshold R and compare the state difference threshold R with the state difference coefficient Perform comparison and generate a third evaluation instruction, including: when When the value is greater than R, it indicates that the image modality and radar modality have detected anomalies in the recognition of traffic participants' behavior states, generating a third strategy, including: increasing the image acquisition frame rate of the image sensor in the abnormal target area by 12%-23%, reducing the behavior recognition weight of the radar sensor in the area by 9%-24%, and shortening the green light passage time by 7%-13%; when When ≤R, it means that the image modality and radar modality are recognizing the behavior status of traffic participants normally, and the current traffic control parameters continue to be maintained and monitoring continues.

9. The traffic control system based on multimodal data fusion according to claim 8, characterized in that: The strategy fusion module includes an association unit and an optimization unit; The associated unit is used to convert the k-th type of traffic target position difference coefficient , speed difference coefficient and behavioral state variance coefficient The strategy fusion coefficient FDC is obtained by calculation.

10. The traffic control system based on multimodal data fusion according to claim 9, characterized in that: The optimization unit is configured to preset a strategy fusion threshold Y, and compare the strategy fusion threshold Y with the strategy fusion coefficient FDC to generate a fourth evaluation instruction, including: When FDC>Y, it indicates that the current multimodal difference fusion is abnormal, and the fourth strategy is generated. This includes reducing the fusion weight of modal identification sources with significant differences by 12%-25% in the control strategy decision-making, while increasing the fusion weight of modalities with small modal differences by 10%-18%, adjusting the green light duration adjustment range by ±15%, adjusting the pedestrian waiting time by ±5%, and adjusting the non-motor vehicle priority passing efficiency by ±3%; When FDC≤Y, it means that the current multimodal difference fusion is normal, and the current traffic control parameters are maintained to ensure normal traffic operation.

Citation Information

Patent Citations

  • Full-dose full-sample real-time traffic data-based multi-parameter fusion method and system

    CN110807924A

  • Intelligent traffic management method and system based on multi-modal perception

    CN117275224A