Smart city multi-modal data acquisition and fusion method and system

By using multimodal sensors for collaborative data acquisition and fusion decision-making, the problems of detection accuracy and response delay in low-visibility environments of smart city traffic management systems have been solved, enabling efficient traffic accident identification and rescue, and reducing traffic congestion and delays.

CN121768191APending Publication Date: 2026-03-31CHANGZHOU INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing smart city traffic management systems suffer from decreased detection accuracy in low visibility environments, insufficient multi-source data fusion capabilities, and a lack of closed-loop response processes, leading to false alarms, missed alarms, and delays in rescue responses, which exacerbate traffic congestion.

Method used

It uses visual sensors, geomagnetic sensors, acoustic sensors and millimeter-wave radar to collect data in a coordinated manner, extracts multimodal features through timestamp alignment and spatial registration, makes decisions using multimodal fusion neural networks, and realizes automatic response and closed-loop control, automatically triggering traffic signal adjustment and rescue dispatch.

Benefits of technology

Effectively identifying accident information in low-visibility environments improves the accuracy of judgment, shortens rescue response time, reduces traffic congestion, achieves deep coordination between traffic signals and rescue, and improves emergency response efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121768191A_ABST
    Figure CN121768191A_ABST
Patent Text Reader

Abstract

The invention discloses a smart city multi-modal data acquisition and fusion method and system, and relates to the technical field of city management methods. The smart city multi-modal data acquisition and fusion method comprises four steps of multi-source sensing and asynchronous acquisition, cross-modal feature extraction and alignment, multi-level linkage fusion decision and automatic response and closed-loop control. According to the smart city multi-modal data acquisition and fusion method, data is acquired asynchronously through a multi-source sensor, and it is ensured that multiple dimensions of a traffic scene are covered. Characteristic extraction and alignment ensure the consistency of different modal data in time and space. And the fusion decision outputs a fusion confidence coefficient through a pre-trained neural network model, and is used for judging whether a traffic accident occurs or not. And the automatic response mechanism triggers traffic control and rescue scheduling according to a decision result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of urban management methods, specifically to a method and system for multimodal data acquisition and fusion in smart cities. Background Technology

[0002] With the acceleration of urbanization and the surge in motor vehicle ownership, smart city traffic management needs to achieve rapid accident identification, efficient emergency response, and minimal traffic impact. In traffic scenarios, the timeliness of accident detection and subsequent handling directly affects personal safety, road traffic efficiency, and the city's emergency management level.

[0003] To achieve automated accident detection and dispatch of rescue forces, existing technologies typically utilize intersection cameras to record video, geomagnetic coils to detect vehicle queue lengths, and video analytics algorithms to independently analyze the video stream and attempt to identify accidents. If the video algorithm triggers an alarm, manual confirmation is required from the backend. After confirmation, traffic lights are manually adjusted or traffic police and ambulances are notified. However, existing traffic incident monitoring and response technologies have the following shortcomings:

[0004] First, single-modal sensing has strong limitations. Most systems rely on visual sensors (cameras), but their detection accuracy drops significantly in low-visibility environments such as heavy fog, heavy rain, and nighttime, making it impossible to effectively distinguish between normal traffic fluctuations and real accidents.

[0005] Second, the ability to fuse multi-source data is insufficient. Even when using multiple types of sensors, the data is often fragmented due to the lack of timestamp alignment and spatial registration. Furthermore, decision-making often relies on fixed rules and cannot dynamically weigh the effectiveness of different modal data, which can easily lead to false alarms or missed alarms.

[0006] Third, the response process lacks a closed loop; after accident detection, manual intervention is required to trigger traffic signal adjustments and rescue dispatch, leading to delays in rescue response and exacerbated traffic congestion. To address the shortcomings of existing technologies, this invention provides a method and system for multimodal data acquisition and fusion in smart cities to solve the above problems. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a method and system for multimodal data acquisition and fusion in smart cities. Through a comprehensive design encompassing multi-source sensing and asynchronous acquisition, cross-modal feature extraction and alignment, multi-level linkage fusion decision-making, and automatic response and closed-loop control, the system employs: First, it utilizes visual sensors, geomagnetic sensors, acoustic sensors, and millimeter-wave radar for collaborative acquisition, avoiding the poor adaptability of single-modal environments and effectively identifying accident information. Second, the asynchronous acquisition mode avoids the system delays of synchronous acquisition while covering multi-dimensional information such as visual, electromagnetic, acoustic, and motion trajectory data from traffic scenarios. Effective extraction of cross-modal feature vectors is achieved through timestamp alignment and spatial registration, followed by outputting fusion confidence scores via a pre-trained multimodal fusion neural network, improving the accuracy of accident determination. Third, it automatically triggers traffic signal control and rescue dispatch commands without manual confirmation, resulting in rapid rescue response and reduced traffic congestion.

[0008] To achieve the above objectives, the present invention provides a method for multimodal data acquisition and fusion in smart cities, comprising the following steps:

[0009] S1. Multi-source sensing and asynchronous acquisition: Asynchronously acquire multimodal raw data of traffic scenarios through visual sensors, geomagnetic sensors, acoustic sensors and millimeter-wave radar;

[0010] S2. Cross-modal feature extraction and alignment: The original multimodal data is time-stamp aligned and spatially registered, and feature vectors for each modality are extracted, including visual feature vector V, geomagnetic feature vector M, acoustic feature vector A, and radar trajectory feature vector R;

[0011] S3. Multi-level linkage fusion decision: The feature vector is input into a pre-trained multimodal fusion neural network model, and the fusion confidence score is output. ;when If a traffic accident is detected, an event tag is generated; otherwise, monitoring continues.

[0012] S4. Automatic Response and Closed-Loop Control: Based on event tags, automatically trigger traffic signal control strategies and emergency dispatch instructions; wherein, the fusion confidence level... Calculated using the following formula:

[0013]

[0014] Where f(V,M) is the visual and geomagnetic co-discrimination function, g(A) is the acoustic anomaly scoring function, h(R) is the radar trajectory conflict detection function, α, β, γ are the weighting coefficients of each mode, and α+β+γ=1. This is a preset threshold.

[0015] Preferably, in step S1, the acoustic sensor is a microphone array used to capture collision sounds and abnormal horns; the millimeter-wave radar is used to detect vehicle speed and trajectory in low-visibility environments.

[0016] Preferably, in step S2, the feature extraction includes:

[0017] The YOLOv7 model was used to detect the vehicle's attitude and motion state based on visual data.

[0018] Short-time Fourier transform is used to extract frequency domain features from acoustic data;

[0019] Kalman filtering is used to predict the trajectory and identify collision points on radar data.

[0020] Preferably, in step S3, the multimodal fusion neural network model is a Transformer structure based on an attention mechanism, used to dynamically weight the importance of each modality feature.

[0021] Preferably, the fusion confidence level The calculation also includes a correction term:

[0022]

[0023] Where t is the duration of the event, T is the maximum allowable response time, and δ is the time decay coefficient.

[0024] Preferably, in step S4, the traffic signal control strategy is to dynamically adjust the signal timing according to the real-time traffic flow status, including:

[0025] Calculate the queue length Li and flow rate Qi in each direction;

[0026] With the goal of minimizing the total waiting time, the green light duration is optimized using the following formula:

[0027]

[0028] Where C is the signal period duration and n is the number of phases.

[0029] Preferably, in the event of an accident, the traffic lights are forcibly switched to "red wave control" mode to prioritize the flow of traffic heading towards the rescue direction.

[0030] Preferably, in step S4, the rescue dispatch instruction includes automatically generating a structured accident report and pushing it to the traffic police and emergency medical service platform, and planning the optimal rescue route based on real-time traffic conditions; the specific steps for planning the optimal rescue route based on real-time traffic conditions include:

[0031] S41. Construct a real-time traffic map network: using road nodes as vertices and road segments as edges, the edge weight W is dynamically calculated based on real-time traffic flow data, using the following formula:

[0032]

[0033] Where L is the road segment length, v is the average speed of the road segment, and ρ is the vehicle density;

[0034] S42. Multi-objective path planning: Starting from the current location of the rescue vehicle and ending at the accident site, the improved Dijkstra algorithm is used to solve for the optimal path. The objective function is to minimize the weighted sum of the total travel time T and the path fluctuation coefficient σ: min(αT+βσ), where α and β are weight coefficients, and σ is the weighted average of the speed variance of each segment in the path.

[0035] S43. Dynamic Route Update: Based on real-time traffic flow data, the optimal route is periodically recalculated. If the estimated time of the new route is reduced by more than a threshold Δt compared to the original route, a route update command is sent to the rescue vehicle.

[0036] Preferably, the method further includes a preliminary estimation of the accident severity level based on acoustic data and radar impact intensity data, specifically including the following steps:

[0037] S91. Feature Extraction and Quantization: Extract the peak sound pressure level SPL_max and duration T_s of the collision event from the acoustic data; extract the speed change Δv of the vehicles involved in the accident from the radar data, where Δv is based on the speed difference before and after the impact;

[0038] S92. Severity Index Calculation: The accident severity index S is calculated using a linear weighted model: S = w1 × normalize(SPL_max) + w2 × normalize(T_s) + w3 × normalize(Δv), where w1, w2, and w3 are preset weights, and normalize is a normalization function;

[0039] S93. Severity classification and priority mapping: The severity index S is mapped to a preset severity level, which is Level 1: extremely serious; Level 2: serious; Level 3: general. The system automatically allocates rescue priorities according to the final classification level. Level 1 priority triggers the highest response level and dispatches all available rescue resources. Level 2 priority dispatches resources within a specified range. Level 3 priority dispatches basic rescue units.

[0040] The second aspect of this invention discloses a smart city multimodal data acquisition and fusion system for implementing the aforementioned smart city multimodal data acquisition and fusion method, comprising:

[0041] The multimodal sensing module is used to asynchronously collect multimodal raw data of traffic scenarios through visual sensors, geomagnetic sensors, acoustic sensors and millimeter-wave radar;

[0042] The data preprocessing module is used to perform timestamp alignment and spatial registration on the multimodal raw data, and extract feature vectors for each modality, including visual feature vector V, geomagnetic feature vector M, acoustic feature vector A, and radar trajectory feature vector R.

[0043] The fusion decision module is used to input the feature vector into a pre-trained multimodal fusion neural network model and output the fusion confidence score. ;when If a traffic accident is detected, an event tag is generated; otherwise, monitoring continues.

[0044] The control execution module is used to automatically trigger traffic signal control strategies and rescue dispatch instructions based on event tags;

[0045] The communication module is used for data exchange with the traffic control center and rescue platform.

[0046] The second aspect of the present invention discloses a

[0047] The technical effects and advantages of this invention are as follows:

[0048] 1. This smart city multimodal data acquisition and fusion method, through a full-process design of multi-source sensing and asynchronous acquisition, cross-modal feature extraction and alignment, multi-level linkage fusion decision-making, and automatic response and closed-loop control, firstly, adopts visual sensors, geomagnetic sensors, acoustic sensors, and millimeter-wave radar for collaborative acquisition, avoiding the poor adaptability of single-modal environments and effectively identifying accident information. Secondly, the asynchronous acquisition mode avoids the system delay of synchronous acquisition and can cover multi-dimensional information such as visual, electromagnetic, acoustic, and motion trajectory information of traffic scenarios. It achieves effective extraction of cross-modal feature vectors through timestamp alignment and spatial registration, and then outputs fusion confidence through pre-trained multimodal fusion neural network to improve the accuracy of accident judgment. Thirdly, it automatically triggers traffic signal control and rescue dispatch instructions without manual confirmation, resulting in fast rescue response and reduced traffic congestion.

[0049] 2. This smart city multimodal data acquisition and fusion method significantly improves the system's robustness to complex traffic environments through targeted sensor selection and feature extraction techniques: the acoustic sensor uses a microphone array to accurately capture transient acoustic signals such as collision sounds and abnormal horns, providing "sound evidence" for accident determination; the millimeter-wave radar is specifically designed for low-visibility environments, stably detecting vehicle speed and trajectory, compensating for the shortcomings of visual sensors in adverse weather conditions. Visual data uses the YOLOv7 model to achieve high-precision detection of vehicle attitude and motion, acoustic data extracts frequency domain features through short-time Fourier transform, and radar data uses Kalman filtering to complete trajectory prediction and collision point identification. The collaboration of multiple technologies ensures that different environments and different types of accidents can be effectively perceived.

[0050] 3. This smart city multimodal data acquisition and fusion method achieves deep collaboration between traffic signal control and emergency dispatch, minimizing the impact of accidents on traffic and shortening rescue time: At the traffic control level, by calculating queue lengths and flow rates in each direction, the green light duration is optimized with the goal of "minimizing total waiting time." When an accident occurs, the "red wave control" mode is forcibly switched to prioritize the passage of emergency vehicles. At the emergency dispatch level, structured accident reports are automatically generated and pushed to traffic police and emergency medical service platforms. Road segment weights are dynamically calculated, and the optimal emergency route is planned and updated periodically using an improved Dijkstra algorithm. At the same time, acoustic and radar data are combined to assess the severity of the accident, achieving precise matching of emergency resources. Attached Figure Description

[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0052] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0053] Figure 2 This is a flowchart of the multi-source sensing and asynchronous acquisition process of the present invention;

[0054] Figure 3 This is a flowchart of the cross-modal feature extraction and alignment process of the present invention;

[0055] Figure 4 This is a flowchart illustrating the multi-level linkage and fusion decision-making process of the present invention;

[0056] Figure 5 This is a flowchart of the automatic response and closed-loop control of the present invention;

[0057] Figure 6This is a flowchart of the rescue route planning process of the present invention;

[0058] Figure 7 This is a flowchart for assessing the severity level of an accident according to the present invention;

[0059] Figure 8 This is a diagram of the overall system architecture of the present invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] This embodiment discloses a method for multimodal data acquisition and fusion in smart cities, such as... Figures 1 to 8 As shown, it includes the following steps:

[0062] S1. Multi-source sensing and asynchronous acquisition: Asynchronously acquire multimodal raw data of traffic scenarios through visual sensors, geomagnetic sensors, acoustic sensors and millimeter-wave radar;

[0063] S2. Cross-modal feature extraction and alignment: Timestamp alignment and spatial registration are performed on the original multimodal data, and feature vectors for each modality are extracted, including visual feature vector V, geomagnetic feature vector M, acoustic feature vector A, and radar trajectory feature vector R;

[0064] S3. Multi-level Linkage Fusion Decision: Input the feature vector into a pre-trained multimodal fusion neural network model and output the fusion confidence score. ;when If a traffic accident is detected, an event tag is generated; otherwise, monitoring continues.

[0065] S4. Automatic Response and Closed-Loop Control: Based on event tags, traffic signal control strategies and emergency dispatch instructions are automatically triggered; among them, confidence levels are integrated. Calculated using the following formula:

[0066]

[0067] Where f(V,M) is the visual and geomagnetic co-discrimination function, g(A) is the acoustic anomaly scoring function, h(R) is the radar trajectory conflict detection function, α, β, γ are the weighting coefficients of each mode, and α+β+γ=1. This is a preset threshold.

[0068] The core process of the entire method includes four steps: multi-source sensing and asynchronous acquisition, cross-modal feature extraction and alignment, multi-level linkage fusion decision-making, and automatic response and closed-loop control. Data is acquired asynchronously through multiple source sensors (visual, geomagnetic, acoustic, and millimeter-wave radar) to ensure coverage of multiple dimensions of the traffic scenario. Feature extraction and alignment guarantee the consistency of different modal data in time and space. Fusion decision-making outputs a fusion confidence score through a pre-trained neural network model to determine whether a traffic accident has occurred. The automatic response mechanism then triggers traffic control and rescue dispatch based on the decision results.

[0069] The above methods enable multimodal information complementarity in smart cities, with different sensors compensating for the limitations of a single sensor, asynchronous acquisition improving efficiency and avoiding the system burden and delay caused by synchronous acquisition, fusion decision-making improving accuracy, and neural networks dynamically weighting the features of each modality to reduce false alarms and missed alarms, and closed-loop control to achieve automation, thus improving emergency response efficiency by automating the entire process from detection to response.

[0070] In step S1, the acoustic sensor is a microphone array used to capture collision sounds and abnormal horns; the millimeter-wave radar is used to detect vehicle speed and trajectory in low-visibility environments.

[0071] Acoustic sensors employ microphone arrays to directionally capture abnormal sounds (such as collision sounds, sudden braking sounds, and abnormal horn blasts), enhancing the accuracy of event detection through sound source localization. Millimeter-wave radar can reliably detect vehicle speed and trajectory even in low-visibility environments such as rain, fog, and smoke, compensating for the limitations of visual sensors.

[0072] Therefore, this method enhances environmental adaptability, with millimeter-wave radar ensuring detection capabilities in adverse weather conditions. Acoustic verification is further used to confirm whether events detected by vision or radar are genuine accidents through sound characteristics.

[0073] In step S2, feature extraction includes:

[0074] The YOLOv7 model was used to detect the vehicle's attitude and motion state based on visual data.

[0075] Short-time Fourier transform is used to extract frequency domain features from acoustic data;

[0076] Kalman filtering is used to predict the trajectory and identify collision points on radar data.

[0077] YOLOv7 is used to detect vehicle position, type, attitude, and motion status in real time, outputting visual feature vectors. Short-Time Fourier Transform (STFT): Converts acoustic signals into a time-frequency spectrum, extracting frequency domain features (such as energy distribution and frequency peaks) for abnormal sound recognition. Kalman Filter: Smooths and predicts radar trajectories, identifying trajectory anomalies (such as sharp turns, rapid deceleration, and collision points).

[0078] YOLOv7 provides high-precision visual target detection. STFT effectively captures transient acoustic events. Kalman filtering improves the accuracy and robustness of trajectory estimation.

[0079] In step S3, the multimodal fusion neural network model is a Transformer structure based on an attention mechanism, which is used to dynamically weight the importance of each modality feature.

[0080] We employ an attention-based Transformer architecture, which automatically learns the correlations between features across different modalities and assigns different attention weights to different modalities, dynamically adjusting their contributions to the fusion decision. This avoids biases introduced by fixed weights and improves the model's generalization ability. The Transformer effectively captures the complex relationships between multimodal features.

[0081] Fusion confidence The calculation also includes a correction term:

[0082]

[0083] Where t is the duration of the event, T is the maximum allowable response time, and δ is the time decay coefficient.

[0084] The confidence level is adjusted by introducing a time decay factor. This increases the urgency of the system's response to events as their duration increases, dynamically improving the fusion confidence level and preventing delays in rescue efforts due to prolonged unresponsiveness. It ensures the system maintains high sensitivity even when events last for extended periods, preventing the overlooking of ongoing events due to low initial confidence.

[0085] In step S4, the traffic signal control strategy involves dynamically adjusting the signal timing based on real-time traffic flow conditions, including:

[0086] Calculate the queue length Li and flow rate Qi in each direction;

[0087] With the goal of minimizing the total waiting time, the green light duration is optimized using the following formula:

[0088]

[0089] Where C is the signal period duration and n is the number of phases.

[0090] Green light times are dynamically allocated based on queue lengths Li and traffic flow Qi in each direction, with the goal of minimizing total waiting time. In the formula, Gi represents the green light time for the i-th phase, and optimal efficiency is achieved through weighted allocation. Signal timing is adjusted in real-time to adapt to changes in traffic flow. Overall waiting time is reduced through mathematical optimization.

[0091] When an accident occurs, the traffic lights are automatically switched to "red wave control" mode to prioritize traffic flow in the direction of emergency response. Upon detecting an accident, the system automatically switches the traffic lights to "red wave control" mode, creating a continuous green wave in the direction of emergency response to ensure priority passage for vehicles involved in the emergency. This ensures that emergency vehicles can quickly pass through the intersection. The system operates automatically without manual intervention.

[0092] In step S4, the rescue dispatch instruction includes automatically generating a structured accident report and pushing it to the traffic police and emergency medical service platform, and planning the optimal rescue route based on real-time traffic conditions; the specific steps for planning the optimal rescue route based on real-time traffic conditions include:

[0093] S41. Construct a real-time traffic map network: using road nodes as vertices and road segments as edges, the edge weight W is dynamically calculated based on real-time traffic flow data, using the following formula:

[0094]

[0095] Where L is the road segment length, v is the average speed of the road segment, and ρ is the vehicle density;

[0096] S42. Multi-objective path planning: Starting from the current location of the rescue vehicle and ending at the accident site, the improved Dijkstra algorithm is used to solve for the optimal path. The objective function is to minimize the weighted sum of the total travel time T and the path fluctuation coefficient σ: min(αT+βσ), where α and β are weight coefficients, and σ is the weighted average of the speed variance of each segment in the path.

[0097] S43. Dynamic Route Update: Based on real-time traffic flow data, the optimal route is periodically recalculated. If the estimated time of the new route is reduced by more than a threshold Δt compared to the original route, a route update command is sent to the rescue vehicle.

[0098] This method constructs a real-time traffic map network, dynamically calculates road segment weights, and uses an improved Dijkstra's algorithm to find the optimal path. The objective function is to minimize the weighted sum of travel time and path volatility. The system periodically updates the path to ensure that rescue vehicles always travel on the optimal route. It employs multi-objective optimization, balancing time and stability. Dynamic updates adapt to real-time traffic changes. This reduces the arrival time of rescue vehicles.

[0099] The method also includes a preliminary estimation of the accident severity level based on acoustic data and radar impact intensity data, with specific steps including:

[0100] S91. Feature Extraction and Quantization: Extract the peak sound pressure level SPL_max and duration T_s of the collision event from the acoustic data; extract the speed change Δv of the vehicles involved in the accident from the radar data, where Δv is based on the speed difference before and after the impact;

[0101] S92. Severity Index Calculation: The accident severity index S is calculated using a linear weighted model: S = w1 × normalize(SPL_max) + w2 × normalize(T_s) + w3 × normalize(Δv), where w1, w2, and w3 are preset weights, and normalize is a normalization function;

[0102] S93. Severity classification and priority mapping: The severity index S is mapped to a preset severity level, which is Level 1: extremely serious; Level 2: serious; Level 3: general. The system automatically allocates rescue priorities according to the final classification level. Level 1 priority triggers the highest response level and dispatches all available rescue resources. Level 2 priority dispatches resources within a specified range. Level 3 priority dispatches basic rescue units.

[0103] By analyzing characteristics such as peak sound pressure level, duration, and vehicle speed variation, the system calculates an accident severity index S and maps it to three levels. Based on these levels, the system automatically allocates rescue priorities and resources, automating the assessment of accident severity and dispatching rescue resources as needed to avoid waste.

[0104] The second aspect of this invention discloses a smart city multimodal data acquisition and fusion system, which implements a smart city multimodal data acquisition and fusion method, comprising:

[0105] The multimodal sensing module is used to asynchronously collect multimodal raw data of traffic scenarios through visual sensors, geomagnetic sensors, acoustic sensors and millimeter-wave radar;

[0106] The data preprocessing module is used to perform timestamp alignment and spatial registration on the multimodal raw data, and extract feature vectors for each modality, including visual feature vector V, geomagnetic feature vector M, acoustic feature vector A, and radar trajectory feature vector R.

[0107] The fusion decision module is used to input feature vectors into a pre-trained multimodal fusion neural network model and output the fusion confidence score. ;when If a traffic accident is detected, an event tag is generated; otherwise, monitoring continues.

[0108] The control execution module is used to automatically trigger traffic signal control strategies and rescue dispatch instructions based on event tags;

[0109] The communication module is used for data exchange with the traffic control center and rescue platform.

[0110] Example 1: Rear-end collision detection and response at intersections;

[0111] Scenario Description: A rear-end collision occurs between two vehicles at an intersection in a city. The system handles the situation through the following process:

[0112] 1. Data Acquisition:

[0113] Visual sensors: detect when the vehicle in front suddenly slows down and the vehicle behind fails to brake in time;

[0114] Geomagnetic sensor: Detected an abnormal dwell time of a metal object in the lane;

[0115] Acoustic sensor: Captures impact sound (peak sound pressure level 85dB, lasting 0.5s);

[0116] Millimeter-wave radar detected a sudden drop in speed between the two vehicles, and their trajectories intersected.

[0117] 2. Feature extraction and alignment:

[0118] Visual: YOLOv7 outputs the vehicle's bounding box and motion state;

[0119] Acoustics: STFT extracts frequency domain features, anomaly score g(A) = 0.8;

[0120] Radar: Kalman filter predicts trajectory collision, h(R)=0.9;

[0121] The collaborative discriminant function f(V,M) = 0.7.

[0122] 3. Fusion decision: Let α=0.4, β=0.3, γ=0.3, θ1=0.75;

[0123] Cf=0.4×0.7+0.3×0.8+0.3×0.9=0.28+0.24+0.27=0.79>0.75;

[0124] If the incident is determined to be a traffic accident, an event tag will be generated.

[0125] 4. Automatic response:

[0126] Signal control: Switch to red wave control to prioritize the direction of rescue efforts;

[0127] Rescue dispatch: Generate an accident report and push it to the traffic police platform;

[0128] Route planning: Calculate the optimal rescue route, which is expected to save 3 minutes of time.

[0129] Example 2: Multi-vehicle collision in low visibility conditions;

[0130] Scene description: A multi-vehicle collision occurred on a section of highway in foggy weather.

[0131] 1. Data Acquisition:

[0132] Visual sensors are limited;

[0133] The geomagnetic sensor detected an abnormal metallic object in the multi-lane road.

[0134] The acoustic sensor captured multiple impact sounds;

[0135] Millimeter-wave radar detected multiple vehicle trajectories intersecting and sudden speed changes.

[0136] 2. Feature extraction:

[0137] The radar trajectory conflict function h(R) = 0.95;

[0138] Acoustic anomaly score g(A) = 0.9;

[0139] Visual features are less affected by weather, f(V,M)=0.6.

[0140] 3. Integrated Decision Making:

[0141] Cf=0.4×0.6+0.3×0.9+0.3×0.95=0.24+0.27+0.285=0.795>0.75;

[0142] It was determined to be a major accident.

[0143] 4. Severity Level Assessment:

[0144] SPL_max=90dB, T_s=2s, Δv=60km / h;

[0145] After normalization, S = 0.8, which is classified as a Level II major accident;

[0146] Dispatch rescue resources within the designated area.

[0147] 5. Path planning:

[0148] The weights are dynamically updated in the real-time traffic map to avoid congested road sections;

[0149] The optimal path is found using an improved Dijkstra algorithm with α=0.7, β=0.3, and path variability coefficient σ=0.1.

[0150] Finally, it should be noted that the above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A smart city multi-modal data collection fusion method, characterized in that, The method comprises the following steps: S1. Multi-source sensing and asynchronous acquisition: Asynchronously acquire multi-modal raw data of the traffic scene through visual sensors, geomagnetic sensors, acoustic sensors, and millimeter wave radars; S2. Cross-modal feature extraction and alignment: Time stamp alignment and spatial registration are performed on the multi-modal raw data, and feature vectors of each modality are extracted, including visual feature vectors V, geomagnetic feature vectors M, acoustic feature vectors A, and radar trajectory feature vectors R; S3. Multi-level linkage fusion decision: input the feature vector into a pre-trained multi-modal fusion neural network model, output fusion confidence When , determine that a traffic accident occurs, generate an event label; otherwise, continue monitoring; S4. Automatic response and closed-loop control: According to the event label, automatically trigger the traffic signal control strategy and rescue dispatch instructions; wherein the fusion confidence is calculated by the following formula: ; wherein f(V,M) is a visual and geomagnetic collaborative discrimination function, g(A) is an acoustic anomaly scoring function, h(R) is a radar track conflict detection function, and α, β, γ are weight coefficients of each modality, and α+β+γ=1, is a preset threshold. 2.The smart city multi-modal data collection and fusion method of claim 1, wherein, In step S1, the acoustic sensor is a microphone array, which is used to capture collision sounds and abnormal sirens; the millimeter wave radar is used to detect vehicle speed and trajectory in low visibility environment. 3.The smart city multi-modal data acquisition and fusion method of claim 1, wherein, In step S2, the feature extraction includes: Using YOLOv7 model to detect vehicle posture and motion state on visual data; Using short-time Fourier transform to extract frequency domain features from acoustic data; Using Kalman filter to perform trajectory prediction and conflict point identification on radar data. 4.The smart city multi-modal data collection and fusion method of claim 1, wherein, In step S3, the multi-modal fusion neural network model is a Transformer structure based on attention mechanism, which is used to dynamically weight the importance of each modality feature. 5.The smart city multi-modal data collection and fusion method of claim 4, wherein, the fusion confidence The calculation also includes a correction term: ; Wherein, t is the event duration, T is the maximum allowed response time, δ is the time decay coefficient. 6.The smart city multi-modal data acquisition and fusion method of claim 1, wherein, In step S4, the traffic signal control strategy is to dynamically adjust the signal timing according to the real-time traffic flow state, including: Calculate the queue length Li and the flow Qi of each direction; Using the following formula to optimize the green light duration with the goal of minimizing the total waiting time: ; Where C is the signal cycle duration, n is the number of phases.

7. The method according to claim 6, wherein, When an accident occurs, the signal light is forced to switch to "red wave control" mode, and the rescue direction traffic flow is preferentially guided. 8.The smart city multi-modal data acquisition and fusion method of claim 1, wherein, In step S4, the rescue dispatching instruction includes automatically generating a structured accident report and pushing it to the traffic police and emergency platform, and planning the optimal rescue path based on real-time road conditions; The specific steps of planning the optimal rescue path based on real-time road conditions include: S41. Construct a real-time road condition graph network: take road nodes as vertices and road segments as edges, and the weight W of the edge is dynamically calculated according to real-time traffic flow data, the calculation formula is: ; Where L is the length of the road segment, v is the average speed of the road segment, and ρ is the vehicle density; S42. Multi-objective path planning: taking the current position of the rescue vehicle as the starting point and the accident site as the end point, using the improved Dijkstra algorithm to solve the optimal path, the objective function is to minimize the weighted sum of the total travel time T and the path fluctuation coefficient σ: min(αT+βσ), where α and β are weight coefficients, and σ is the weighted average of the speed variance of each road segment in the path; S43. Dynamic path update: periodically recalculate the optimal path according to the real-time received traffic flow data, and if the estimated time of the new path is saved more than the threshold Δt than the original path, send the path update instruction to the rescue vehicle. 9.The smart city multi-modal data acquisition and fusion method of claim 1, wherein, The method further comprises preliminarily estimating the accident severity level according to acoustic data and radar impact intensity data, and the specific steps include: S91. Feature extraction and quantization: extract the peak sound pressure level SPL_max and the duration T_s of the collision event from the acoustic data; extract the speed change Δv of the vehicle involved in the accident from the radar data, Δv is based on the speed difference before and after the impact; S92. Severity index calculation: the accident severity index S is calculated using a linear weighted model: S = w1 x normalize(SPL_max) + w2 x normalize(T_s) + w3 x normalize(Av), where w1, w2, w3 are preset weights, and normalize is a normalization function; S93. Grade division and priority mapping: the severity index S is mapped to a preset severity grade, and the severity grade is: first grade: extremely serious; second grade: serious; third grade: general; the system automatically assigns rescue priorities according to the final grade, the first priority triggers the highest response level, and dispatches all available rescue resources; the second priority dispatches resources within a specified range; the third priority dispatches basic rescue units.

10. A smart city multi-modal data acquisition and fusion system for implementing a smart city multi-modal data acquisition and fusion method according to any one of claims 1-9, characterized in that, Comprise: a multi-modal sensing module for asynchronously collecting multi-modal raw data of a traffic scene through a visual sensor, a geomagnetic sensor, an acoustic sensor, and a millimeter wave radar; a data preprocessing module for timestamp alignment and spatial registration of the multi-modal raw data, and extracting feature vectors of each modality, including a visual feature vector V, a geomagnetic feature vector M, an acoustic feature vector A, and a radar trajectory feature vector R; a fusion decision module, configured to input the feature vector into a pre-trained multi-modal fusion neural network model to output a fusion confidence ; when , determining that a traffic accident occurs, and generating an event label; otherwise, continuing to monitor; a control execution module for automatically triggering traffic signal control strategies and rescue dispatch instructions according to event labels; a communication module for data interaction with a traffic control center and a rescue platform.