Traffic safety real-time monitoring and early warning method and device based on infrasound sensor
Through the infrasound sensor combined with improved weighted noise adaptive modal decomposition and self-supervised comparison learning model, the weak disturbance identification problem of the traffic safety monitoring system in complex environments is solved, high-precision and low-cost real-time monitoring and multi-level response are achieved, and the system's intelligence level is improved.
Patent Information
- Application Number
- CN202510522281.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing traffic safety monitoring system has reduced the recognition accuracy and response capabilities in low visibility environments, making it difficult to adapt to weak disturbance recognition in complex traffic environments, and lacks adaptive and low-cost training mechanisms and multi-level response strategies, and the level of system intelligence is limited.
Infrasound sensors are used for signal acquisition, combined with improved weighted noise adaptive modal decomposition and self-supervised comparison learning model, real-time identification and hierarchical response to traffic abnormal events is achieved through multi-scale decomposition, timing enhancement and dynamic early warning linkage technology.
It significantly improves the recognition accuracy and real-time performance in complex traffic environments, reduces deployment costs, realizes the adaptability and intelligence of the system, and has multi-level linkage response capabilities.
Smart Images

Figure CN120279711A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of infrasound sensors, and particularly to a real-time traffic safety monitoring and early warning method and device based on infrasound sensors. Background Art
[0002] With the continuous expansion and complexity of urban traffic systems, traffic safety issues have become increasingly prominent. Especially in high-risk sections such as highways, tunnels, bridges, and mountain roads, traffic accidents occur frequently, not only threatening the lives and property safety of the people, but also posing higher requirements for road traffic efficiency and management scheduling. Traditional traffic safety monitoring systems mainly rely on perception means such as video surveillance, millimeter-wave radar, vehicle-mounted GPS information, and road-embedded induction coils. These technologies have certain practicality in conventional scenarios, but still face various challenges in actual applications. For example, video surveillance systems are easily interfered by external factors such as lighting, weather, occlusion, and blind spots. Especially in low visibility environments such as at night, in foggy weather, or inside tunnels, the recognition accuracy and response ability significantly decrease; although millimeter-wave radar is not affected by light, its ability to classify complex traffic behaviors is weak, and the cost is relatively high; the construction period of embedded coils is long, the maintenance cost is high, and it cannot be flexibly adapted to different road structures and traffic environments.
[0003] In recent years, with the development of artificial intelligence and Internet of Things technologies, some traffic perception systems based on multi-sensor fusion have been gradually proposed, attempting to improve the detection accuracy of traffic abnormal events through data fusion. However, most of the solutions rely on video, image, or high-frequency sensor signals, and the research on the utilization of low-frequency, non-structural signals is still in its infancy. In contrast, environmental infrasound signals have natural advantages such as long propagation distance, strong penetration ability, and being unaffected by visible light and obstacles, and are particularly suitable for monitoring sudden or hidden traffic abnormal behaviors, such as flat tires, rear-end collisions, collisions, sudden brakes, vehicle instability, etc. However, due to the extremely low frequency of infrasound signals (usually between 0.01 Hz and 20 Hz), their non-linear and non-stationary characteristics are obvious, and the difficulty of signal acquisition and feature extraction is much higher than that of conventional audio or image signals. How to accurately extract the features of abnormal infrasound events from complex backgrounds is still a technical difficulty in the current field.
[0004] In the existing research on traffic anomaly detection, there have also been a small number of explorations to apply acoustic sensing technology to anomaly behavior recognition. For example, some literature has proposed sound recognition systems based on on-vehicle microphones or environmental sound sources, which classify traffic events through machine learning methods. However, most of these methods rely on medium and high-frequency audible sound waves and lack support for sampling and modeling low-frequency sounds. At the same time, existing technologies generally adopt supervised learning algorithms and rely on a large number of manually labeled anomaly event samples, resulting in poor model generalization ability and high deployment costs, and it is difficult to meet the traffic event recognition requirements under different regions and road conditions. In addition, conventional machine learning models such as support vector machines, decision trees, and CNNs often ignore the evolution law of traffic anomaly events in the time dimension, cannot capture the dynamic change trend of traffic behavior, and the recognition results lack real-time, continuity, and interpretability.
[0005] Regarding the modeling technology of infrasound signals, some current research has adopted methods such as wavelet transform and empirical mode decomposition for time-frequency analysis, but there are problems such as low decomposition accuracy, serious mode mixing, and poor anti-noise ability. As a relatively advanced improved algorithm in recent years, the complete ensemble empirical mode decomposition can alleviate the mode mixing problem to a certain extent, but it still has defects such as mode mismatch, many redundant components, and lack of discriminability in complex traffic environments. In addition, existing models often do not fully utilize the cross-structural features of multi-modal infrasound signals and lack a flexible enhancement mechanism to simulate the weak perturbation features under various anomaly events, resulting in low recognition rates of the models when facing weak signal events such as weak collisions, local tire bursts, and local slips.
[0006] In addition, in terms of anomaly event modeling, the current mainstream recognition algorithms mostly adopt fully supervised training methods and rely heavily on large-scale labeled datasets. However, traffic accident data has characteristics such as scarcity, non-reproducibility, and high annotation costs. Some self-supervised contrast learning methods that have emerged in recent years provide a new path for modeling unlabeled time series data, but most of the applications focus on the fields of images, speech, or financial time series, and the self-supervised representation learning for low-frequency acoustic signals, especially in traffic scenarios, has not yet formed a complete technical system. In addition, existing self-supervised methods mostly generate positive and negative samples with fixed augmentation means and lack a dedicated enhancement mechanism for modal structure perturbations in traffic anomaly behaviors, making it difficult to obtain stable representations in signal environments with modal asynchrony and uneven perturbations.
[0007] In the application process of the abnormal event recognition results, most existing systems only complete event recognition or alarm actions. They lack a perfect multi-level response strategy and linkage mechanism, and cannot classify and respond or handle events at different risk levels. They lack the comprehensive judgment ability of the spatio-temporal aggregation trend of events, which is prone to false alarms, missed alarms or response lags. In addition, the current system is usually a passive response mechanism, unable to achieve closed-loop control among event recognition, response, policy feedback, and model optimization, and the overall intelligence level of the system is limited.
[0008] To sum up, in the real-time monitoring and early warning of traffic safety abnormal events, the existing technologies mainly have the following technical deficiencies: (1) The ability to collect, model, and extract multi-modal features of low-frequency sound signals is weak, and it is difficult to meet the recognition requirements of weak disturbances in complex traffic environments; (2) Strongly dependent on supervised learning and manual annotation, lacking an adaptive and low-cost training mechanism, with poor generalization ability; (3) Lack of efficient and structured positive and negative sample enhancement and contrast learning strategies, unable to effectively model the dynamic evolution mode of traffic events; (4) The response mechanism lacks functions of grading, linkage, and policy feedback, with weak system closed-loop ability and limited practicality. Therefore, there is an urgent need for a real-time monitoring and early warning method for traffic safety that integrates a highly robust signal decomposition method, an advanced self-supervised representation modeling mechanism, and a hierarchical response control strategy to effectively solve the above problems in the existing technologies.
[0009] Therefore, how to provide a real-time monitoring and early warning method and device for traffic safety based on infrasound sensors is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0010] An object of the present invention is to propose a real-time monitoring and early warning method and device for traffic safety based on infrasound sensors. The present invention makes full use of the propagation characteristics of environmental infrasound signals, the feature modeling ability of the self-supervised contrast learning model, and the modal decomposition and dynamic early warning linkage technology, and details the whole process from low-frequency sound signal acquisition, modal decomposition, time series enhancement, contrast learning training to traffic abnormal event recognition and multi-level linkage early warning, with the advantages of strong adaptability, high recognition accuracy, good real-time performance, and low deployment cost.
[0011] According to the real-time monitoring and early warning method and device for traffic safety based on infrasound sensors of the embodiments of the present invention, the following steps are included:
[0012] 1. A real-time monitoring and early warning method for traffic safety based on infrasound sensors, characterized by including the following steps:
[0013] S1. Deploy multiple infrasound sensors in the monitoring and early warning area, and continuously and real-time collect the original infrasound signals;
[0014] S2. Preprocess the collected original infrasound signal to obtain a purified infrasound signal sequence;
[0015] S3. Use the complete ensemble empirical mode decomposition algorithm to perform multi-scale decomposition on the purified infrasound signal sequence to obtain multiple intrinsic mode function sequences;
[0016] S4. Construct the time series structure for each intrinsic mode function sequence, and use methods such as time perturbation, window clipping, and amplitude scaling to generate enhanced sample pairs, and construct positive and negative samples;
[0017] S5. Input the enhanced samples into the self-supervised contrast learning model, and use the time series encoder to perform feature representation learning on the input samples to obtain the feature vector representation;
[0018] S6. Use a preset contrast loss function to train the self-supervised contrast learning model, and optimize the model parameters by minimizing the distance difference between the enhanced sample pairs to perform the ability of normal traffic state and abnormal traffic state;
[0019] S7. Process the real-time collected infrasound signal into input samples and input them into the trained model for real-time identification of traffic abnormal events;
[0020] S8. Generate corresponding warning signals according to the recognition results and perform linkage warning through communication methods.
[0021] Optionally, the S2 specifically includes: performing normalization processing on the original infrasound signal sequence collected by the infrasound acquisition module to make its mean zero and variance one, and then applying a band-pass filter to the normalized signal for filtering. After filtering, use the multi-window moving average method for noise suppression to eliminate random interference and pseudo-signals in the background environment, and finally obtain a purified infrasound signal sequence for subsequent decomposition and feature extraction.
[0022] Optionally, the S3 specifically includes:
[0023] S31. Define the purified infrasound signal sequence as S clean (t), and input it into the improved weighted noise adaptive mode decomposition module, and add multiple groups of weighted Gaussian white noise sequences ω i (t)·ε i (t), where i = 1, 2,..., N, ω i (t) is the dynamic noise weight function, and ε i (t) is a standard normal distribution random variable, and construct a perturbation signal sequence:
[0024] S i (t) = S clean (t) + ωi (t)·ε i (t)
[0025] S32. Apply the empirical mode decomposition algorithm to each group of disturbance signals S i (t) to extract the j-th order intrinsic mode function IMF i,j (t), where j = 1, 2,..., M, and M is the maximum decomposition order;
[0026] S33. Calculate the signal-to-noise ratio η of each order of mode function j , defined as:
[0027]
[0028] where, R(t) is the current residual signal;
[0029] S34. According to the preset threshold η th , only retain the mode function IMF j ≥η th , and discard the redundant mode components with relatively low signal-to-noise ratio; j (t);
[0030] S35. When the residual signal R(t) no longer has local extreme points or meets the stop criterion, terminate the decomposition process, and finally output the optimized intrinsic mode sequence group {IMF1(t), IMF2(t),..., IMF K (t)}, where K ≤ M, which is the number of effective modes screened by the signal-to-noise ratio.
[0031] Optionally, the S4 specifically includes:
[0032] S41. Perform time-domain alignment and normalization splicing on the intrinsic mode function sequence group {IMF1(t), IMF2(t),..., IMF K (t)} to construct a multi-channel time series input matrix X(t), where K is the number of modes and T is the time step length;
[0033] S42. Perform a multi-level mode perception enhancement operation based on the matrix X(t) to generate structured positive and negative sample pairs;
[0034] S43. Assemble the positive and negative samples generated by enhancement into a sample pair set
[0035] Optionally, the S42 specifically includes:
[0036] S421. Perform mode attention perturbation enhancement to construct a mode attention weight sequence w = {w1, w2,..., w K}, where each weight w k ∈[0.5, 1.5] is a weight factor dynamically sampled in the traffic infrasound frequency band, and performs multiplicative perturbation on each modal sequence IMF k (t):
[0037]
[0038] All combinations constitute enhanced positive samples
[0039] S422. Perform segmented perturbation enhancement and local structure preservation operations. Divide X(t) along the time axis into m segments, denoted as {X (1) , X (2) ,..., X (m)}}. For each segment, use different amplitude perturbation coefficients α i ∈[0.8, 1.2] and different masking methods to generate enhanced positive samples During the reconstruction process, retain the original mean and coefficient of variation of each segment to maintain the local dynamic pattern;
[0040] S423. Perform structure distortion enhancement. Map X(t) to the frequency domain based on the short-time Fourier transform Apply a perturbation function in the frequency domain The perturbed spectrum is:
[0041]
[0042] Perform the inverse short-time Fourier transform on it to obtain enhanced negative samples
[0043] S424. Perform cross-modal recombination enhancement. Combine and rearrange the IMF components of different modalities, including swapping the modal channel order and intercepting discontinuous modal segments, to form enhanced negative samples with the characteristics of potential abnormal behavior signals
[0044] Optionally, the specific steps of S5 are as follows:
[0045] S51. Construct a self-supervised contrastive learning model architecture. The model includes an encoder module, a projection head module, and a sample similarity evaluation module. Among them, the encoder module adopts a bidirectional time series modeling structure, which consists of a group of stacked bidirectional long short-term memory networks. The projection head module is a fully connected mapping network with a non-linear transformation layer for mapping feature vectors to the contrastive learning space. The sample similarity evaluation module is used to calculate the normalized similarity between different sample representations;
[0046] S52. For the sample pair set For each sample x in it, input the encoder module to extract the temporal feature representation h;
[0047] S53. Input the encoder output h into the projection head module to perform non-linear mapping to obtain the feature vector z in the contrast space. The transformation formula is:
[0048] z = σ(W2·GAU(W1·h + b1) + b2);
[0049] Among them, GAU(·) is a gated activation unit that integrates a dual-channel feature gating mechanism and is applicable to the deep transformation of non-stationary traffic infrasound signals. W1 and W2 are trainable weights, b1 and b2 are bias terms, and σ represents a normalization or activation function;
[0050] S54. For each pair of positive samples and negative samples extract their encoded projection representations as Construct a set of contrast sample pairs based on cosine similarity;
[0051] S55. Construct a complete self-supervised representation training structure to support the input of multiple positive samples and multiple negative samples to the contrast loss function.
[0052] Optionally, the gated activation unit in S53 specifically includes: This gated activation unit is used to enhance the sensitivity of the model to abnormal components in traffic infrasound signals. The calculation method includes the following operations:
[0053]
[0054] Among them, y is the feature vector output by the bidirectional long short-term memory network encoder, W t , W g are the learnable parameters of the gating path, b t , b g are bias vectors, tanh(·) and σ(·) are the hyperbolic tangent function and the Sigmoid function respectively, and ⊙ represents an element-wise multiplication operation, which is used to perform dynamic gating adjustment on the main activation path;
[0055] The gated activation unit also includes the following structural features: Introduce a modal selection weight adjustment mechanism in the modal combination representation of each temporal sample, making the high-frequency modal more sensitive to events such as tire blowouts and impacts, and the low-frequency modal more sensitive to behaviors such as slippage and instability;
[0056] The gating structure has a residual path structure, that is, introduce a shortcut connection to retain the stability of the feature distribution of the original time series;
[0057] The GAU module is configured to be embedded in the projection head structure, placed between the linear mapping and normalization layers, serving as a non-linear conversion bridge from the encoder output to the contrast representation space;
[0058] During the training process, the structure automatically learns the adjustment function of each channel gating through the contrast loss function, enabling the model to better identify the subtle non-stable perturbations induced in infrasound signals;
[0059] The GAU output feature vector z is a representation in the high-dimensional contrast space, used to calculate the semantic similarity between traffic states.
[0060] Optionally, S6 specifically includes:
[0061] S61. Train the self-supervised contrast learning model using a preset modal weighted gating dynamic contrast loss function, which is used to optimize the feature separation ability of the model in the traffic infrasound event scenario, and is defined as follows:
[0062]
[0063]
[0064] Among them, is the contrast space feature vector mapped from the positive sample pair, is the contrast space feature vector of the negative sample, represents the cosine similarity between vectors, ‖·‖2 represents the L2 norm, λ i ∈[0,1] is the positive sample gating confidence factor adaptively learned by the model, used to dynamically adjust the positive pair loss intensity, α∈[0,1] is the weighted coefficient of the cosine and Euclidean metrics, w j ∈[0,1] is the modal sensitivity weighted coefficient, calculated based on the IMF modal signal energy obtained by CEEMDAN decomposition, is the penalty factor for the negative sample contrast term, and N is the number of negative samples.
[0065] S62. Define the overall loss function as the average loss value of all positive sample pairs:
[0066]
[0067] where n is the total number of positive sample pairs;
[0068] S63. Based on the loss function Use the backpropagation and gradient optimization algorithm to update the parameters of the self-supervised contrast learning model. The updated parameters include encoder parameters, gating activation unit parameters, projection head module parameters, gating confidence factor generation network parameters, and parameters of the modal sensitivity weighted function;
[0069] S64. When the loss function converges stably on the training set and the validation set, and the semantic discrimination ability of the model in different abnormal traffic scenarios reaches the set threshold, save the trained self-supervised contrast model.
[0070] Optionally, S7 specifically includes: After purifying and modal decomposing the real-time collected infrasound signal sequence, input it into the trained self-supervised contrast learning model, and through the processing of the time series encoder, gated activation unit and projection head module, extract the contrast feature representation vector of the traffic state, match the feature representation with the pre-established traffic event representation space, and perform real-time inference according to the similarity to determine whether the current signal belongs to the abnormal traffic event type. The abnormal events include collision, rear-end collision, flat tire, emergency braking or vehicle skidding, etc. When the similarity score exceeds the set threshold, immediately determine the abnormal event type, record the time stamp and event attributes of the recognition result, and use it as the input for subsequent early warning decision-making.
[0071] Optionally, S8 specifically includes:
[0072] S81. Use the identified abnormal traffic event category as the input for early warning decision-making, and generate multi-level early warning instructions based on the type, confidence score and modal perturbation characteristics of the abnormal event. The early warning levels include:
[0073] High-intensity early warning: When the event recognition confidence is higher than the preset high threshold and the event type is high-risk behaviors such as collision, chain rear-end collision, or serious flat tire, generate a first-level high-intensity early warning instruction;
[0074] Medium-level early warning: When the recognition confidence is in the middle threshold range, or the event type is potential risk behaviors such as emergency braking or skidding, generate a second-level medium-level early warning instruction;
[0075] Low-level reminder: When the event recognition confidence is low, or the modal perturbation characteristics are unstable but have suspicious patterns, generate a third-level early warning reminder and conduct continuous observation;
[0076] S82. Trigger a multi-level linkage early warning mechanism based on the event level, recognition confidence and traffic environment parameters. The linkage response devices include:
[0077] In-vehicle terminal response: Broadcast voice and visual prompt information to the in-vehicle systems of vehicles near the event occurrence area to alert drivers to take avoidance actions;
[0078] Road information device linkage: Real-time release risk tips or traffic guidance information through infrastructure such as road information screens and signal lights;
[0079] Traffic management platform synchronization: Synchronously upload the abnormal event category, recognition time, signal source ID, and geographical location information to the traffic management cloud platform for analysis and processing by the background command system;
[0080] Regional linkage mechanism: When multiple medium- and high-level events occur continuously in the same region within a short period of time, trigger a regional-level linkage warning for temporary traffic dispatching or restricted passage;
[0081] S83. Store the triggered warning events and their linkage response actions in the event database, mark the warning trigger conditions, linkage strategies, and response results for subsequent event attribution and system optimization;
[0082] S84. Periodically update the warning strategy set based on the linkage response log and event recognition history, and construct an adaptive warning strategy template for different road types, time periods, environmental conditions, etc., to achieve dynamic adjustment;
[0083] S85. When multi-source sensors or multi-channel infrasound input nodes make consistent warning judgments on the same event type, merge the warning levels and enhance the response level to ensure the system's comprehensive response ability to high-intensity and multi-point collaborative events;
[0084] S86. Use the warning strategy execution process and feedback results for the training optimization module to evaluate indicators such as strategy coverage rate, false alarm rate, and linkage timeliness, and adjust the model parameters and trigger thresholds to achieve closed-loop optimization of the warning system.
[0085] A real-time traffic safety monitoring and warning method and device based on infrasound sensors, including the following modules:
[0086] Infrasound acquisition and preprocessing module, used to collect infrasound signals through infrasound sensors deployed in scenarios such as traffic roads, bridges, or tunnels;
[0087] Modal decomposition and enhancement module, used to perform improved weighted noise adaptive modal decomposition on the purified infrasound signal sequence, extract multiple effective intrinsic mode function sequences, and perform operations such as structural perturbation, frequency perturbation, and cross-modal recombination on the sequences to generate positive and negative sample pairs for self-supervised learning training;
[0088] Self-supervised feature modeling module, used to perform feature representation learning on positive and negative sample pairs based on a time series encoder, gated activation unit, and non-linear projection structure, and train the model through a preset modal weighted gated dynamic contrast loss function to obtain a contrast learning feature model with the ability to distinguish traffic states;
[0089] Real-time anomaly recognition module, used to identify whether there are abnormal events and their corresponding types in the current traffic state;
[0090] The multi - level early warning linkage module is used to generate multi - level early warning instructions according to the identified abnormal event types and confidence levels, and link the in - vehicle terminal, road information equipment and traffic control platform for hierarchical response, adapting to different road environments and traffic safety strategies;
[0091] The closed - loop optimization and strategy feedback module is used to record and evaluate the early warning response process and model output results, dynamically update the early warning strategy template and model parameters, and perform closed - loop optimization on the event recognition and early warning mechanism.
[0092] The beneficial effects of the present invention are as follows:
[0093] (1) By introducing an improved weighted noise adaptive mode decomposition algorithm, the present invention performs high - precision multi - scale decomposition on the original infrasound signal, can effectively extract the intrinsic mode functions with abnormal characteristics in the traffic scene, significantly improves the modeling ability of low - frequency non - stationary signals, solves the problems of serious existing mode aliasing, low signal - to - noise ratio, and fuzzy feature distribution, and provides a more discriminative signal basis for subsequent abnormal event recognition.
[0094] (2) The present invention constructs a self - supervised contrast learning model based on a gated activation structure and a modal weight mechanism. Through a structured time - series enhancement strategy and a modal - perception contrast loss function, it effectively improves the recognition robustness of the model for weak - perturbation events such as tire blowouts, rear - end collisions, collisions, and sudden braking, solves the defects of existing methods that overly rely on manual annotation and lack dynamic modeling ability, and has wide adaptability in the actual traffic environment with scarce data and complex scenarios.
[0095] (3) The present invention constructs a multi - level closed - loop early warning mechanism covering "recognition - early warning - response - optimization", realizes real - time hierarchical judgment, intelligent linkage response and strategy dynamic feedback after the occurrence of abnormal events, not only improves the timeliness and accuracy of traffic safety event response, but also provides an intelligent early warning solution for the traffic management system with traceable strategies, updatable rules and adjustable responses, significantly enhancing the practicality, flexibility and continuous optimization ability of the system. Description of the Drawings
[0096] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0097] Figure 1 is a flowchart of the traffic safety real - time monitoring and early warning method and device based on infrasound sensors proposed by the present invention;
[0098] Figure 2Flow chart of feature modeling and training for the self-supervised contrast learning model of the traffic safety real-time monitoring and early warning method and device based on infrasound sensors proposed by the present invention;
[0099] Figure 3 Schematic diagram of the real-time recognition and multi-level early warning linkage mechanism for traffic abnormal events of the traffic safety real-time monitoring and early warning method and device based on infrasound sensors proposed by the present invention. Detailed implementation manners
[0100] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0101] Refer to Figures 1-3 , the traffic safety real-time monitoring and early warning method and device based on infrasound sensors include the following steps:
[0102] S1. Deploy multiple infrasound sensors in the monitoring and early warning area to continuously and real-time collect raw infrasound signals;
[0103] S2. Preprocess the collected raw infrasound signals to obtain a purified infrasound signal sequence;
[0104] S3. Use the complete ensemble empirical mode decomposition algorithm to perform multi-scale decomposition on the purified infrasound signal sequence to obtain multiple intrinsic mode function sequences;
[0105] S4. Construct the temporal structure for each intrinsic mode function sequence, generate enhanced sample pairs using time perturbation, window cropping, and amplitude scaling methods, and construct positive samples and negative samples;
[0106] S5. Input the enhanced samples into the self-supervised contrast learning model, use a time series encoder to perform feature representation learning on the input samples, and obtain a feature vector representation;
[0107] S6. Use a preset contrast loss function to train the self-supervised contrast learning model, optimize the model parameters by minimizing the distance difference between the enhanced sample pairs, and perform the ability of normal traffic states and abnormal traffic states;
[0108] S7. Process the real-time collected infrasound signals into input samples and input them into the trained model for real-time recognition of traffic abnormal events;
[0109] S8. Generate corresponding warning signals according to the recognition results and perform linkage warning through communication means.
[0110] This method can achieve efficient recognition and intelligent early warning of traffic abnormal events without manual annotation by constructing a full - process structure from infrasound acquisition, modal decomposition, enhanced training to early warning output, and has the advantages of strong real - time performance, system closed - loop, and high adaptive ability.
[0111] In this embodiment, the specific steps of S2 are as follows: normalize the original infrasound signal sequence collected by the infrasound acquisition module so that its mean is zero and variance is one. Then, apply a band - pass filter to the normalized signal for filtering. After filtering, use the multi - window moving average method for noise suppression to eliminate random interference and pseudo - signals in the background environment, and finally obtain a purified infrasound signal sequence for subsequent decomposition and feature extraction.
[0112] By normalizing, band - pass filtering, and multi - window moving average noise reduction processing on the original infrasound signal, the signal - to - noise ratio of the infrasound signal can be significantly improved, background interference can be reduced, and a stable and clean input data basis for subsequent modal decomposition and feature modeling can be provided.
[0113] In this embodiment, the specific steps of S3 are as follows:
[0114] S31: Define the purified infrasound signal sequence as S clean (t), and input it into the improved weighted noise adaptive modal decomposition module, and add multiple groups of weighted Gaussian white noise sequences ω i (t)·ε i (t), where i = 1, 2,..., N, ω i (t) is a dynamic noise weight function, and ε i (t) is a standard normal distribution random variable, and construct a perturbation signal sequence:
[0115] S i (t)=S clean (t)+ω i (t)·ε i (t)
[0116] S32: Apply the empirical mode decomposition algorithm to each group of perturbation signals S i (t) to extract the j - th order intrinsic mode function IMF i,j (t), where j = 1, 2,..., M, and M is the maximum decomposition order;
[0117] S33: Calculate the signal - to - noise ratio η j of each order of modal function, which is defined as:
[0118]
[0119] Among them, R(t) is the current residual signal;
[0120] S34. According to the preset threshold η th , only retain the intrinsic mode functions IMF j that satisfy η th ≥η j (t), and discard the redundant modal components with relatively low signal-to-noise ratio;
[0121] S35. When the residual signal R(t) no longer has local extreme points or satisfies the stopping criterion, terminate the decomposition process, and finally output the optimized intrinsic mode sequence group {IMF1(t), IMF2(t),..., IMF K (t)}, where K ≤ M, which is the number of effective modes screened by the signal-to-noise ratio.
[0122] By introducing an improved weighted noise CEEMDAN decomposition strategy and adopting a Gaussian perturbation and modal signal-to-noise ratio screening mechanism, the present invention can effectively avoid modal aliasing and redundancy, enhance the interpretability of the structure of low-frequency non-stationary signals, and provide high-quality characteristic modes for anomaly event detection.
[0123] In this embodiment, the specific steps of S4 are as follows:
[0124] S41. Align and normalize the splicing of the intrinsic mode function sequence group {IMF1(t), IMF2(t),..., IMF K (t)} in the time domain to construct a multi-channel time series input matrix X(t), where K is the number of modes and T is the time step length;
[0125] S42. Perform a multi-level modal perception enhancement operation based on the matrix X(t) to generate structured positive and negative sample pairs;
[0126] S43. Assemble the enhanced positive and negative samples into a sample pair set
[0127] By uniformly aligning, splicing, and multi-dimensionally enhancing the IMF modes, the present invention generates a set of positive and negative sample pairs, provides sample support from multiple angles and scales for self-supervised training, and significantly improves the model's ability to distinguish abnormal perturbation patterns and generalization performance.
[0128] In this embodiment, the specific steps of S42 are as follows:
[0129] S421. Perform modal attention perturbation enhancement to construct a modal attention weight sequence w = {w1, w2,..., w K}, where each weight w k ∈[0.5, 1.5] is a weight factor dynamically sampled in the infrasonic frequency band of traffic, and for each modal sequence IMF k(t) Perform multiplicative perturbation:
[0130]
[0131] All combinations form enhanced positive samples
[0132] S422. Perform segmented perturbation enhancement and local structure preservation operation. Divide X(t) into m segments along the time axis, denoted as {X (1) , X (2) ,..., X (m)}. For each segment, use different amplitude perturbation coefficients α i ∈[0.8, 1.2] and different masking methods to generate enhanced positive samples During the reconstruction process, retain the original mean and coefficient of variation of each segment to maintain the local dynamic pattern;
[0133] S423. Perform structure distortion enhancement. Map X(t) to the frequency domain based on the short-time Fourier transform Apply a perturbation function in the frequency domain The perturbed spectrum is:
[0134]
[0135] Perform the inverse short-time Fourier transform on it to obtain enhanced negative samples
[0136] S424. Perform cross-modal recombination enhancement. Combine and rearrange the IMF components of different modalities, including swapping the modality channel order and intercepting discontinuous modality segments, to form enhanced negative samples with the characteristics of potential abnormal behavior signals
[0137] By introducing various enhancement strategies such as modal attention perturbation, segmented perturbation, frequency domain distortion, and cross-modal recombination, the present invention can generate rich and structurally diverse positive and negative sample pairs, significantly improve the model's learning ability and discrimination ability for weak abnormal patterns in traffic infrasound signals, and enhance the effectiveness and generalization ability of self-supervised training.
[0138] In this embodiment, the specific steps of S5 are as follows:
[0139] S51. Construct a self-supervised contrastive learning model architecture. The model includes an encoder module, a projection head module, and a sample similarity evaluation module. Among them, the encoder module adopts a bidirectional time series modeling structure, which is composed of a group of stacked bidirectional long short-term memory networks. The projection head module is a fully connected mapping network with a non-linear transformation layer for mapping feature vectors to the contrastive learning space. The sample similarity evaluation module is used to calculate the normalized similarity between different sample representations;
[0140] S52. For each sample x in the sample pair set , input it into the encoder module to extract the temporal feature representation h;
[0141] S53. Input the encoder output h into the projection head module to perform non-linear mapping to obtain the feature vector z in the contrast space. The transformation formula is:
[0142] z = σ(W2 · GAU(W1 · h + b1) + b2);
[0143] where GAU(·) is a gated activation unit that integrates a two-channel feature gating mechanism and is suitable for deep transformation of non-stationary traffic infrasound signals. W1 and W2 are trainable weights, b1 and b2 are bias terms, and σ represents a normalization or activation function;
[0144] S54. For each pair of positive samples and negative samples , extract their encoded projection representations as respectively, and construct a contrast sample pair set based on cosine similarity;
[0145] S55. Construct a complete self-supervised representation training structure to support the input of multiple positive samples and multiple negative samples to the contrast loss function.
[0146] By constructing a self-supervised contrast learning model including a bidirectional time series encoder, a gated activation mechanism, and a non-linear projection structure, the present invention can effectively capture the temporal dynamics and modal perturbation features in infrasound signals, improve the model's expression ability and sample discrimination ability for traffic states, and provide a high-quality feature basis for anomaly event recognition under unsupervised conditions.
[0147] In this embodiment, the gated activation unit in S53 specifically includes: The gated activation unit is used to enhance the sensitivity of the model to abnormal components in traffic infrasound signals, and its calculation method includes the following operations:
[0148] GAU(y) = tanh(W t y + b t ) ⊙ σ(W g y + b g );
[0149] where y is the feature vector output by the bidirectional long short-term memory network encoder, W t , W g are the learnable parameters of the gating path, b t , b g are bias vectors, tanh(·) and σ(·) are the hyperbolic tangent function and the Sigmoid function respectively, and ⊙ represents an element-wise multiplication operation for dynamically gating and adjusting the main activation path;
[0150] The gating activation unit further includes the following structural features: a modal selection weight adjustment mechanism is introduced in the modal combination representation of each time series sample, making the high-frequency modes more sensitive to events such as tire blowouts and impacts, and the low-frequency modes more sensitive to behaviors such as slippage and instability;
[0151] The gating structure has a residual path structure, that is, a shortcut connection is introduced to preserve the stability of the feature distribution of the original time series;
[0152] The GAU module is configured to be embedded in the projection head structure, placed between the linear mapping and the normalization layer, serving as a non-linear transformation bridge from the encoder output to the contrast representation space;
[0153] During the training process, the structure automatically learns the adjustment function of each channel gating through the contrast loss function, enabling the model to better identify the subtle non-stable perturbations induced in the infrasound signals;
[0154] The GAU output feature vector z is a representation in the high-dimensional contrast space, used to calculate the semantic similarity between traffic states.
[0155] By designing a gating activation unit with modal selection adjustment, residual connection, and non-linear mapping capabilities, the present invention significantly enhances the sensitivity and expression ability of the model to abnormal infrasound features in different frequency bands, improves the recognition accuracy of micro-perturbation events such as traffic tire blowouts and slippage, and at the same time maintains the stability of the feature distribution, providing a more discriminative deep semantic representation for contrastive learning.
[0156] In this embodiment, the S6 specifically includes:
[0157] S61. Train the self-supervised contrastive learning model using a preset modal weighted gating dynamic contrast loss function, and the loss function is used to optimize the feature separation ability of the model in the traffic infrasound event scenario, defined as follows:
[0158]
[0159] Among them, is the contrast space feature vector mapped from the positive sample pair, is the contrast space feature vector of the negative sample, represents the cosine similarity between vectors, ‖·‖2 represents the L2 norm, λ i ∈ [0,1] is the positive sample gating confidence factor adaptively learned by the model, used to dynamically adjust the positive pair loss intensity, α ∈ [0,1] is the weighted coefficient of the cosine and Euclidean metrics, w j ∈ [0,1] is the modal sensitivity weighted coefficient, calculated according to the IMF modal signal energy obtained by CEEMDAN decomposition, is the penalty factor for the negative sample comparison item, and N is the number of negative samples.
[0160] S62. Define the overall loss function as the average loss value of all positive sample pairs:
[0161]
[0162] where n is the total number of positive sample pairs;
[0163] S63. Based on the loss function Use the backpropagation and gradient optimization algorithm to update the parameters of the self-supervised contrast learning model. The updated parameters include the encoder parameters, gated activation unit parameters, projection head module parameters, gated confidence factor generation network parameters, and the parameters of the modal sensitivity weighting function;
[0164] S64. When the loss function converges stably on the training set and the validation set, and the semantic discrimination ability of the model in different abnormal traffic scenarios reaches the set threshold, save the trained self-supervised contrast model.
[0165] The present invention uses the trained self-supervised model to perform feature matching and event discrimination on the real-time input infrasound modality, and can quickly and efficiently complete the recognition of traffic abnormal behaviors, having the advantages of low computational cost, strong real-time performance, and flexible deployment.
[0166] In this embodiment, the S7 specifically includes: After the real-time collected infrasound signal sequence is purified and mode decomposed, it is input into the trained self-supervised contrast learning model. After being processed by the time series encoder, gated activation unit and projection head module, the contrast feature representation vector of the traffic state is extracted. Match the feature representation with the pre-established traffic event representation space, and perform real-time inference according to the similarity to determine whether the current signal belongs to the abnormal traffic event type. The abnormal events include collision, rear-end collision, flat tire, emergency braking, or vehicle skidding, etc. When the similarity score exceeds the set threshold, immediately determine the abnormal event type, record the time stamp and event attributes of the recognition result, and use it as the input for subsequent early warning decision-making.
[0167] In this embodiment, the S8 specifically includes:
[0168] S81. Use the recognized abnormal traffic event category as the input for early warning decision-making, and generate multi-level early warning instructions based on the type, confidence score, and modal perturbation characteristics of the abnormal event. The early warning levels include:
[0169] High-intensity early warning: When the event recognition confidence is higher than the preset high threshold, and the event type is high-risk behaviors such as collision, chain rear-end collision, or serious flat tire, generate a first-level high-intensity early warning instruction;
[0170] Medium-level warning: When the recognition confidence level is within the intermediate threshold range, or the event type is potential risk behaviors such as emergency braking, skidding, etc., generate a secondary medium-level warning instruction;
[0171] Low-level reminder: When the event recognition confidence level is low, or the modal disturbance characteristics are unstable but have suspicious patterns, generate a tertiary warning reminder and conduct continuous observation;
[0172] S82. Trigger a multi-level linkage warning mechanism based on the event level, recognition confidence level, and traffic environment parameters. The linkage response devices include:
[0173] In-vehicle terminal response: Broadcast voice and visual reminder information to the in-vehicle systems of vehicles near the event occurrence area to alert drivers to take evasive actions;
[0174] Road information device linkage: Real-time publish risk reminders or traffic guidance information through infrastructure such as road information screens and signal lights;
[0175] Traffic management platform synchronization: Synchronously upload the abnormal event category, recognition time, signal source ID, and geographical location information to the traffic management cloud platform for the background command system to analyze and process;
[0176] Regional linkage mechanism: When multiple medium- and high-level events occur continuously in the same region within a short period of time, trigger a regional-level linkage warning for temporary traffic dispatching or restricted passage;
[0177] S83. Store the triggered warning events and their linkage response actions in the event database, mark the warning trigger conditions, linkage strategies, and response results for subsequent event attribution and system optimization;
[0178] S84. Periodically update the warning strategy set based on the linkage response logs and event recognition history, and construct an adaptive warning strategy template for different road types, time periods, environmental conditions, etc. to achieve dynamic adjustment;
[0179] S85. When multi-source sensors or multi-channel infrasound input nodes make consistent warning judgments on the same event type, merge the warning levels and enhance the response level to ensure the system's comprehensive response ability to high-intensity, multi-point collaborative events;
[0180] S86. Use the warning strategy execution process and feedback results for the training and optimization module, evaluate indicators such as strategy coverage rate, false alarm rate, and linkage timeliness, and adjust the model parameters and trigger thresholds to achieve closed-loop optimization of the warning system.
[0181] A real-time traffic safety monitoring and warning method and device based on infrasound sensors, including the following modules:
[0182] Infrasound acquisition and preprocessing module, which is used to collect infrasound signals through infrasound sensors deployed in scenarios such as traffic roads, bridges or tunnels;
[0183] Modal decomposition and enhancement module, which is used to perform improved weighted noise adaptive modal decomposition on the purified infrasound signal sequence, extract multiple effective intrinsic mode function sequences, and perform operations such as structural perturbation, frequency perturbation and cross-modal recombination on the sequences to generate positive and negative sample pairs for self-supervised learning training;
[0184] Self-supervised feature modeling module, which is used to perform feature representation learning on positive and negative sample pairs based on a time series encoder, a gated activation unit and a non-linear projection structure, and train the model through a preset modal weighted gated dynamic contrast loss function to obtain a contrast learning feature model with the ability to distinguish traffic states;
[0185] Real-time anomaly recognition module, which is used to identify whether there are abnormal events and their corresponding types in the current traffic state;
[0186] Multi-level early warning linkage module, which is used to generate multi-level early warning instructions according to the identified abnormal event types and confidence levels, and link vehicle-mounted terminals, road information devices and traffic control platforms for hierarchical responses to adapt to different road environments and traffic safety strategies;
[0187] Closed-loop optimization and policy feedback module, which is used to record and evaluate the early warning response process and model output results, dynamically update the early warning policy template and model parameters, and perform closed-loop optimization on the event recognition and early warning mechanism.
[0188] This device integrates five modules: acquisition, processing, modeling, recognition and response, forming a complete software and hardware closed-loop system, with the capabilities of modular deployment, intelligent analysis and linkage control, significantly improving the automation level of safety monitoring in traffic scenarios.
[0189] Embodiment 1:
[0190] To verify the feasibility of the present invention in implementation, the present invention is applied to the section from K128+300 to K128+800 of the east section of G45 Expressway in a certain city of a certain province. This section is a five-lane divided expressway with a large annual traffic volume and frequent accidents. Especially at night and in rainy and foggy weather, sudden abnormal events such as flat tires, rear-end collisions, sudden braking and vehicle skidding often occur. To improve the traffic safety control level in this area, a real-time traffic safety monitoring and early warning system based on infrasound sensors proposed by the present invention is deployed.
[0191] A total of 12 infrasound sensors were deployed at multiple key locations such as the central isolation belt and the right - hand guardrail area of this section of the road. The acquisition frequency band ranges from 0.01 Hz to 20 Hz, the sampling frequency is 100 Hz, and it operates continuously all - weather. The deployment period was from October 1, 2024 to February 28, 2025, lasting for 5 months. The data was pre - processed by edge devices and then uploaded to the back - end server for modal decomposition, self - supervised modeling, and anomaly recognition. This system uses an improved CEEMDAN algorithm to decompose the signals and enhances the feature sensitivity of the self - supervised contrast learning model through a gated activation unit. Compared with traditional video + rule - triggered systems, accurate and efficient anomaly recognition and real - time early warning are achieved without relying on manually labeled data.
[0192] During the implementation period, a total of 8356 infrasound event samples were recorded, of which 2634 were automatically identified by the system as potential traffic anomalies. After manual comparison and verification with traffic accident reports, 312 real events were confirmed. The detection hit rate of the system was 91.3%, and the false - alarm rate was 5.7%. Especially in rainy and foggy weather, the system can effectively detect weak infrasound signals generated by abnormal tire friction, identifying 63 flat - tire accidents, 42 of which occurred at night. The average response time was 4.8 seconds, much faster than the average delay of 16.4 seconds for traditional manual dispatching response.
[0193] In addition, the system triggered 127 high - level early - warning instructions in total, successfully linked the in - vehicle terminal voice broadcast and the road lane - changing reminder screen 81 times. After two consecutive major rear - end accident potential events occurred, the automatic regional lockdown mechanism was implemented, and the longest early - warning lead time reached 42 seconds, providing a key reaction window for accident prevention.
[0194] The following table shows the typical anomaly detection and response statistics after the system deployment. The data sources include multi - dimensional cross - verification data such as traffic management department accident bulletins, road video playback comparison, sensor raw signal analysis, and system log records.
[0195] Table 1: Statistical Table of Typical Anomaly Event Recognition and Response after System Deployment
[0196]
[0197] From the data in Table 1, it can be seen that the real - time traffic safety monitoring and early - warning method based on infrasound sensors proposed in the present invention shows good recognition accuracy and response efficiency under multiple different meteorological conditions and different types of traffic anomaly events. In five typical events from October 2024 to February 2025, the system covered various complex weather scenarios such as clear at night, rainy at night, foggy during the day, and snowy at night, demonstrating good environmental adaptability. Especially in low - visibility conditions such as at night and in rainy and foggy weather where traditional video systems are difficult to play a role, the system still operates stably.
[0198] In terms of event types, typical sudden traffic abnormalities such as tire blowouts, rear-end collisions, sudden braking, and skidding were successfully identified. Among them, tire blowouts were detected twice, occurring under extreme conditions of clear night and snowy night, respectively. The system recognition confidence was above 0.92, and the response time was 4.3 seconds and 4.2 seconds, respectively, which was much faster than the 15.8 seconds and 17.9 seconds required for manual confirmation, fully demonstrating the sensitive capture ability of infrasound signals for weak disturbances such as tire ruptures. The rear-end collision occurred during the rainy night, with a system recognition confidence of up to 0.95, and triggered an early warning within 5.1 seconds, successfully linking the road lane change prompt system to avoid the possible expansion of chain accidents. Both sudden braking and skidding occurred in low visibility environments. The system accurately judged and issued voice prompts to assist the vehicles in front and behind to quickly respond to avoidance.
[0199] In addition, from the comparison between "system response time" and "manual confirmation time", it can be seen that the system's average response time in the five events was 4.56 seconds, while the average confirmation required for manual intervention was 17.24 seconds, a difference of nearly 13 seconds, indicating that the system has obvious real-time warning advantages and can greatly improve the efficiency of accident response. At the same time, no false alarms occurred in any of the five sample events, further confirming the high robustness of the system's self-supervised comparative learning model in feature extraction and anomaly recognition.
[0200] The "Whether the warning was successful" column shows that the system successfully triggered the warning mechanism in all events, and the linkage effect after the warning was triggered was good, especially in continuous risk events such as skidding and rear-end collisions, which effectively promoted the rapid response collaboration between the on-board terminal, information screen and traffic management system, reflecting the system's linkage closed-loop capability.
[0201] In summary, the system not only maintains high accuracy and low latency in complex environments, but also has strong scene adaptability, event recognition capabilities and collaborative response mechanisms, showing significant application value in improving traffic operation safety and reducing accident rates. Each data in the table comes from cross-validation results in actual deployment scenarios, which truly reflects the feasibility and advancement of the invention in actual engineering.
[0202] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A real-time traffic safety monitoring and early warning method and device based on infrasound sensors, characterized in that, It includes the following steps: S1. Deploy multiple infrasound sensors in the monitoring and early warning area to continuously and real-time collect the original infrasound signals; S2. Preprocess the collected original infrasound signals to obtain a purified infrasound signal sequence; S3. Use the complete ensemble empirical mode decomposition algorithm to perform multi-scale decomposition on the purified infrasound signal sequence to obtain multiple intrinsic mode function sequences; S4. Construct the time series structure for each intrinsic mode function sequence, and use methods such as time perturbation, window clipping, and amplitude scaling to generate enhanced sample pairs, and construct positive samples and negative samples; S5. Input the enhanced samples into the self-supervised contrast learning model, and use a time series encoder to perform feature representation learning on the input samples to obtain a feature vector representation; S6. Use a preset contrast loss function to train the self-supervised contrast learning model, and optimize the model parameters by minimizing the distance difference between the enhanced sample pairs, and perform the ability of normal traffic state and abnormal traffic state; S7. Process the real-time collected infrasound signals into input samples and input them into the trained model for real-time identification of traffic abnormal events; S8. Generate corresponding early warning signals according to the identification results and perform linkage early warning through communication methods.
2. The traffic safety real-time monitoring and warning method and device based on infrasound sensors according to claim 1, characterized in that The specific content of S2 includes: perform normalization processing on the original infrasound signal sequence collected by the infrasound acquisition module, apply a band-pass filter for filtering, and use the multi-window moving average method for noise suppression after filtering to obtain a purified infrasound signal sequence.
3. The traffic safety real-time monitoring and warning method and device based on infrasound sensors according to claim 1, characterized in that, The specific content of S3 includes: S31. Define the purified infrasound signal sequence as S clean (t), and input it into the improved weighted noise adaptive mode decomposition module. Add multiple groups of weighted Gaussian white noise sequences to construct the perturbation signal sequence S i (t); S32. Apply the empirical mode decomposition algorithm to each group of disturbance signals S i (t), and extract the j-th order intrinsic mode function IMF i,j (t) of each group based on the residual signal; S33. Calculate the signal-to-noise ratio η of each order modal function based on the intrinsic mode function j ; S34. According to the preset threshold η th , only retain the mode functions that satisfy η j ≥η th ; S35. When the residual signal no longer has local extreme points or meets the stopping criterion, terminate the decomposition process and finally output the optimized group of intrinsic mode sequences {IMF1(t), IMF2(t),..., IMF K (t)}, where K ≤ M.
4. The real-time traffic safety monitoring and early warning method and device based on infrasound sensors according to claim 1, characterized in that, The specific content of S4 includes: S41. Align the intrinsic mode function sequence group in the time domain and perform normalized splicing to construct a multi-channel time series input matrix X(t); S42. Perform a multi-level modal perception enhancement operation based on the matrix X(t) to generate structured positive and negative sample pairs; S43. Assemble the generated positive and negative samples into a set of sample pairs 5. The traffic safety real-time monitoring and warning method and device based on infrasound sensors according to claim 4, characterized in that, The specific content of S42 includes: S421. Perform modal attention perturbation enhancement to construct a modal attention weight sequence w = {w1, w2,..., w K}, where each weight w k ∈[0.5, 1.5] is a weight factor dynamically sampled in the traffic infrasound frequency band, and perform multiplicative perturbation on each modal sequence IMF k (t): All combinations constitute enhanced positive samples S422. Perform segmented perturbation enhancement and local structure preservation operations. Divide X(t) along the time axis into m segments, denoted as {X (1) , X (2) ,..., X (m)}. For each segment, use different amplitude perturbation coefficients α i ∈ [0.8, 1.2] and different masking methods to generate enhanced positive samples During the reconstruction process, retain the original mean and coefficient of variation of each segment to maintain the local dynamic pattern; S423. Perform structural distortion enhancement and map X(t) to the frequency domain based on the short-time Fourier transform. Apply a perturbation function in the frequency domain. The perturbed spectrum is: Then perform the inverse short-time Fourier transform to obtain enhanced negative samples S424. Perform cross-modal recombination enhancement, combine and rearrange the IMF components of different modalities, including swapping the order of modal channels and intercepting discontinuous modal segments, to form enhanced negative samples with the characteristics of potential abnormal behavior signals 6. The traffic safety real-time monitoring and early warning method and device based on infrasound sensors according to claim 1, characterized in that, The specific content of S5 includes: S51. Construct a self-supervised contrast learning model architecture. The model includes an encoder module, a projection head module, and a sample similarity evaluation module. Among them, the encoder module adopts a bidirectional time series modeling structure, which is composed of a group of stacked bidirectional long short-term memory networks. The projection head module is a fully connected mapping network with a non-linear transformation layer for mapping the feature vector to the contrast learning space. The sample similarity evaluation module is used to calculate the normalized similarity between different sample representations; S52. For each sample x in the sample pair set , input it into the encoder module to extract the temporal feature representation h; S53. Input the encoder output h into the projection head module to perform non-linear mapping to obtain the feature vector z in the contrast space. The transformation formula is: z = σ(W2·GAU(W1·h + b1)+b2); where GAU(·) is a gated activation unit that integrates a dual-channel feature gating mechanism and is suitable for deep transformation of non-stationary traffic infrasound signals. W1 and W2 are trainable weights, b1 and b2 are bias terms, and σ represents a normalization or activation function; S54. For each pair of positive samples and negative samples extract their encoded projection representations as Construct a set of contrast sample pairs based on cosine similarity; S55. Construct a complete self-supervised representation training structure to support the contrast loss function to input multiple positive samples and multiple negative samples.
7. The traffic safety real-time monitoring and warning method and device based on infrasound sensors according to claim 1, characterized in that The specific content of S53 includes: This gated activation unit is used to enhance the sensitivity of the model to abnormal components in traffic infrasound signals. The calculation method includes the following operations: GAU(y) = tanh(W t y + b t ) ⊙ σ(W g y + b g ) ; Among them, y is the feature vector output by the bidirectional long short-term memory network encoder, W t , W g are the learnable parameters of the gating path, b t , b g are the bias vectors, tanh(·) and σ(·) are the hyperbolic tangent function and the Sigmoid function respectively, and ⊙ represents the element-wise multiplication operation.
8. The traffic safety real-time monitoring and early warning method and device based on infrasound sensors according to claim 1, characterized in that The specific steps of S6 are as follows: S61. Train the self-supervised contrast learning model using a preset modal weighted gated dynamic contrast loss function. The loss function is used to optimize the feature separation ability of the model in the traffic infrasound event scenario and is defined as follows: Among them, is the contrast space feature vector mapped from the positive sample pair, is the contrast space feature vector of the negative sample, represents the cosine similarity between vectors, ‖·‖2 represents the L2 norm, and λ i ∈[0,1] is the positive sample gating confidence factor adaptively learned by the model, used to dynamically adjust the intensity of the positive pair loss. α∈[0,1] is the weighted coefficient of the cosine and Euclidean metrics, and w j ∈[0,1] is the modal sensitivity weighted coefficient, calculated according to the IMF modal signal energy obtained by CEEMDAN decomposition, is the penalty factor for the negative sample contrast term, and N is the number of negative samples; S62. Define the overall loss function as the average loss value of all positive sample pairs: where n is the total number of positive sample pairs; S63. Based on the loss function Use the backpropagation and gradient optimization algorithm to update the parameters of the self-supervised contrastive learning model. The updated parameters include the encoder parameters, the gated activation unit parameters, the projection head module parameters, the gated confidence factor generation network parameters, and the parameters of the modal sensitivity weighting function; S64. When the loss function converges stably on the training set and the validation set, and the semantic discrimination ability of the model in different abnormal traffic scenarios reaches the set threshold, save the trained self-supervised contrast model.
9. The traffic safety real-time monitoring and early warning method and device based on infrasound sensors according to claim 1, characterized in that, The specific steps of S8 are as follows: S81. Use the identified abnormal traffic event category as the input of the early warning decision. Generate multi-level early warning instructions based on the type, confidence score, and modal perturbation characteristics of the abnormal event. The early warning levels include: High-intensity warning: When the event recognition confidence is higher than the preset high threshold, and the event type is collision, chain rear-end collision, or serious tire blowout behavior, generate a first-level high-intensity warning instruction; Medium-level warning: When the recognition confidence is in the middle threshold range, or the event type is a potential risk behavior, generate a second-level medium-level warning instruction; Low-level reminder: When the event recognition confidence is low, or the modal perturbation characteristics are unstable but have a suspicious pattern, generate a third-level warning reminder for continuous observation; S82. Trigger a multi-level linkage early warning mechanism based on the event level, recognition confidence, and traffic environment parameters. The linkage response devices include: In-vehicle terminal response: Broadcast voice and visual reminder information to the in-vehicle systems of vehicles near the event occurrence area to alert drivers to take evasive actions; Road information device linkage: Real-time publish risk reminder or traffic guidance information through road information screens and signal infrastructure; Traffic management platform synchronization: Synchronize the abnormal event category, recognition time, signal source ID, and geographical location information to the traffic management cloud platform for analysis and processing by the background command system; Regional linkage mechanism: When multiple medium- and high-level events occur continuously in the same region within a short period of time, trigger a regional-level linkage early warning for temporary traffic dispatching or restricted passage; S83. Store the triggered early warning events and their linkage response actions in the event database, marking the early warning trigger conditions, linkage strategies, and response results; S84. Periodically update the early warning strategy set based on the linkage response logs and event recognition history, and construct an adaptive early warning strategy template for different road types, time periods, and environmental conditions to achieve dynamic adjustment; S85. When multiple-source sensors or multiple infrasound input nodes make consistent early warning judgments on the same event type, merge the early warning levels and enhance the response level; S86. Use the early warning strategy execution process and feedback results for training and optimization modules, evaluate indicators such as strategy coverage rate, false alarm rate, and linkage timeliness, and adjust model parameters and trigger thresholds to achieve closed-loop optimization of the early warning system.
10. A real-time traffic safety monitoring and warning device based on infrasound sensors, which executes the real-time traffic safety monitoring and warning method based on infrasound sensors described in any one of claims 1 to 8, characterized in that, It includes the following modules: Infrasound acquisition and preprocessing module, which is used to collect infrasound signals through infrasound sensors deployed in traffic roads, bridges, tunnels, and other scenarios; The modal decomposition and enhancement module is used to perform improved weighted noise adaptive modal decomposition on the purified infrasound signal sequence, extract multiple effective intrinsic mode function sequences, and perform operations such as structural perturbation, frequency perturbation, and cross-modal recombination on the sequences to generate positive and negative sample pairs for self-supervised learning training; The self-supervised feature modeling module is used to perform feature representation learning on the positive and negative sample pairs based on a time series encoder, a gated activation unit, and a non-linear projection structure, and train the model through a preset modal weighted gated dynamic contrast loss function to obtain a contrastive learning feature model with the ability to distinguish traffic states; The real-time anomaly recognition module is used to identify whether there are abnormal events and their corresponding types in the current traffic state; The multi-level early warning linkage module is used to generate multi-level early warning instructions according to the identified abnormal event types and confidence levels, and link the in-vehicle terminal, road information equipment, and traffic control platform for hierarchical responses to adapt to different road environments and traffic safety strategies; The closed-loop optimization and policy feedback module is used to record and evaluate the early warning response process and the model output results, dynamically update the early warning policy template and model parameters, and perform closed-loop optimization on the event recognition and early warning mechanism.