Traffic accident early warning method and system based on prediction residual and bidirectional attention

CN122821769APending Publication Date: 2026-09-25SHANDONG JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611264458.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]此外,交通状态参数与预测残差的关联未被充分利用,未能建模二者的相互印证与增强关系,而且残差信号易退化,当模型预测能力下降或数据缺失时,残差失效会导致识别性能不稳定,现有预测模型与下游识别任务缺乏协同设计,残差的判别潜力未被充分发挥

Benefits of technology

本发明对每一个监测节点分别采用中位数与中位数绝对偏差进行节点级自适应归一化,消除了预测残差在不同监测节点之间的尺度差异,使较弱的残差信号在统一的标准尺度下变得可比;并且,中位数与中位数绝对偏差属于稳健统计量,对交通事故等极端值不敏感,可有效降低事故样本对归一化参数的污染,从而提升残差特征的可用性与稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821769A_ABST
    Figure CN122821769A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of intelligent transportation and deep learning, and particularly relates to a traffic accident early warning method and system based on predicted residual error and bidirectional attention. The method comprises: obtaining historical time series data of traffic state parameters, and obtaining predicted residual error through a prediction model; performing node self-adaptive normalization on the predicted residual error and decomposing the predicted residual error into multiple residual sub-channels, and splicing the residual sub-channels with traffic state channels to form a multi-channel fusion input tensor; dividing the input tensor into a traffic state channel group and an abnormal channel group, and interacting and modulating the traffic state channel group and the abnormal channel group through a bidirectional gate attention module; and finally splicing the initial traffic state channel group, the modulated traffic state channel group and the abnormal channel group, and inputting the spliced groups into a classification network to identify and warn traffic accidents. The present application can still achieve accurate traffic accident identification and warning under the conditions of weak residual error, cross-node incomparability and residual error degradation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent transportation and deep learning technology, and particularly relates to a traffic accident early warning method and system based on prediction residuals and bidirectional attention. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] The timely and accurate identification of traffic accidents and road anomalies is of great significance for improving road network operational safety, shortening accident response time, and ensuring road traffic efficiency. With the widespread deployment of traffic detection equipment on urban roads and highways, the automatic identification of traffic accidents based on time-series data such as traffic flow, occupancy rate, and speed collected from road network monitoring sections has become an important research direction in the field of intelligent transportation.

[0004] Existing methods based on predicted residuals obtain predicted residuals by training a model using normal data and then distinguish them based on abnormal fluctuations in residuals caused by accidents. However, the predicted residual signals are weak and easily drowned out by noise, and normal fluctuations and accident anomalies can overlap, resulting in insufficient identification capabilities. At the same time, the predicted residuals are not directly comparable between different monitoring nodes, and the baselines and fluctuation amplitudes of different cross-sections differ significantly. Uniform processing will introduce scale bias.

[0005] Furthermore, the correlation between traffic state parameters and prediction residuals has not been fully utilized, and the mutual confirmation and enhancement relationship between the two has not been modeled. Moreover, residual signals are prone to degradation. When the model's predictive ability declines or data is missing, residual failure can lead to unstable recognition performance. Existing prediction models and downstream recognition tasks lack collaborative design, and the discrimination potential of residuals has not been fully realized.

[0006] Therefore, how to achieve robust and accurate traffic accident identification and early warning under conditions of weak residual signals, incomparability across nodes, underutilization of cross-channel correlation, and easy degradation of residuals is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0007] To overcome the shortcomings of the existing technologies, this invention provides a traffic accident early warning method based on predictive residuals and bidirectional attention. It aims to enhance residual representation capabilities through node-level adaptive normalization and multi-sub-channel decomposition, and to achieve cross-channel mutual verification and enhancement between traffic state parameters and residual anomalies using a bidirectional gating attention mechanism with bias initialization. Furthermore, it combines adaptive backoff gating based on residual time-series statistics to suppress the impact of residual signal degradation. Thus, it achieves robust and accurate traffic accident identification and early warning even under conditions of weak residual signals, incomparability across nodes, underutilization of cross-channel correlation, and easy residual degradation.

[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: Firstly, a traffic accident early warning method based on predictive residuals and bidirectional attention is disclosed, including: Historical time-series data of various traffic state parameters of each monitoring node in the traffic monitoring road network are obtained, and prediction models are used to predict the historical time-series data and obtain prediction residuals. The predicted residual is subjected to node adaptive normalization to obtain the residual standard score, which is then decomposed into multiple residual sub-channels and spliced ​​with the traffic state channels corresponding to various traffic state parameters to form a multi-channel fusion input tensor. The multi-channel fusion input tensor is divided into traffic state channel group and abnormal channel group, and the interactively modulated traffic state channel group and abnormal channel group are obtained through a bidirectional gating attention module. Calculate the temporal statistics of the residual sub-channels in the abnormal channel group. When the temporal statistics are lower than a preset threshold, replace the gating weights used to modulate the traffic state channel group in the bidirectional gating attention module with learnable backoff gating weights. The initial traffic state channel group, the modulated traffic state channel group, and the modulated abnormal channel group are concatenated and input into the classification network to obtain traffic accident identification results and issue warnings accordingly.

[0009] Furthermore, the prediction model is a dual-path temporal prediction model, including a parallel Mamba branch and a temporal convolutional network branch, as well as a gated fusion unit for adaptive weighted fusion of the outputs of the two branches.

[0010] Furthermore, the prediction head of the prediction model has a dual-head structure, including a point prediction head and an uncertainty prediction head. The specific expression for the standardized prediction residual is as follows:

[0011] in, Indicates the actual observed value; This indicates the point prediction value given by the point prediction header; The uncertainty is the prediction uncertainty output by the uncertainty prediction head corresponding to the predicted value of the point, and its value is always positive; ε represents a preset positive number to avoid the denominator being zero. This represents the standardized prediction residual.

[0012] Furthermore, the node adaptive normalization includes: using the median of the predicted residuals of the monitoring nodes as the location parameter and the median absolute deviation as the scale parameter, standardizing and truncating the predicted residuals to obtain the residual standard score.

[0013] Furthermore, the plurality of residual sub-channels include: residual standard score, residual absolute value, residual difference, and residual moving average; wherein, the residual absolute value is the element-wise absolute value of the residual standard score, the residual difference is the first-order difference of the residual standard score along the time dimension, and the residual moving average is the result of the residual standard score smoothed by a sliding window.

[0014] Furthermore, the bidirectional gating attention module includes two branches: an anomaly-to-traffic state gating branch and a traffic state-to-anomaly gating branch.

[0015] Furthermore, the anomaly-to-traffic state gating branch encodes the anomaly channel group via convolution, then maps it to the (0,1) interval using a sigmoid function. This mapping is then added to the gating bias parameter and divided by two to obtain the first gating weight.

[0016] Wherein, A is the abnormal channel group; (·) represents the convolutional coding operation of the anomaly-to-traffic state gating branch; σ(·) represents the sigmoid activation function, which maps the input to the (0,1) interval; The gating bias parameter represents the gating branch from anomaly to traffic state, and its initial value is a preset constant; the first gating weight generated by the gating branch from anomaly to traffic state is represented.

[0017] Furthermore, the time-series statistics are the standard deviations of specified residual sub-channels along the time dimension in the abnormal channel group.

[0018] Furthermore, the bidirectional gating attention module and the classification network are trained using R-Drop regularization during the training phase. The same input sample is forward-propagated twice to obtain a first prediction distribution and a second prediction distribution, respectively. In addition to the classification loss, the bidirectional KL divergence between the first prediction distribution and the second prediction distribution is superimposed as a regularization term in the loss function.

[0019] Secondly, a traffic accident early warning system based on prediction residuals and bidirectional attention is disclosed, including: The data acquisition and residual generation module is used to acquire historical time-series data of various traffic state parameters of each monitoring node in the traffic monitoring road network, and use the prediction model to predict the historical time-series data and obtain the prediction residual. The residual feature construction and fusion module is used to perform node adaptive normalization on the predicted residual to obtain the residual standard score, decompose it into multiple residual sub-channels, and splice it with the traffic state channels corresponding to various traffic state parameters to form a multi-channel fusion input tensor. A bidirectional gated attention modulation module is used to divide the multi-channel fused input tensor into traffic state channel group and abnormal channel group, and obtain the interactively modulated traffic state channel group and abnormal channel group through the bidirectional gated attention module; The residual degradation detection and backoff module is used to calculate the temporal statistics of the residual sub-channels in the abnormal channel group. When the temporal statistics are lower than a preset threshold, the learningable backoff gating weights are used to replace the gating weights in the bidirectional gating attention module used to modulate the traffic state channel group. The feature splicing and classification early warning module is used to splice the initial traffic state channel group, the modulated traffic state channel group, and the modulated abnormal channel group, input them into the classification network to obtain traffic accident identification results, and issue early warnings based on them.

[0020] The above one or more technical solutions have the following beneficial effects: This invention performs node-level adaptive normalization for each monitoring node using the median and median absolute deviation, eliminating scale differences in prediction residuals between different monitoring nodes and making weak residual signals comparable under a unified standard scale. Furthermore, the median and median absolute deviation are robust statistics that are insensitive to extreme values ​​such as traffic accidents, effectively reducing the contamination of normalization parameters by accident samples, thereby improving the usability and stability of residual features.

[0021] This invention decomposes the standard residual score into multiple sub-channels, characterizing the predicted residual from multiple perspectives such as the residual's amplitude level, amplitude intensity, rate of change, and local trend. This amplifies the discrimination information related to traffic accidents contained in the predicted residual and alleviates the problem of weak residual signals.

[0022] This invention utilizes bidirectional gating attention interaction between traffic state channels and abnormal channels to enable traffic state parameter features and residual abnormal features to mutually verify and reinforce each other, making full use of the correlation between cross channels. Furthermore, after bias initialization, the gating weights are close to identity mapping in the initial training stage, avoiding the destruction of the original cross-channel representation by the gating mechanism in the early training stage, making the training process more stable.

[0023] This invention employs an adaptive mechanism that combines residual missing detection with learnable backoff gating. This enables the identification model to automatically reduce its dependence on residual channels when predicting residual signal degradation or failure, significantly improving the robustness of the method under residual failure conditions.

[0024] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0025] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0026] Figure 1 This is a flowchart of a traffic accident early warning method based on prediction residuals and bidirectional attention according to Embodiment 1 of the present invention. Figure 2 This is a schematic diagram of the structure of the Mamba-TCN dual-path temporal prediction model in the traffic accident early warning method based on prediction residuals and bidirectional attention according to Embodiment 1 of the present invention. Figure 3 This is a schematic diagram of the residual feature construction in the traffic accident early warning method based on prediction residuals and bidirectional attention according to Embodiment 1 of the present invention. Figure 4 This is a schematic diagram of the bidirectional gating attention module in the traffic accident early warning method based on prediction residuals and bidirectional attention according to Embodiment 1 of the present invention. Detailed Implementation

[0027] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0028] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0029] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0030] Example 1 Through long-term research and engineering practice, it has been found that traffic accident identification methods based on prediction residuals still have the following technical defects in practical applications.

[0031] First, the prediction residual signal is weak and easily drowned out by noise. Traffic flow data itself has strong random fluctuations. The prediction error of the prediction model contains both abnormal components caused by traffic accidents and a large number of components caused by normal traffic fluctuations. These two components overlap, resulting in a weak discrimination signal of the prediction residual for traffic accidents. If the prediction residual is used directly for classification without processing, the model's ability to identify accident-type samples is often insufficient.

[0032] Second, the prediction residuals are not directly comparable across different monitoring nodes. The baseline levels and fluctuation amplitudes of traffic flow differ significantly across different road sections and monitoring cross-sections, meaning that the same prediction residual represents different degrees of anomaly at different monitoring nodes. Treating the residuals of all monitoring nodes uniformly without differentiation introduces significant inter-node scaling bias, affecting identification accuracy.

[0033] Third, the correlation between traffic state channels and residual channels is not fully utilized. Traffic accidents are usually reflected simultaneously in traffic state parameters such as traffic flow, occupancy, and speed, as well as the predicted residuals calculated from them. Existing methods either simply concatenate the channels and feed them directly into the classifier, or only perform one-way feature interactions between channels, failing to fully model the mutually corroborating and reinforcing correlation between traffic state parameters and residual anomalies, resulting in insufficient utilization of cross-channel information.

[0034] Fourth, the prediction residual signal may degrade or even fail under certain operating conditions. When the prediction model's predictive ability decreases in a specific road segment or time period, or when the residual channel becomes approximately constant in the time dimension due to data loss, sensor failure, or other reasons, the prediction residual will not be able to provide effective accident discrimination information. If the identification model still relies heavily on the residual channel at this time, it will lead to unstable identification performance and a lack of necessary robustness.

[0035] Furthermore, most existing traffic accident identification methods based on prediction residuals employ prediction models with improving traffic volume prediction accuracy as their sole optimization objective. However, high prediction accuracy does not necessarily mean that the generated residuals have good discriminative power for traffic accidents. The lack of targeted collaborative design between the prediction model and downstream accident identification tasks prevents the discriminative potential of residual features from being fully utilized.

[0036] Based on the above issues, such as Figure 1 As shown, this embodiment discloses a traffic accident early warning method based on prediction residuals and bidirectional attention, including: Step 1: Obtain historical time-series data of various traffic state parameters of each monitoring node in the traffic monitoring network, use the prediction model to predict the historical time-series data and obtain the prediction residuals.

[0037] In step 1, historical time-series data of each monitoring node in the traffic monitoring network are first acquired, including three types of traffic state parameters: traffic flow, occupancy rate, and speed. Assume the monitoring network contains N monitoring nodes. Each monitoring node collects traffic flow, occupancy rate, and speed parameters at fixed time intervals (e.g., 5 minutes). Within a time window of length T, the observation data of the nth monitoring node can be represented as a time-series matrix, where the time index t ranges from 1 to T, and the monitoring node index n ranges from 1 to N.

[0038] Based on the acquired historical time-series data, a prediction model is used to perform time-series predictions on the aforementioned traffic state parameters. In this embodiment, the preferred method is as follows: Figure 2 The Mamba-TCN dual-channel time-domain prediction model shown is used as the prediction model.

[0039] The prediction model takes the traffic state parameters of each monitoring node within a historical time window as input and outputs the predicted value of that traffic state parameter. The prediction residual is obtained by subtracting the actual observed value corresponding to each monitoring node from the predicted value, as shown below:

[0040] in, This represents the prediction residual of the nth monitoring node at time t; This represents the actual observed value of the traffic state parameters of the nth monitoring node at time t; This represents the predicted value of the traffic state parameters of the nth monitoring node at time t; the subscript t is the time number, ranging from 1 to T, where T is the length of the time window; the subscript n is the monitoring node number, ranging from 1 to N, where N is the total number of monitoring nodes.

[0041] In this embodiment, the prediction residual can be calculated separately for traffic flow, for occupancy or speed, or calculated separately for multiple traffic state parameters and then used together.

[0042] The predictive model works by training on normal traffic data without traffic accidents to learn the normal evolution of traffic volume. When a traffic accident occurs at a monitoring node during a certain period, its traffic operation deviates from the normal pattern, and there is a large deviation between the predicted value and the actual observed value. The prediction residual then becomes abnormal, thus providing clues for subsequent traffic accident identification.

[0043] like Figure 2As shown, the Mamba-TCN dual-path temporal prediction model includes a parallel Mamba branch, a temporal convolutional network (TCN) branch, and a gated fusion unit. The Mamba branch uses a unidirectional causal state-space model to model the long-range evolution of traffic volume, while the TCN branch uses a causal dilated convolution stack with sequentially increasing dilation rates to model local abrupt changes in traffic volume. The gated fusion unit adaptively weights and fuses the outputs of the two branches, and then maps the results to the prediction head to obtain the predicted value.

[0044] The historical time-series data of the input traffic state parameters are first passed through an input linear projection layer and a layer normalization layer to be mapped to the hidden feature space, thus obtaining the input features. Subsequently, the input features are simultaneously fed into the parallel Mamba branch and the Temporal Convolutional Network (TCN) branch.

[0045] The Mamba branch adopts a one-way causal state-space model to model the long-range evolution of traffic volume (such as trends and periods) and outputs the first intermediate feature, as shown in formula (2):

[0046] in, This represents the input features obtained after processing historical traffic volume time-series data through the input linear projection layer and layer normalization; Mamba(·) represents the one-way causal state-space model operation; This represents the first intermediate feature of the Mamba branch output.

[0047] The temporal convolutional network branches are stacked with causal dilated convolutions of successively increasing dilation rates (e.g., 1, 2, 4, 8) to capture local abrupt changes in traffic volume at multiple time scales and output a second intermediate feature, as shown in formula (3):

[0048] in, The meaning of is the same as in formula (2), representing the input feature; TCN(·) represents the temporal convolutional network operation composed of stacked causal dilated convolutions with successively increasing dilation rates; This represents the second intermediate feature of the output of a branch in a temporal convolutional network.

[0049] Real-world traffic accidents often cause instantaneous changes in parameters such as flow rate and speed, as well as an abnormal evolutionary process that lasts for a period of time. Relying solely on convolutions with local receptive fields makes it difficult to detect long-term trend anomalies; while using only global sequence modeling can easily smooth out key local abrupt change signals. This embodiment uses a parallel design of Mamba branches and temporal convolutional network branches to enable the model to retain sensitivity to local abrupt changes while possessing the ability to capture long-range dependencies, effectively covering the multi-timescale manifestations of traffic accidents.

[0050] The outputs of the two branches are adaptively weighted and fused by a gated fusion unit. Specifically, the first intermediate feature and the second intermediate feature are concatenated along the feature dimension, mapped by a multilayer perceptron and activated by the sigmoid function to obtain the fusion gate coefficients with values ​​in the interval (0,1), as shown in formula (4).

[0051] in, , These are the first and second intermediate features, respectively; [·; ·] indicates the operation of concatenating the two features within the brackets along the feature dimension; and σ(·) represents the weight matrix and bias vector of the multilayer perceptron in the gated fusion unit, respectively; · represents matrix multiplication; σ(·) represents the sigmoid activation function, which maps the input to the interval (0,1); g represents the fusion gating coefficient, which takes values ​​in the interval (0,1).

[0052] Subsequently, the intermediate features of the two branches are summed element-wise according to the fusion gating coefficient to obtain the fusion features, as shown in formula (5):

[0053] in, The fusion gating coefficients obtained from formula (4); and These are the first and second intermediate features, respectively; the symbol ⊙ represents element-wise multiplication. and represents the element-wise weighting coefficients of the first intermediate feature and the second intermediate feature, respectively; H represents the fused feature output by the gated fusion unit.

[0054] Finally, the fused feature H is input into the prediction head, and the predicted values ​​of the traffic state parameters are obtained through mapping. In a preferred embodiment, the prediction head adopts a dual-head structure, including a point prediction head and an uncertainty prediction head; the point prediction head outputs the point predicted value of the traffic state parameter, and the uncertainty prediction head outputs the prediction uncertainty corresponding to that point predicted value. In this preferred embodiment, the prediction residual obtained in step 1 is the standardized prediction residual, as shown in formula (6):

[0055] in, Represents the actual observed value; This indicates the point prediction value given by the point prediction header; This represents the prediction uncertainty output by the uncertainty prediction head, corresponding to the predicted value of the point, and its value is always positive. is a preset, extremely small positive number used to avoid the denominator being zero; represents the standardized prediction residual.

[0056] By calculating the standardized residuals, we can identify locations where the uncertainty of the prediction model itself is relatively high. When the value is large, the corresponding residuals are automatically decayed, thereby suppressing spurious anomalies introduced by insufficient predictive power; while for predictive models with high confidence ( The residuals corresponding to the positions where the deviation is relatively small are amplified. These positions are more likely to correspond to actual traffic accidents, thereby enhancing the residuals' ability to distinguish traffic accidents. It should be noted that the dual-head structure is an optional preferred implementation. When the dual-head structure is not used, step 1 calculates the prediction residuals according to formula (1).

[0057] Step 2: Perform node adaptive normalization on the predicted residuals to obtain the residual standard score, decompose it into multiple residual sub-channels, and splice them with the traffic state channels corresponding to various traffic state parameters to form a multi-channel fusion input tensor.

[0058] The residual features of the predicted residuals obtained in step 1 are constructed, such as... Figure 3 As shown, it includes node-level adaptive normalization, residual multi-sub-channel decomposition, and splicing with traffic state parameter channels.

[0059] First, node-level adaptive normalization is performed. Considering that the prediction residuals of different monitoring nodes differ significantly in baseline level and fluctuation amplitude, this embodiment calculates the normalization parameters for each monitoring node separately.

[0060] For the nth monitoring node, the median of all valid prediction residuals over the historical time interval is used as the location parameter of the monitoring node, as shown in formula (7):

[0061] in, This represents the prediction residual of the nth monitoring node at time t, which is the prediction residual obtained in step 1; {·} represents the operation of taking the median along the time dimension, that is, taking the median of the prediction residuals of the nth monitoring node at all valid moments within the historical time interval; This represents the location parameter of the nth monitoring node.

[0062] The scale parameter of the monitoring node is the larger of the median absolute deviation of all effective prediction residuals of the monitoring node with respect to the location parameter multiplied by a preset proportional coefficient and a preset lower limit value, as shown in formula (8).

[0063] in, Let be the prediction residual for the nth monitoring node; The position parameter of the monitoring node is obtained by formula (7); the symbol |·| indicates taking the absolute value; {·} represents the operation of taking the median along the time dimension; 1.4826 is the preset scaling factor, which is used to make the absolute deviation of the median consistent with the standard deviation under the assumption of normal distribution; δ represents the preset lower limit value, which is used to prevent the scaling parameter from being too small; max(·,·) means taking the larger of the two values ​​within the parentheses; This represents the scale parameter of the nth monitoring node.

[0064] Using the median and median absolute deviation as normalization parameters is a robust statistic. Compared to the mean and standard deviation, it is not sensitive to extreme values ​​such as traffic accidents, and can effectively avoid the contamination of normalization parameters by accident samples. Based on the above node-level location and scale parameters, the prediction residuals of the monitoring node are standardized, and the results are truncated to a preset interval to obtain the residual standard score, as shown in formula (9):

[0065] in, To predict residuals; and These are the location and scale parameters of the nth monitoring node, respectively; ε is a preset positive number used to avoid the denominator being zero. This indicates that the value within the parentheses is truncated to a closed interval. Within ], that is, the value is greater than Time to take Less than Time to take ; A preset cutoff threshold (e.g., 10) is used to suppress extreme outliers; Let represent the standard residual score of the nth monitoring node at time t.

[0066] After node-level adaptive normalization, the residuals of different monitoring nodes are unified to a comparable standard scale, solving the problem of incomparability of prediction residuals across nodes.

[0067] Based on this, in order to characterize the accident-related information contained in the prediction residuals from multiple perspectives, this embodiment further decomposes the residuals into multiple sub-channels based on the obtained residual standard scores, thereby deriving multiple residual sub-channels.

[0068] The first residual sub-channel is the residual standard score itself, i.e. It reflects the magnitude and direction of the residual.

[0069] The second residual sub-channel is the absolute value of the residual, which is equal to the element-wise absolute value of the standard residual score. It reflects the magnitude of the residual deviation from the normal level without distinguishing the direction, as shown in formula (10):

[0070] in, The standard residual score is obtained from formula (9); the symbol |·| represents taking the absolute value of each element; This represents the sub-channel of the absolute value of the residual.

[0071] The third residual subchannel is the residual difference, which is equal to the first-order difference of the residual standard score along the time dimension. It reflects the instantaneous rate of change of the residual and is more sensitive to the sudden changes in state caused by traffic accidents, as shown in formula (11):

[0072] in, and Represent the nth monitoring node at time t and Standard score of residuals at time step; This represents the residual molecular channel.

[0073] The fourth residual sub-channel is the residual moving average, which is equal to the result of convolution smoothing the standard residual score through a sliding window of preset length K. It reflects the local trend of the residual, can suppress random noise and highlight persistent anomalies, as shown in formula (12):

[0074] in, This represents the standard residual score of the nth monitoring node at time t+k; K represents the preset length of the sliding window. This represents the residual moving average sub-channel. It should be understood that in formulas (10), (11), and (12), , , Time t and node n are not labeled, but each variable has the same time and node dimensions as the residual standard score. That is, each residual sub-channel has the same time and node dimensions as the residual standard score.

[0075] The four residual sub-channels described above characterize the predicted residuals from four perspectives: amplitude level, amplitude intensity, rate of change, and local trend. This maximizes the amplification of traffic accident-related discrimination information contained within the residual signal, even when it is weak. Those skilled in the art will understand that the number and types of residual sub-channels are not limited to the four mentioned above, and other derived sub-channels can be added as needed.

[0076] Finally, channel splicing is performed. The above-mentioned multiple residual sub-channels are spliced ​​along the channel dimension to form an abnormal channel group; the three types of traffic state parameters, namely traffic flow, occupancy rate, and speed, are spliced ​​along the channel dimension to form a traffic state channel group; then the traffic state channel group and the abnormal channel group are spliced ​​along the channel dimension to obtain the multi-channel fused input tensor, as shown in formula (13):

[0077] In the formula, P represents a traffic state channel group composed of traffic flow, occupancy rate, and speed parameters spliced ​​along the channel dimension; A represents an abnormal channel group composed of residual sub-channels such as residual standard score, residual absolute value, residual difference, and residual moving average spliced ​​along the channel dimension; Concat(·) represents the splicing operation along the channel dimension; X represents the multi-channel fusion input tensor, whose shape is the time window length T multiplied by the total number of channels C, where the total number of channels C is equal to the sum of the number of traffic state channels and the number of residual sub-channels.

[0078] Step 3: Divide the multi-channel fusion input tensor into traffic state channel group and abnormal channel group, and obtain the interactively modulated traffic state channel group and abnormal channel group through the bidirectional gating attention module.

[0079] The multi-channel fused input tensor input bidirectional gated attention module obtained in step 2 is modulated. The structure of the bidirectional gated attention module is as follows: Figure 4 As shown.

[0080] First, the multi-channel fused input tensor X is divided along the channel dimension into a traffic state channel group P and an abnormal channel group A. The traffic state channel group corresponds to three types of traffic state parameters: traffic flow, occupancy rate, and speed. The abnormal channel group corresponds to each residual sub-channel constructed in step 2. The traffic state channel group and the abnormal channel group are encoded by their respective convolutional coding branches. The convolutional coding branches can adopt a structure that includes pointwise convolution, batch normalization, activation functions, and depthwise separable convolution.

[0081] The division is the inverse operation of the channel splicing operation in formula (13) in step 2: according to the channel order during splicing in formula (13), the channels corresponding to the three types of traffic state parameters in the multi-channel fusion input tensor X are segmented to form traffic state channel group P, and the remaining channels corresponding to each residual sub-channel are segmented to form abnormal channel group A. The resulting traffic state channel group and abnormal channel group are completely consistent with the traffic state channel group and abnormal channel group before splicing in step 2 in terms of channel composition, channel order and values, without any re-division. By first splicing into a single tensor and then segmenting according to a fixed channel order, it is convenient to transfer data between processing steps with a unified tensor interface, and the semantic boundaries of traffic state channels and residual sub-channels remain stable throughout the entire processing flow, thereby facilitating the bidirectional gating attention module to perform gating modulation on the two semantically clear groups of channels respectively.

[0082] The bidirectional gating attention module includes two branches: anomaly-to-traffic state gating branch and traffic state-to-anomaly gating branch, thereby realizing bidirectional interaction between traffic state parameter features and residual abnormal features.

[0083] For the abnormal-to-traffic-state gated branch, the abnormal channel group is convolutionally encoded and then mapped to the (0,1) interval by the sigmoid function. Then, it is added to the gate bias parameter and divided by two to obtain the first gate weight, as shown in formula (14):

[0084] Where A is the abnormal channel group described in formula (13); (·) represents the convolutional coding operation of the gated branch from abnormality to traffic state; σ(·) represents the sigmoid activation function, which maps the input to the (0,1) interval; This represents the gating bias parameter for the gated branch from the abnormal to the traffic state. This represents the first gating weight generated by the gating branch from the abnormal to the traffic state.

[0085] The gate bias parameters The initial value is set to a preset constant (e.g., 1). Since the output of the sigmoid function is in the (0,1) interval, when the initial value of the gating bias parameter is 1, the first gating weight... The initial value of is in the range (0.5,1) and biased towards 1, which makes the subsequent gating modulation operation close to the identity mapping in the initial stage of training, thereby avoiding the gating mechanism from damaging the original representation across channels in the early stage of training and making the training process more stable.

[0086] Subsequently, the traffic state channel group is modulated element by element using the first gating weight to obtain the modulated traffic state channel group, as shown in formula (15):

[0087] Wherein, P is the traffic state channel group described in formula (13); The first gate weight obtained by formula (14); the symbol ⊙ represents element-wise multiplication; This represents the traffic state channel group after the first gating weight modulation.

[0088] For traffic state to abnormal gated branches, a symmetrical approach is adopted. The traffic state channel group is convolutionally encoded, then mapped by the sigmoid function and added to the gate bias parameter and divided by two to obtain the second gate weight, as shown in formula (16):

[0089] Where P represents the traffic state channel group; (·) represents the convolutional coding operation from traffic state to the abnormal gated branch; σ(·) represents the sigmoid activation function; The gating bias parameter representing the traffic condition to the abnormal gated branch is also set to a preset constant. This represents the second gating weight generated from the traffic condition to the abnormal gating branch.

[0090] Subsequently, the abnormal channel group is modulated element by element using the second gating weight to obtain the modulated abnormal channel group, as shown in formula (17):

[0091] Where A represents the abnormal channel group; The second gate weight is obtained by formula (16); the symbol ⊙ represents element-wise multiplication. This indicates the abnormal channel group after being modulated by the second gating weight.

[0092] Through the bidirectional interaction between the anomaly-to-traffic state gating branch and the traffic state-to-anomaly gating branch, traffic state parameter characteristics and residual anomaly characteristics mutually corroborate and reinforce each other: on the one hand, the anomaly information carried by the anomaly channel group is used to modulate the traffic state channel group, strengthening the anomaly-related parts of the traffic state parameter characteristics; on the other hand, the operational status information carried by the traffic state channel group is used to modulate the anomaly channel group, strengthening the parts of the residual anomaly characteristics related to actual accidents. Compared to the simple splicing or one-way interaction between channels in existing methods, this embodiment fully utilizes the correlation between traffic state parameters and residual anomalies.

[0093] Step 4: Calculate the temporal statistics of the residual sub-channels in the abnormal channel group. When the temporal statistics are lower than a preset threshold, replace the gating weights used to modulate the traffic state channel group in the bidirectional gating attention module with learnable backoff gating weights.

[0094] To address the problem of predicted residual signals degrading or even failing under certain operating conditions, this embodiment introduces a residual missing detection and adaptive backoff mechanism.

[0095] Specifically, the standard deviation along the time dimension is calculated for the specified residual sub-channels in the abnormal channel group (e.g., the first residual sub-channel, i.e., the residual standard score channel), and used as a time-series statistic. When the time-series statistic of a sample is lower than a preset threshold, it indicates that the residual of the sample within the time window is approximately constant, i.e., the predicted residual signal has degraded or failed and cannot provide effective accident discrimination information. Based on this, the residual missing indicator is defined as shown in formula (18):

[0096] in, This refers to a specified residual sub-channel (e.g., a residual standard score channel) within an abnormal channel group. (·) represents the operation of calculating the standard deviation along the time dimension; τ represents the preset threshold; 1[·] represents the indicator function, which takes the value of 1 when the condition in the square brackets is true, and takes the value of 0 otherwise; m represents the residual missing quantity, where m takes the value of 1 to indicate that the predicted residual signal of the sample is degraded, and m takes the value of 0 to indicate that the residual signal is normal.

[0097] When the predicted residual signal is determined to be degraded, the first gating weight generated by the anomaly-to-traffic state gating branch is adaptively switched to a learnable backoff gating weight. Specifically, using the missing indicator as the interpolation coefficient, linear interpolation is performed between the first gating weight and the backoff gating weight to obtain the adaptive backoff gating weight, as shown in formula (19):

[0098] Where m is the residual missing indicator obtained by formula (18); The first gate weight obtained by formula (14); Denotes a learnable parameter; σ(·) represents the sigmoid activation function. This refers to the learnable backoff gating weight; the symbol ⊙ represents element-wise multiplication. This represents the gating weight after adaptive backoff.

[0099] As can be seen from formula (19), when the missing indicator m is 0 (i.e. the residual signal is normal), the gating weight after adaptive backoff is equal to the original first gating weight. The model utilizes residual information normally; when the missing indicator m is 1 (i.e., the residual signal degrades), the gated weights after adaptive backoff are switched to learnable backoff gated weights. The back-off gating weights do not depend on the degraded residual signal, but are automatically learned through training to obtain a reasonable modulation intensity that is independent of the residual.

[0100] Therefore, the model can adaptively reduce its dependence on the residual channel under the condition of residual signal degradation, thereby significantly improving the robustness of the method.

[0101] Accordingly, the calculation of the modulated traffic state channel group shown in formula (15) in step 3 is replaced by modulation with the gated weights after adaptive backoff when the adaptive backoff mechanism of this step is adopted, as shown in formula (20):

[0102] Where P represents the traffic state channel group; The gating weights obtained by adaptive backoff from formula (19); the symbol ⊙ represents element-wise multiplication. This indicates the modulated traffic state channel group when an adaptive backoff mechanism is used.

[0103] Step 5: Concatenate the initial traffic state channel group, the modulated traffic state channel group, and the modulated abnormal channel group, input them into the classification network to obtain the traffic accident identification results, and issue warnings accordingly.

[0104] The traffic state channel group modulated by the bidirectional gating attention module, the original traffic state channel group without modulation, and the abnormal channel group modulated are spliced ​​along the channel dimension and input into the classification network backbone for feature extraction and mapping. The output is the traffic accident identification result that represents whether the sample belongs to the accident category or the normal category.

[0105] The features modulated by the bidirectional gated attention module are fused, and the traffic accident recognition result is output through the classification network backbone. The classification network backbone can be any one or a combination of gated multilayer perceptron networks, temporal convolutional networks, one-dimensional convolutional neural networks, residual networks, long short-term memory networks, gated recurrent units, and Transformer encoders based on self-attention mechanisms. The specific type of classification network backbone does not constitute a limitation of the present invention; this embodiment preferably adopts the following structure based on a gated multilayer perceptron.

[0106] Specifically, the modulated traffic state channel group, the unmodulated original traffic state channel group, and the modulated abnormal channel group are spliced ​​along the channel dimension to obtain the fused features, as shown in formula (21):

[0107] in, P represents the traffic state channel group modulated by the bidirectional gating attention module; P represents the original traffic state channel group without modulation. represents the abnormal channel group modulated by the bidirectional gating attention module; Concat(·) represents the concatenation operation along the channel dimension; F represents the fusion feature of the input classification network backbone.

[0108] The original traffic state channel group P, which is not modulated, is retained in the fusion feature F. Its purpose is that even if the modulation of the bidirectional gated attention module deviates, the original traffic state channel group can still serve as an identity bypass to provide the classification network backbone with unchanged traffic state parameter information, thereby further enhancing the stability of the method.

[0109] Subsequently, the fused feature F is input into the classification network backbone for feature extraction and mapping. In this embodiment, the classification network backbone can adopt a structure including an embedding layer, several gated multilayer perceptron blocks, a normalization layer, and a classification head connected in sequence. The fused feature F is first mapped to the model hidden dimension through the embedding layer, then sequentially passed through several gated multilayer perceptron blocks for interactive modeling of the time dimension and feature dimension, then processed by the normalization layer and aggregated along the time dimension, and finally mapped by the classification head to the prediction scores of each category. After the output of the classification network backbone is normalized by the softmax function, the probability distribution of the sample belonging to the accident class and the normal class is obtained, as shown in formula (22):

[0110] Where F is the fusion feature obtained by formula (21); Backbone(·) represents the feature extraction and mapping operation performed by the backbone of the classification network; softmax(·) represents the normalization exponential function, which normalizes the output of the backbone of the classification network into a probability distribution; This indicates the predicted probability distribution of whether the sample belongs to the accident or normal category.

[0111] As another optional implementation, the backbone of the classification network can also adopt a phase pooling structure, that is, the fused feature F is divided into three stages along the time dimension: the pre-accident stage, the accident stage, and the post-accident stage. Mean pooling and maximum pooling are performed on each stage respectively, and standard deviation pooling is additionally performed on the accident stage. The pooling results of each stage are concatenated and input into the classification head for classification, thereby explicitly utilizing the stage structure of traffic accidents in the time dimension.

[0112] Finally, based on the probability distribution The category with the largest value is selected to obtain the traffic accident identification result for the sample, that is, to determine whether the sample belongs to the accident category or the normal category. Based on the identification result, corresponding traffic early warning mechanisms can be triggered, such as reporting accident warning information to the monitoring center and issuing prompts on variable message signs, so as to achieve rapid detection and proactive response to traffic accidents.

[0113] In this embodiment, the bidirectional gating attention module and the classification network backbone are jointly trained in an end-to-end manner. The training samples are constructed from historical traffic data with accident labels.

[0114] As a preferred implementation, R-Drop regularization is used during the training phase to alleviate overfitting. Specifically, the same input sample is subjected to two forward propagations. Due to the randomness of the model, such as random deactivation, the two forward propagations yield different first and second prediction distributions. In addition to the classification loss, the training loss function also includes the bidirectional KL divergence between the first and second prediction distributions as a regularization term, as shown in formula (23).

[0115] in, and denoted as the first and second prediction distributions obtained by performing two forward propagations on the same input sample, respectively; y represents the true accident label of the sample. (·,·) denotes the cross-entropy classification loss function; α denotes the weight coefficient of the regularization term; The coefficient represents the two-way KL divergence between the first and second predicted distributions. Used to average the classification losses of the two items within the parentheses; L represents the total loss function used for model training.

[0116] The two-way KL divergence The calculation method is to average the KL divergence in two directions, as shown in formula (24):

[0117] in, and These represent the first and second prediction distributions, respectively; KL(·‖·) denotes the KL divergence operation, where the symbol ‖ is used to separate the two distribution parameters of the KL divergence. This represents the average of the KL divergences in the two directions mentioned above, i.e., the two-way KL divergence between the first and second prediction distributions.

[0118] Furthermore, to address the class imbalance issue in traffic accident data, where the number of accident-class samples is typically far less than that of normal-class samples, class weights can be introduced into the cross-entropy classification loss, coupled with label smoothing techniques, to improve the model's ability to identify accident-class samples. During training, an adaptive moment estimation optimizer can be employed, along with a learning rate scheduling strategy combining learning rate warm-up and cosine annealing. Data augmentation operations such as random zeroing can also be applied to abnormal channel groups during training to improve the model's robustness under residual missing conditions. This random zeroing operation, combined with the residual missing detection mechanism in step 4, allows the model to observe residual degradation during the training phase.

[0119] Example 2 Based on the method described in Embodiment 1, the purpose of this embodiment is to provide a traffic accident early warning system based on prediction residuals and bidirectional attention, including: The data acquisition and residual generation module is used to acquire historical time-series data of various traffic state parameters of each monitoring node in the traffic monitoring road network, and use the prediction model to predict the historical time-series data and obtain the prediction residual. The residual feature construction and fusion module is used to perform node adaptive normalization on the predicted residual to obtain the residual standard score, decompose it into multiple residual sub-channels, and splice it with the traffic state channels corresponding to various traffic state parameters to form a multi-channel fusion input tensor. A bidirectional gated attention modulation module is used to divide the multi-channel fused input tensor into traffic state channel group and abnormal channel group, and obtain the interactively modulated traffic state channel group and abnormal channel group through the bidirectional gated attention module; The residual degradation detection and backoff module is used to calculate the temporal statistics of the residual sub-channels in the abnormal channel group. When the temporal statistics are lower than a preset threshold, the learningable backoff gating weights are used to replace the gating weights in the bidirectional gating attention module used to modulate the traffic state channel group. The feature splicing and classification early warning module is used to splice the initial traffic state channel group, the modulated traffic state channel group, and the modulated abnormal channel group, input them into the classification network to obtain traffic accident identification results, and issue early warnings based on them.

[0120] The steps and methods involved in the apparatus of the above embodiments correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0121] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0122] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A traffic accident early warning method based on prediction residuals and bidirectional attention, characterized in that, include: Historical time-series data of various traffic state parameters of each monitoring node in the traffic monitoring road network are obtained, and prediction models are used to predict the historical time-series data and obtain prediction residuals. The predicted residual is subjected to node adaptive normalization to obtain the residual standard score, which is then decomposed into multiple residual sub-channels and spliced ​​with the traffic state channels corresponding to various traffic state parameters to form a multi-channel fusion input tensor. The multi-channel fusion input tensor is divided into traffic state channel group and abnormal channel group, and the interactively modulated traffic state channel group and abnormal channel group are obtained through a bidirectional gating attention module. Calculate the temporal statistics of the residual sub-channels in the abnormal channel group. When the temporal statistics are lower than a preset threshold, replace the gating weights used to modulate the traffic state channel group in the bidirectional gating attention module with learnable backoff gating weights. The initial traffic state channel group, the modulated traffic state channel group, and the modulated abnormal channel group are concatenated and input into the classification network to obtain traffic accident identification results and issue warnings accordingly.

2. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 1, characterized in that, The prediction model is a dual-path temporal prediction model, which includes a parallel Mamba branch and a temporal convolutional network branch, as well as a gated fusion unit for adaptive weighted fusion of the outputs of the two branches.

3. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 1, characterized in that, The prediction head of the prediction model has a dual-head structure, including a point prediction head and an uncertainty prediction head. The specific expression for the standardized prediction residual is as follows: in, Indicates the actual observed value; This indicates the point prediction value given by the point prediction header; The uncertainty is the prediction uncertainty output by the uncertainty prediction head corresponding to the predicted value of the point, and its value is always positive; ε represents a preset positive number to avoid the denominator being zero. This represents the standardized prediction residual.

4. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 1, characterized in that, The node adaptive normalization includes: using the median of the predicted residuals of the monitoring nodes as the location parameter and the median absolute deviation as the scale parameter, standardizing and truncating the predicted residuals to obtain the standard residual score.

5. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 1, characterized in that, The plurality of residual sub-channels include: residual standard score, residual absolute value, residual difference, and residual moving average; wherein, the residual absolute value is the element-wise absolute value of the residual standard score, the residual difference is the first-order difference of the residual standard score along the time dimension, and the residual moving average is the result of the residual standard score smoothed by a sliding window.

6. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 1, characterized in that, The bidirectional gating attention module includes two branches: an anomaly-to-traffic state gating branch and a traffic state-to-anomaly gating branch.

7. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 6, characterized in that, The abnormal-to-traffic-state gating branch encodes the abnormal channel group via convolution, then maps it to the (0,1) interval using a sigmoid function. This gating weight is then added to the gating bias parameter and divided by two to obtain the first gating weight. Wherein, A is the abnormal channel group; (·) represents the convolutional coding operation of the gated branch from abnormality to traffic state; σ(·) represents the sigmoid activation function, which maps the input to the (0,1) interval; The gating bias parameter represents the gating branch from abnormal to traffic state, and its initial value is a preset constant; This represents the first gating weight generated by the gating branch from the abnormal to the traffic state.

8. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 1, characterized in that, The time-series statistics are the standard deviations of specified residual sub-channels along the time dimension in the abnormal channel group.

9. The traffic accident early warning method based on prediction residuals and bidirectional attention as described in claim 1, characterized in that, The bidirectional gating attention module and the classification network are trained using R-Drop regularization during the training phase. The same input sample is forward-propagated twice to obtain a first prediction distribution and a second prediction distribution, respectively. In addition to the classification loss, the bidirectional KL divergence between the first prediction distribution and the second prediction distribution is superimposed as a regularization term in the loss function.

10. A traffic accident early warning system based on prediction residuals and bidirectional attention, characterized in that, include: The data acquisition and residual generation module is used to acquire historical time-series data of various traffic state parameters of each monitoring node in the traffic monitoring road network, and use the prediction model to predict the historical time-series data and obtain the prediction residual. The residual feature construction and fusion module is used to perform node adaptive normalization on the predicted residual to obtain the residual standard score, decompose it into multiple residual sub-channels, and splice it with the traffic state channels corresponding to various traffic state parameters to form a multi-channel fusion input tensor. A bidirectional gated attention modulation module is used to divide the multi-channel fused input tensor into traffic state channel group and abnormal channel group, and obtain the interactively modulated traffic state channel group and abnormal channel group through the bidirectional gated attention module; The residual degradation detection and backoff module is used to calculate the temporal statistics of the residual sub-channels in the abnormal channel group. When the temporal statistics are lower than a preset threshold, the learningable backoff gating weights are used to replace the gating weights in the bidirectional gating attention module used to modulate the traffic state channel group. The feature splicing and classification early warning module is used to splice the initial traffic state channel group, the modulated traffic state channel group, and the modulated abnormal channel group, input them into the classification network to obtain traffic accident identification results, and issue early warnings based on them.