Highway network risk identification method based on time-space diagram attention network

By using a spatiotemporal graph attention network-based approach and combining multi-source data for highway network risk identification, the problem of false alarms and missed alarms caused by a single data source is solved. This enables accurate identification and prediction of risk spread paths, improving the accuracy of risk identification and early warning capabilities.

CN121743970APending Publication Date: 2026-03-27HEBEI PROVINCIAL COMM PLANNING & DESIGN INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies for identifying traffic risks on highways rely on a single data source, leading to false alarms or missed alarms. They also struggle to identify risk propagation paths and transmission mechanisms, and lack modeling of road network topology, thus affecting the accuracy and predictive ability of risk identification.

Method used

A spatiotemporal graph attention network-based approach is adopted, combining radar detector data, gantry/toll station transaction data, and meteorological data. Through data preprocessing, road network map construction, modal feature extraction, and spatiotemporal graph attention network modeling, quality labels and spatiotemporal interpolation mechanisms are introduced to enhance robustness to local data missing and anomalies and improve the identifiability of risk propagation links across road segments.

Benefits of technology

It enables continuous perception of risks in the highway network and dynamic tracking of evolving links, improving the accuracy of risk identification and early warning lead time, supporting online risk assessment and evolution detection, and enhancing the proactive decision-making capabilities of traffic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743970A_ABST
    Figure CN121743970A_ABST
Patent Text Reader

Abstract

The invention discloses a highway network risk identification method based on a space-time diagram attention network. The method comprises the steps of S1, data acquisition; s2, data preprocessing and quality diagnosis; s3, constructing a road network map; s4, modal feature extraction; s5, constructing a space-time diagram attention network; s6, risk scoring and evolution analysis; and S7, outputting and storing a result. The method comprises the steps of data preprocessing and quality diagnosis, road network-based graph construction, modal feature extraction, space-time diagram attention network modeling, expression-based risk scoring and evolution analysis and the like, a quality label weighting mechanism and graph-based space-time interpolation are introduced, and robustness to local data missing and abnormity is enhanced; through combined use of a charge transaction true value and a weather correction factor, the identifiability of a cross-road-section risk propagation link is improved; online risk assessment and evolution detection are supported, timely intervention of a traffic management department when a risk occurs is facilitated, and the active decision-making capability of road safety operation is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of road traffic technology, specifically to a method for risk identification of highway networks based on spatiotemporal graph attention networks. Background Technology

[0002] With the increasing informatization of highway traffic, the focus of research and engineering has shifted from relying on traditional structural upgrades (such as buffer zones, guardrails, and ramp optimization) to real-time situation prediction and risk identification. Existing road traffic operation risk identification schemes include analyzing historical accident statistics to conduct offline "black spot analysis" and "network screening" of the entire network or individual road segments to identify high-risk road sections; or using a single data source, including data from highway surveillance cameras or unmanned aerial vehicles (UAVs), based on traditional time series models (such as ARIMA) and deep learning networks (such as LSTM) to conduct real-time traffic flow monitoring or conflict early warning.

[0003] Based on the above analysis of existing technologies, several key pain points still exist in engineering applications, affecting the feasibility and accuracy of the methods and resulting in low economic and social benefits. First, data detection methods rely on a single data source, leading to problems such as sparse sensor deployment, environmental interference, and risk blind spots. These problems often cause false alarms or missed alarms. The data space is sparse and discontinuous across sections, making it difficult to form a continuous macro-level situation based on data from a single section, and also making it difficult to identify risk diffusion paths, thus affecting the accurate identification of risk evolution. Second, time-series analysis and prediction methods can only handle risk change relationships at a single section or dimension, lacking modeling of road network topological spatial relationships. Existing methods struggle to accurately identify and predict the transmission and diffusion paths of macro-level traffic risks on the network, failing to provide risk identification results for traffic managers, and consequently, making it difficult to prevent potential hazards.

[0004] The unique advantage of spatiotemporal graph attention networks lies in their ability to iteratively update information from the neighborhood of nodes and their sensitivity to connections within spatial structures. Highway networks, characterized by strong closure and regular topological structures, can be leveraged by defining the structural features of highways and interchanges as a graph and fusing multi-source information to achieve complementary traffic data, thereby outputting traffic situation and risk evolution results. However, current methods are still rarely implemented in engineering applications.

[0005] Therefore, it is necessary to research and develop a highway network risk identification method based on spatiotemporal graph attention network to solve the above problems. Summary of the Invention

[0006] This invention provides a highway network risk identification method based on spatiotemporal graph attention networks. Using radar detector data, gantry / toll station transaction data, and meteorological data as core inputs, the method includes steps such as data preprocessing and quality diagnosis, road network-based graph construction, modal feature extraction, spatiotemporal graph attention network modeling, and representation-based risk scoring and evolution analysis. It introduces a quality label weighting mechanism and graph-based spatiotemporal interpolation to enhance robustness to local data gaps and anomalies. Furthermore, by jointly using toll transaction ground truth and weather correction factors, the method improves the identifiability of cross-segment risk propagation links.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] A method for risk identification of highway networks based on spatiotemporal graph attention networks includes the following steps:

[0009] S1. Data Acquisition: Collect structured data including radar detector data, gantry / toll station transaction data and meteorological data, and align and serialize the various types of data in the structured data according to a unified time base.

[0010] S2. Data Preprocessing and Quality Diagnosis: The collected data is processed to unify the units and sampling frequency, and missing and abnormal data are detected and repaired. Multi-class data is matched spatiotemporally and consistently, and a quality label is generated for each data record.

[0011] S3. Road network map construction: Based on the research road segments, construct a regional road network structure map containing multiple nodes, map adjacent road segments and weaving / merging relationships as graph edges, and determine edge weights based on real-time traffic characteristics and external meteorological conditions.

[0012] S4. Modal feature extraction: Extract the features of nodes and edges from the preprocessed data. The features include at least speed, flow, occupancy rate, segment travel time residual and weather factors.

[0013] S5. Spatiotemporal Graph Attention Network Construction: Using graph attention mechanism and temporal aggregation technology, a spatiotemporal graph attention network is constructed. Within a sliding time window, the spatiotemporal dynamic representation of nodes and edges is learned. The regional road network structure map constructed in S3 and the features in S4 are input to generate and output the situation vector at the road segment level.

[0014] S6. Risk Scoring and Evolution Analysis: Based on the situation vector output by the spatiotemporal graph attention network, calculate the segment-level risk score sequence, calculate the segment-level risk score, analyze the dynamic evolution of the risk score, and identify the temporal and spatial distribution of risk events.

[0015] S7. Results Output and Storage: Output and display the results of risk evolution analysis in the form of time series and spatial heatmaps, and store them for subsequent visualization and offline analysis.

[0016] Preferably, in S1, the structured data collected during the data collection and integration process includes:

[0017] Radar detector data includes fields such as timestamp, detector ID, lane number, traffic flow, average vehicle speed, occupancy rate, average headway, and sampling period.

[0018] Gantry / Toll Station Transaction Data: Includes transaction time, gantry / toll station ID (as a unique identifier, providing the longitude and latitude of each gantry / toll station as location information), vehicle unique identification number, vehicle category, lane number, and vehicle's entry / exit direction.

[0019] Meteorological data includes timestamps, meteorological observation point IDs, visibility, precipitation, road surface temperature, and road surface conditions.

[0020] Preferably, in S2, the data preprocessing and quality diagnosis process is as follows:

[0021] S2.1 Unit Normalization: The collected data of various types is input by the data access module. The units of the input fields are unified, and time calibration and sorting are performed. The timestamps of various types of data are unified and matched to the corresponding nodes in the graph in space.

[0022] S2.2 Unified sampling frequency: The normalized data is time-aligned according to a unified sampling period to ensure that the sampling frequency of various data sources is consistent.

[0023] S2.3 Data Quality Assessment and Diagnosis: Robust statistical methods are used to identify outliers in continuous quantitative data; topological consistency verification is used to diagnose consistency anomalies in cross-section velocity data; when a velocity change occurs in an adjacent section within a short period of time, and the velocity change is not accompanied by a corresponding change in flow data, the velocity data of that cross-section is marked as abnormal data.

[0024] S2.4 Data Repair and Missing Data Imputation: For gantry / toll station transactions, a travel time threshold test is used to remove obviously erroneous transactions and eliminate the interference of extreme values ​​on data quality.

[0025] Determine whether it is necessary to repair missing outliers. If necessary, use spatiotemporal interpolation to integrate historical databases for repair. After repair, proceed to the next step for the generated samples. If not necessary, proceed directly to the next step.

[0026] S2.5 Quality Label Generation: For each data record in the dataset, a quality level label is generated. The quality level labels are divided into three categories: high, medium, and low. During the model input stage of the graph attention network, for records carrying low quality level labels, the corresponding input features are processed in the following ways: the weight contribution of the corresponding input features in the graph attention network is reduced, or the input features are masked.

[0027] Preferably, in S2, the quality label The definition is as follows:

[0028]

[0029] in: This represents a missing mask, where 1 represents missing data and 0 represents complete data. This represents the consistency score with neighboring nodes / historical sequences. This indicates the stability index of continuous multi-period fluctuations. Represents the weighting coefficients, satisfying .

[0030] Preferably, in S3, the mapping rules between nodes and edges during the road network map construction process are as follows:

[0031] The locations where radar detector data and gantry / toll station transaction data are acquired, as well as the starting and ending points of roads, are mapped to nodes in the road network map; adjacent road segments and traffic weaving / merging areas are mapped to edges in the map, with edge weights determined by factors including real-time average speed, number of lanes, road segment length, and weather correction factors.

[0032] Preferably, in S4, the selection criteria for features during modal feature extraction include:

[0033] The magnitude of the influence factor for each feature is calculated using a normalization function.

[0034] The correlation between features is assessed using analysis of covariance.

[0035] Construct a feature scoring system that comprehensively considers influence and relevance, and prioritize features with high scores.

[0036] Preferably, in S5, the spatiotemporal graph attention network construction process is as follows:

[0037] S5.1 Constructing the graph structure: Layout nodes, edges and their characteristics to generate a road network structure map of the highway research area.

[0038] S5.2 Feature Extraction and Aggregation: Utilize graph attention mechanism to aggregate and update node information, and train a risk model through supervised learning to extract high-dimensional features of graph nodes.

[0039] S5.3 Spatiotemporal Update: Within a sliding time window, the representation of nodes is spatiotemporally aggregated to reflect the local traffic status and historical evolution trend at the current moment.

[0040] S5.4 Output Layer and Risk Scoring: Calculate road segment-level risk scores, perform risk surge detection and trend analysis, and mark the time and spatial scope of risk events.

[0041] The technical effects and advantages provided by the present invention in the above technical solution are as follows:

[0042] 1. Using radar detector data, gantry / toll station transaction data, and meteorological data as core inputs, the method includes data preprocessing and quality diagnosis, road network-based graph construction, modal feature extraction, spatiotemporal graph attention network modeling, and representation-based risk scoring and evolution analysis. It further introduces a quality label weighting mechanism and graph-based spatiotemporal interpolation to enhance robustness against local data gaps and anomalies. By combining toll transaction ground truth with weather correction factors, the identifiability of cross-segment risk propagation links is improved. Compared to traditional methods relying on single-section perception and prediction, graph neural networks can more accurately capture the propagation effects and structural risk diffusion mechanisms between adjacent sections, thereby significantly improving the accuracy of risk identification and the lead time for early warning.

[0043] 2. By integrating the inherent topology of the road network with multi-source dynamic data and using graph neural networks for joint spatiotemporal modeling, continuous perception of macro-level road network traffic operation risks and dynamic tracking of evolution links can be achieved. This method maintains the continuity of the macro-level road network situation and improves the ability to accurately identify and predict risk diffusion even under conditions of data sparsity and external disturbances. This method supports online risk assessment and evolution detection, which helps traffic management departments to intervene in a timely manner when risks occur and enhances their proactive decision-making ability for safe road operation. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 The flowchart illustrates the operation of the highway network risk identification method provided by this invention.

[0046] Figure 2 This is a simplified example of the road network structure diagram for a highway research area.

[0047] Figure 3 This is a schematic diagram of the data preprocessing and quality control module operation.

[0048] Figure 4 This is a schematic diagram of the training and inference framework structure of a spatiotemporal graph attention network.

[0049] Figure 5 This is a thermal distribution map of a high-risk event on a highway before control measures are implemented, as shown in this embodiment of the invention.

[0050] Figure 6 This is a thermal distribution map of a high-risk event on a highway after control measures were implemented, as shown in this embodiment of the invention. Detailed Implementation

[0051] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below.

[0052] like Figures 1 to 4 As shown, this invention provides a method for identifying risks in highway networks based on spatiotemporal graph attention networks, comprising the following steps:

[0053] S1. Data Acquisition: Collect structured data including radar detector data, gantry / toll station transaction data and meteorological data, and align and serialize the various types of data in the structured data according to a unified time base.

[0054] The structured data collected during the data collection and integration process includes:

[0055] Radar detector data includes timestamp (i.e., data acquisition time), detector ID (as a unique identifier, providing the longitude and latitude of each detector as positioning information), lane number, traffic flow, average vehicle speed, occupancy rate, average headway, and sampling period.

[0056] Gantry / toll station transaction data includes transaction time (i.e., vehicle passage time), gantry / toll station ID (as a unique identifier, providing the longitude and latitude of each gantry / toll station as location information), vehicle unique identification number, vehicle category, lane number, and vehicle entry / exit direction; used to calculate segment travel time.

[0057] Meteorological data includes timestamps, meteorological observation point IDs (i.e., unique identifiers for meteorological stations / roadside sensors, providing longitude and latitude as location information for each observation point), visibility, precipitation, road surface temperature, and road surface condition. Correction factors are used to map these data to edge weights.

[0058] Radar detector data provides high-frequency local information for situation characterization and is the basis for macro-situation projection; gantry / toll station transaction data can provide information on segment travel time and cross-section connectivity, which can make up for the spatiotemporal gaps caused by the sparseness of detectors; meteorological data is an external disturbance factor and can be used as a risk amplification or propagation correction factor.

[0059] The fields / specifications of the three types of data input in this embodiment—radar detector data (obtained through radar detectors), gantry / toll station transaction data, and meteorological data—are shown in Table 1.

[0060] Table 1. Input Multi-Source Data Fields Table

[0061]

[0062] Sample data for different risk scenarios is required, including but not limited to data samples collected in low-light conditions at night, under severe weather conditions, in areas affected by traffic accidents, in heavily congested scenarios, and baseline samples under normal traffic conditions. The proportion of samples from each scenario should be balanced to ensure the model's generalization ability and to meet the requirements shown in Table 2 as much as possible.

[0063] Table 2. Proportion of Sample Data for Different Risk Scenarios

[0064]

[0065] The definition of severe weather can be based on the classification standards of the China Meteorological Administration, including heavy snow, freezing, low temperature, strong wind, high temperature and heat, heavy rainfall, and continuous rainfall.

[0066] S2. Data Preprocessing and Quality Diagnosis: The collected data is processed to unify the units and sampling frequency, and missing and abnormal data are detected and repaired. Multi-class data is matched spatiotemporally and consistently, and a quality label is generated for each data record.

[0067] The data preprocessing and quality diagnosis process is as follows:

[0068] S2.1 Unit Normalization: The collected data of various types is input into the data access module. Units are standardized for input fields, such as speed being standardized to km / h or m / s, and time calibration and sorting are performed. This ensures that all data sources are synchronized and serialized according to a unified time base, facilitating subsequent processing. Specifically, timestamps for various data types are standardized and spatially matched to corresponding nodes in the graph. If the data requires special semantic transformations, corresponding transformation rules are further set to modify the required fields, ensuring consistency of all data in time and space. The standardized data is then transmitted back to the historical database and output to the next module.

[0069] S2.2 Unified Sampling Frequency: The normalized data is time-aligned according to a unified sampling period to ensure consistent sampling frequencies across various data sources. In specific applications, the data sampling frequency can be adjusted as needed to ensure consistency of the time series.

[0070] S2.3 Data Quality Assessment and Diagnosis: For continuous quantities (such as continuous quantitative indicators such as speed and flow time series data of a single cross section), outliers in the dataset are identified by robust statistical methods (such as median absolute deviation, robust regression, etc.) to eliminate the interference of extreme values ​​on data quality.

[0071] For velocity data spanning multiple monitoring sections, the focus is on verifying its topological consistency: based on the spatial topological relationship between sections (such as upstream and downstream connections and spacing distribution), it is determined whether the velocity data of different sections conforms to reasonable transmission logic, and anomalies of inconsistent velocities across sections are identified.

[0072] For velocity data of adjacent monitoring sections, a dual judgment condition is set: if a significant change occurs in the velocity of adjacent sections within the same time window (short time scale), and the flow data of the corresponding section does not change in a matching manner within the time interval corresponding to the change (i.e., the flow does not fluctuate synchronously), then the cross-section velocity data under this condition is judged as abnormal and marked.

[0073] S2.4 Data Repair and Missing Data Imputation: Use travel time threshold checks to remove obviously erroneous transactions from gantry / toll station transaction data and eliminate the interference of extreme values ​​on data quality.

[0074] Determine whether it is necessary to repair missing outliers. If necessary, use spatiotemporal interpolation methods to integrate the historical database for repair, such as minimizing the problem by combining graph Laplacian regularization with time smoothing, to ensure that the interpolation results are consistent with neighboring nodes and historical sequences. After repair, proceed to the next step for the generated samples. If it is not necessary, proceed directly to the next step.

[0075] S2.5 Quality Label Generation: For each independent data record in the dataset, a corresponding quality level label is generated based on the preset quality evaluation criteria (which can be combined with the data quality assessment results mentioned above, such as outliers, consistency anomalies, etc.). The quality level labels are uniformly divided into three categories: "high", "medium" and "low", and each record corresponds to only one category of quality label.

[0076] Model Input Feature Processing: During the model input phase of the graph attention network, records marked as "low-quality" undergo targeted processing of their input features used for model training / inference. This involves appropriately reducing the contribution of the corresponding features of the low-quality record to the attention calculation in the graph attention network through a weight adjustment mechanism, or directly masking the input features corresponding to the low-quality record (i.e., blocking the effective input of that feature) to reduce the interference of low-quality data on the model's performance. This reduces noise propagation.

[0077] Quality Label The definition is as follows:

[0078]

[0079] in: This represents a missing mask, where 1 represents missing data and 0 represents complete data. This represents the consistency score with neighboring nodes / historical sequences, ranging from 0 to 1; This is an indicator of the stability of continuous multi-period fluctuations; the smaller the fluctuation, the higher the score. Represents the weighting coefficients, satisfying . The value is set by human definition.

[0080] Substituting the above parameters into the formula and calculating, the result is normalized to obtain the mass weight. This is used to adjust feature contributions during subsequent spatiotemporal aggregation. Specifically, in the subsequent step S5, features of low-quality records are weighted / masked during modeling to adjust the contribution weights of node features.

[0081] S3. Road Network Graph Construction: Based on the research road segments, define relevant issues and construct a structural graph of the highway research area road network containing multiple nodes; the mapping rules between nodes and edges are as follows:

[0082] The locations where radar detector data and gantry / toll station transaction data are acquired, as well as the road start and end points, are mapped to nodes in the road network map. Adjacent road segments and traffic weaving / merging areas are mapped to edges in the map. Edge weights are determined by factors including real-time average speed, number of lanes, road segment length, and weather correction factors. For example, edge weights can be determined by a combination of real-time average speed, number of lanes, road segment length, and weather correction factors.

[0083] This embodiment uses To represent a traffic network diagram, where: Represents a set of traffic nodes. Represents a set of features. This represents the set of connected edges.

[0084] Node construction: When mapping graph nodes, to accurately distinguish the different node facilities and data characteristics, the node set U is further decomposed into several disjoint type subsets:

[0085]

[0086] in: Indicates a radar detector node (radar, coil, microwave); Indicates the gantry / toll station node; Indicates the starting point, ending point, or exit node of a road or ramp.

[0087] Each node The features are represented using a structure of "general sub-vector + type-specific sub-vector":

[0088]

[0089] in: It represents common features (such as node type, geographic coordinates, etc.). Indicates node type Its unique features (see S4 for details).

[0090] Edge construction: A set E of edges, including unidirectional and bidirectional edges, is mapped to physically adjacent road segments, ramps, and weaving / merging / separating locations. Its properties are strictly mapped to the actual road network structure by the adjacency matrix A. The adjacency matrix... The definition is as follows:

[0091]

[0092] The eigenvector of edge (i,j) at time t can be further expressed as: .

[0093] Based on actual road network data, one-way and two-way road sections can be further divided into one-way edges and two-way edges:

[0094] Traffic flow direction on one-way road sections is only allowed to be in the following directions: ,Right now .

[0095] Traffic flow direction is permitted on two-way road sections and ,Right now .

[0096] High-speed mainline: One-way or two-way connections are established through continuous radar detectors and gantry / toll station nodes.

[0097] Ramp / Interchange: A one-way edge is established between corresponding nodes to indicate the merging or diverging direction.

[0098] Toll station / terminal: connects upstream and downstream nodes and reflects import and export constraints.

[0099] Based on the above definition, a highway network can be abstracted as a spatial graph, where nodes and edges all have real-world physical meaning. An example of the graph structure is shown below. Figure 2 .

[0100] When constructing the graph structure, edge weights (i.e., edge weights) are determined based on real-time traffic characteristics and external weather conditions.

[0101] S4. Modal feature extraction: Extract the features of nodes and edges from the preprocessed data. The features include at least speed, flow, occupancy rate, segment travel time residual and weather factors.

[0102] Node features are defined as follows:

[0103] It has high-frequency observation fields, such as speed, flow rate, occupancy rate, instantaneous headway, etc., and is represented as ; The main information provided is the segment travel time and throughput through single-vehicle level transaction timestamps / entry / exit records, represented as... ; It mainly includes attributes such as connectivity, directionality, and statistical throughput, and is represented as And the latter two types of nodes typically do not directly possess instantaneous aggregated fields such as speed / occupancy.

[0104] Edge features consist of the union of three feature sets:

[0105]

[0106] Static characteristics of roads This includes information such as the number of lanes, curve curvature, and road segment length (derived from highway network design drawings).

[0107] Static characteristics of roads .

[0108] in, This includes information such as the number of lanes, curve curvature, and road segment length (derived from highway network design drawings).

[0109] Real-time dynamic features .

[0110] in, This includes speed and flow determined by the dynamic characteristics of nodes; travel time estimation; speed dispersion; and speed difference between adjacent lanes.

[0111] Weather mapping factor .

[0112] These represent indicators such as visibility, precipitation, road surface temperature, and road surface condition. Weather mapping factors. Defined as mapping weather fields to function values ​​that affect the intensity of risk propagation: .

[0113] in, It is a learnable or pre-defined monotonic mapping function.

[0114] Perform normalization on each node and edge feature and output quality weights.

[0115] In multi-source traffic data modeling, the number of initial candidate features for connecting edges is usually large, but not all features contribute equally to the risk identification task. On the one hand, some features have low correlation with risk indicators, and using them as input not only fails to improve identification performance but may also introduce noise. On the other hand, the computational cost of the subsequent model attention mechanism is very high, so it is necessary to consider the appropriate dimensionality of the model parameters. Some features are highly correlated (high redundancy), and repeated input in graph neural networks will cause information redundancy, increase training complexity, and lead to model overfitting. Therefore, correlation comparison and deduplication mechanisms are introduced to improve the model's risk identification accuracy and reduce the risk of the model being interfered with by noise and redundant features.

[0116] To avoid redundant input and highlight risk contribution, this invention performs a dimension-by-dimensional analysis of edge features and calculates the first... The original weight vector of each candidate feature The definition is as follows:

[0117]

[0118] in: Indicates the first There are 10 candidate features, which belong to the set of edge features; Representing candidate features Relevance to the true risk label; This represents the average quality score of the data; the more complete the data, the higher the score. The average correlation coefficient with other features is used to measure redundancy. This represents a weighting parameter that adjusts the relative importance of risk correlation, data stability, and redundancy penalty.

[0119]

[0120] in: Indicates the first The magnitude of each characteristic influence factor; Indicates the first The original score or weight of each feature; Indicates the first The original score or weight of each feature.

[0121] Calculate the correlation coefficient between the two vectors. For features with highly redundant risk contributions, retain only the more representative terms:

[0122]

[0123] in: , indicating except All external features; This indicates the calculation of features within each detection period. The autovariance; Represents covariance; This indicates the calculation of features within each detection period. The independent variance.

[0124]

[0125] To comprehensively consider relevance and influence factors, a composite score is defined. By combining the two, features with a stronger risk contribution can be screened out.

[0126]

[0127] Will Sort in reverse order and select the first few. One feature is preserved, among which The method for determining the value is as follows:

[0128]

[0129] in Optimize the validation set performance (e.g., maximize the F1 score for risk identification) and select the optimal result. The value is determined as the number of features.

[0130] The selection of edge features adopts a "static-dynamic hierarchical + global sorting" strategy:

[0131] First, in the static feature set Dynamic feature set Independent computing Before each screening Each feature is then combined into a unified candidate set to form the final input vector. .

[0132] This method ensures that geometric design features (such as curve radius and number of lanes) are not completely masked by dynamic traffic flow indicators, while retaining the independent influence of weather factors, thus taking into account both the inherent road conditions and the real-time operating situation.

[0133] S5. Spatiotemporal Graph Attention Network Construction: Using graph attention mechanism and temporal aggregation technology, a spatiotemporal graph attention network is constructed. Within a sliding time window, the spatiotemporal dynamic representation of nodes and edges is learned. The regional road network structure map constructed in S3 and the features in S4 are input to generate and output the situation vector at the road segment level.

[0134] like Figure 4As shown, a series of graph attention layers (GAT) connected to a temporal aggregation layer capture short-term and medium-term dynamics. The input is the graph and corresponding feature sequence within a sliding time window, and the output is the spatiotemporal representation vector of each node / segment. Compared to the original GAT, the key improvement of this invention lies in introducing risk priors and considering data quality weights, so that the attention score not only depends on the feature similarity of neighboring nodes, but is also directly linked to the risk propagation potential. During the training phase, the model is trained using historical sequences (including known risk event annotations); during the inference phase, the output representation is inferred online using a sliding window; quality labels are used for feature weighting or loss weighting during training or inference.

[0135] The specific steps for constructing a spatiotemporal graph attention network are as follows:

[0136] S5.1 Constructing the graph structure: Layout nodes, edges and their characteristics to generate a road network structure map of the highway research area.

[0137] S5.2 Feature Extraction and Aggregation: Utilize graph attention mechanism to aggregate and update node information, and train a risk model through supervised learning to extract high-dimensional features of graph nodes.

[0138] Input data and feature definitions, based on the spatial graph Node features Edge features During input, node and edge features are encoded into vector form.

[0139] In the spatiotemporal graph attention network of this invention, let:

[0140] This is the node feature mapping matrix, used to map the original feature vectors of the nodes. Linear projection onto the hidden feature space:

[0141]

[0142] in, Input the feature dimension for the node. The hidden dimension of the attention layer.

[0143] This is the edge feature mapping matrix, used to map edge features (such as road segment geometric attributes, real-time traffic volume, weather correction factors, etc.). Projected into the same hidden space:

[0144]

[0145] in Input the feature dimension for the edge.

[0146] , All of these are learnable parameter matrices, automatically obtained through model training, without relying on manual settings.

[0147] The graph attention mechanism at each time step For nodes Aggregate Neighbors Features, calculate attention score :

[0148]

[0149] in: This is a learnable attention parameter vector. This represents the transpose of a vector. This represents vector concatenation, which yields the relevance score for each edge. ,in This is the leakage coefficient (typically taken as 0.01–0.2); used to avoid traditional... The problem that the gradient of a function is zero in the negative interval allows the model to maintain a weak response to negative inputs, thus enhancing convergence stability.

[0150] Subsequently passed Functions on the same node within the field The attention score is normalized to obtain the normalization coefficient:

[0151]

[0152] in, Indicates time ,node When aggregating neighbor information, neighbors Normalized attention weights. This represents an exponential function. To ensure robustness, data quality weights are introduced. Low-quality data will be downgraded in weighting.

[0153] In the subsequent spatiotemporal aggregation, nodes It selectively absorbs the temporal features of its neighbors based on attention weights, thereby achieving adaptive learning of the risk propagation intensity.

[0154] S5.3 Spatiotemporal Update: Within a sliding time window, the representation of nodes is spatiotemporally aggregated to reflect the local traffic status and historical evolution trend at the current moment; the spatiotemporal update process is as follows: within a time length of... Within the sliding time window, for nodes The representation is used for spatiotemporal aggregation:

[0155]

[0156]

[0157]

[0158] in: Indicates the end time of the current sliding time window; This is a historical time index, representing each historical sampling point within a time window, and satisfying... ; This represents the time decay weight (exponential decay or learnable parameter), which can be defined here as an exponential distribution that highlights recent data; Represents a node In time eigenvectors; Representing an edge In time eigenvectors; This represents a non-linear activation function, here it is... function. This represents the linear transformation matrix learned during model training, used to transform the features of the input nodes. Projected to dimension A unified feature space; Indicates the feature dimension of the input node; This represents the dimension of the hidden layer after mapping (usually 32, 64, or 128).

[0159] This process achieves two-layer aggregation: a weighted summation of neighbor node information in the spatial dimension; and a weighted summation of historical information within a time window in the temporal dimension. Ultimately... It also reflects the local traffic status of the node at the current moment and its evolution trend over a period of time.

[0160] S5.4 Output Layer and Risk Scoring: Calculate road segment-level risk scores, perform risk surge detection and trend analysis, and mark the time and spatial scope of risk events.

[0161] Updated node representation Input risk identification module, calculate Risk score at any moment:

[0162]

[0163] To ensure that the model training and prediction results are consistent with the actual risk situation, this invention adopts the following annotation and classification process:

[0164] Data fusion: The collected accident records (including the time, location, and type of the accident) are fused with the detector data, gantry / toll station section travel time, and weather data of the corresponding period to construct the spatiotemporal feature vector of that period.

[0165] Manual annotation: Traffic safety experts or data annotation personnel verify the location and time of the accident, mark the period as an "abnormal event" period, and record the event category (accident, congestion, safety hazards caused by severe weather, etc.) and main causes (sudden increase in traffic, decreased road visibility, traffic wave transmission, etc.).

[0166] Risk index calculation: The features obtained by fusing data and annotations are input into the spatiotemporal graph attention network to obtain the risk score of each road segment in that period. For example, recording the accident and detection data, the corresponding predicted risk index is... This indicates that the road section was in an extremely high-risk state before and after the accident.

[0167] Risk score distribution analysis and grading: Statistically analyze the risk score distribution of historical samples and set grading thresholds according to percentile or clustering methods. :

[0168] High risk: (e.g., above the 90th percentile).

[0169] Medium risk: .

[0170] Low risk: .

[0171] This hierarchical structure ensures a one-to-one correspondence between the labels and the model's output range, facilitating supervised learning during training.

[0172] Verification and iteration: Compare the classification results with subsequent statistics on actual accidents and conflicts. If there are significant deviations, adjust the thresholds or model parameters to ensure the interpretability and accuracy of the risk index.

[0173] Through the above process, this invention not only uses real accident data as annotation, but also establishes a continuous risk measurement by combining the risk index (such as low, medium and high risk) output by the model, so that the model can learn the transition process from normal state to risk state, thereby realizing early warning and gradual evolution identification in future prediction.

[0174] During the training phase, incidents / conflicts labeled in historical data are used as supervision signals, and weighted cross-entropy or focal loss is used as the loss function:

[0175]

[0176] in: Indicates the true risk label; This represents the probability of risk predicted by the model; This represents the class weight, used to address the class imbalance problem caused by the scarcity of risk samples.

[0177] The reason for choosing the cross-entropy loss function is that risk identification is essentially a classification problem, requiring road segments to be classified into different risk levels. Cross-entropy can effectively measure the difference between the predicted probability distribution and the true label distribution. By optimizing the above loss through gradient descent, the model can learn the correspondence between spatial propagation, temporal evolution, and risk events on the training set, and achieve early risk identification in the testing or online phase.

[0178] S6. Risk Scoring and Evolution Analysis: Based on the situation vector output by the spatiotemporal graph attention network, calculate the segment-level risk score sequence, calculate the segment-level risk score, and analyze the dynamic evolution of the risk score to identify the temporal and spatial distribution of risk events. The risk score is obtained by a weighted combination of normalized features of nodes / segments (e.g., speed dispersion, density abrupt changes, travel time residuals, and weather mapping factors), and the weights can be obtained through learning or set by human experience. The score is recorded in time series form. Evolution analysis includes the detection of sudden increases in the score sequence and the cumulative trend analysis, thereby marking the time window of risk event occurrence and the possible propagation direction (i.e., risk evolution link). Corresponding evaluation indicators are set for each task.

[0179] Specifically, the previous The input data matrix X for each normalized detection cycle The network, which has been trained, is input into a predefined highway network map structure, and a multi-head output model is designed for the risk identification task.

[0180] The assessment indicators used in the risk scoring and evolution analysis process include:

[0181] Accuracy metrics for risk identification include recall, false positive rate, F1 score for maximizing risk identification, and precision.

[0182] The accuracy indicators for traffic parameter prediction include the accuracy of average speed prediction and the accuracy of segment travel time prediction.

[0183] The timeliness indicators of risk warnings include the lead time for risk index warnings and the warning response time.

[0184] The stability metrics of system operation include the system's continuous operational stability and its tolerance to data loss.

[0185] S7. Results Output and Storage: Output and display the results of risk evolution analysis in the form of time series and spatial heatmaps, and store them for subsequent visualization and offline analysis.

[0186] The output content, as specified in the task, can be used for online heatmap visualization or stored in the project log for offline analysis and model backtesting.

[0187] This embodiment also provides a case comparison of high-risk events, such as... Figure 5 As shown in the image: A significant surge in traffic flow at the merging zone caused congestion, with the risk escalating upstream and potentially leading to accidents. The image also shows the evolution of the highway heatmap after implementing traffic control measures such as entrance ramp flow control and emergency lane opening, based on the anticipated congestion. Figure 6 As shown, the congestion situation has significantly improved, and the overall risk situation has been effectively managed.

[0188] Management measures can be triggered for road sections that show evolution signs, such as variable speed limits, lane control, detour guidance, and temporary traffic control at toll stations. Management personnel can implement tiered management measures according to the actual situation.

[0189] To demonstrate the necessity of the three types of core data and the effectiveness of the method, the following comparative implementation designs are presented:

[0190] The study selected a section of a domestic expressway from K15+000 to K45+000, covering a period of June 2024, including various weather conditions such as sunny, rainy, and foggy days.

[0191] Evaluation metrics include recall rate, false alarm rate, F1 score, and risk index advance warning, among other key metrics such as average speed and segment travel time.

[0192] Example A (Baseline): Training and testing were completed using only radar detector data from the XX highway section, serving as the baseline. The accuracy of the model in predicting traffic conditions and risks was evaluated.

[0193] Example B (Radar Detector + Gantry / Toll Station): Based on A, add gantry / toll station transaction data features and compare the improvement in indicators such as recall rate and advance timing of risk detection.

[0194] Example C (Radar Detector + Weather Factors): Based on A, meteorological data is added to compare the improvement in indicators such as recall rate and lead time changes in risk detection under severe weather scenarios.

[0195] Example D (Complete): Using radar detectors + gantry / toll station transaction data features + meteorological data, and enabling quality diagnostics, compare the performance differences with Example A, Example B, and Example C.

[0196] Experiments have shown that this method can improve the risk identification F1 value by about 15% to 20% compared with the baseline threshold method, and provide early warnings 1 to 3 minutes in advance, thus gaining reaction time for proactive management.

[0197] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for risk identification of highway networks based on spatiotemporal graph attention networks, characterized in that: Includes the following steps: S1. Data Acquisition: Collect structured data including radar detector data, gantry / toll station transaction data and meteorological data, and align and serialize the various types of data in the structured data according to a unified time base; S2. Data Preprocessing and Quality Diagnosis: The collected data is processed to unify the units and sampling frequency, and missing and abnormal data are detected and repaired. Multi-class data is matched spatiotemporally and consistently, and a quality label is generated for each data record. S3. Road network map construction: Based on the research road segments, construct a regional road network structure map containing multiple nodes, map adjacent road segments and weaving / merging relationships as graph edges, and determine edge weights based on real-time traffic characteristics and external meteorological conditions; S4. Modal feature extraction: Extract the features of nodes and edges from the preprocessed data. The features should include at least speed, flow, occupancy rate, segment travel time residual, and weather factors. S5. Spatiotemporal graph attention network construction: Using graph attention mechanism and temporal aggregation technology, a spatiotemporal graph attention network is constructed. Within a sliding time window, the spatiotemporal dynamic representation of nodes and edges is learned. The regional road network structure map constructed in S3 and the features in S4 are input to generate and output a road segment-level situation vector. S6. Risk Scoring and Evolution Analysis: Based on the situation vector output by the spatiotemporal graph attention network, calculate the road segment-level risk score sequence, calculate the road segment-level risk score, analyze the dynamic evolution of the risk score, and identify the temporal and spatial distribution of risk events. S7. Results Output and Storage: Output and display the results of risk evolution analysis in the form of time series and spatial heatmaps, and store them for subsequent visualization and offline analysis.

2. The method for identifying highway network risks based on spatiotemporal graph attention networks according to claim 1, characterized in that: In S1, the structured data collected during the data collection and integration process includes: Radar detector data includes fields such as timestamp, detector ID, lane number, traffic flow, average vehicle speed, occupancy rate, average headway, and sampling period. Gantry / Toll Station Transaction Data: Includes transaction time, gantry / toll station ID, vehicle unique identification number, vehicle category, lane number, and vehicle's entry / exit direction; Meteorological data includes timestamps, meteorological observation point IDs, visibility, precipitation, road surface temperature, and road surface conditions.

3. The method for risk identification of highway networks based on spatiotemporal graph attention networks according to claim 1, characterized in that: In S2, the data preprocessing and quality diagnosis process is as follows: S2.1 Unit Normalization: The collected data of various types is input by the data access module. The units of the input fields are unified, and time calibration and sorting are performed. The timestamps of various types of data are unified and matched to the corresponding nodes in the graph in space. S2.2 Unified sampling frequency: The normalized data is time-aligned according to a unified sampling period to make the sampling frequency of various data sources consistent. S2.3 Data Quality Assessment and Diagnosis: Robust statistical methods are used to identify outliers in continuous quantitative data; topological consistency verification is used to diagnose cross-sectional velocity data in case of consistency anomalies; when a velocity change occurs between adjacent sections in a short period of time and the velocity change is not accompanied by a corresponding change in flow data, the velocity data of that cross-section is marked as abnormal data. S2.4 Data Repair and Missing Data Imputation: Use travel time threshold checks to remove obviously erroneous transaction data from gantry / toll station transaction data, and eliminate the interference of extreme values ​​on data quality; Determine whether it is necessary to repair missing outliers. If necessary, use spatiotemporal interpolation to integrate historical databases for repair. After repair, proceed to the next step for the generated samples. If not necessary, proceed directly to the next step. S2.5 Quality Label Generation: For each data record in the dataset, a quality level label is generated. The quality level labels are divided into three categories: high, medium, and low. During the model input stage of the graph attention network, for records carrying low-quality labels, the corresponding input features are processed in the following ways: the weight contribution of the corresponding input features in the graph attention network is reduced, or the input features are masked.

4. The method for risk identification of highway networks based on spatiotemporal graph attention networks according to claim 1, characterized in that: In S2, the quality label The definition is as follows: ; in: This represents a missing mask, where 1 represents missing data and 0 represents complete data. This represents the consistency score with neighboring nodes / historical sequences. This indicates the stability index of continuous multi-period fluctuations. Represents the weighting coefficients, satisfying .

5. The method for identifying highway network risks based on spatiotemporal graph attention networks according to claim 1, characterized in that: In S3, the mapping rules between nodes and edges during the road network graph construction process are as follows: The locations where radar detector data and gantry / toll station transaction data are acquired, as well as the starting and ending points of roads, are mapped to nodes in the road network map; adjacent road segments and traffic weaving / merging areas are mapped to edges in the map, with edge weights determined by factors including real-time average speed, number of lanes, road segment length, and weather correction factors.

6. The method for risk identification of highway networks based on spatiotemporal graph attention networks according to claim 1, characterized in that: In S4, the selection criteria for modal feature extraction include: The magnitude of the influence factor for each feature is calculated using a normalization function; The correlation between features is assessed using analysis of covariance. Construct a feature scoring system that comprehensively considers influence and relevance, and prioritize features with high scores.

7. The method for risk identification of highway networks based on spatiotemporal graph attention networks according to claim 1, characterized in that: In S5, the construction process of the spatiotemporal graph attention network is as follows: S5.1 Constructing the graph structure: Layout nodes, edges and their characteristics to generate a road network structure map of the highway study area; S5.2 Feature Extraction and Aggregation: Utilize graph attention mechanism to aggregate and update node information, and train a risk model through supervised learning to extract high-dimensional features of graph nodes; S5.3 Spatiotemporal Update: Within the sliding time window, the representation of the node is spatiotemporally aggregated to reflect the local traffic status and historical evolution trend at the current moment; S5.4 Output Layer and Risk Scoring: Calculate road segment-level risk scores, perform risk surge detection and trend analysis, and mark the time and spatial scope of risk events.