Urban area traffic accident risk assessment method based on multi-source data fusion
By unifying spatial reference coordinates on the ArcGIS platform and constructing the Risk-STGNet model, the problem of insufficient fusion of multi-source heterogeneous data was solved, enabling refined and dynamic assessment of urban traffic accident risks and improving the accuracy and practicality of the assessment.
Patent Information
- Application Number
- CN202510822375.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies lack the ability to fuse multi-source heterogeneous data, have coarse risk assessment granularity, lack a systematic evaluation framework, and are unable to fully reflect the dynamic changes in the urban traffic environment, resulting in inaccurate traffic accident risk assessments.
By unifying spatial reference coordinates through the ArcGIS platform, dividing multi-dimensional spatiotemporal grids, constructing the Risk-STGNet model, extracting multi-source data features using LSTM, GCN, and ConvLSTM networks, and optimizing the model using a weighted loss function, the fusion of multi-source data and risk prediction are achieved.
It enables refined and dynamic assessment of traffic accident risks in urban areas, improving the accuracy and practicality of the assessment and providing scientific decision support for traffic management.
Smart Images

Figure CN120911940A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of traffic safety, and particularly relates to a city area traffic accident risk assessment method based on multi-source data fusion. BACKGROUND
[0002] With the acceleration of urbanization, the complexity of urban traffic system is increasing day by day, and the frequent traffic accidents bring serious challenges to the safety of residents and urban operation. The traditional city traffic risk assessment method mostly depends on a single data source, such as only according to historical accident data, or only focusing on the static characteristics of the traffic system, like the basic design parameters of the road, etc., which is difficult to fully reflect the dynamic change characteristics in the complex and changeable urban traffic environment. In recent years, with the development of big data technology, multi-source data (such as road network structure, POI, weather information, etc.) fusion and big data analysis method have gradually emerged, which brings new opportunities for dynamic prediction and fine management of traffic risk. However, there are still many problems in the existing method: firstly, the multi-source heterogeneous data fusion capability is insufficient, because different types of data have differences in format, structure and semantics, it is difficult to effectively integrate spatial, temporal and semantic information; secondly, the risk assessment granularity is coarse, which cannot accurately locate the high-risk areas in the city, resulting in lack of pertinence of management measures; thirdly, there is a lack of systematic evaluation framework, which is difficult to be widely applied in actual urban traffic management. Therefore, it is urgent to develop a comprehensive risk assessment method which can fuse multi-source traffic data, consider the spatio-temporal characteristics and urban function factors, so as to improve the evaluation accuracy and practicability, and provide more effective decision support for urban traffic management.
[0003] In view of this, the present application is proposed. SUMMARY
[0004] The present application aims to overcome the shortcomings of the prior art, and provides a city area traffic accident risk assessment method based on multi-source data fusion, which solves the problems of insufficient multi-source heterogeneous data fusion capability, coarse risk assessment granularity and lack of systematic evaluation framework in the prior art, realizes fine and dynamic assessment of city area traffic accident risk, and provides scientific decision support for urban traffic management.
[0005] The purpose of the present application is solved by the following technical scheme: The present application provides a city area traffic accident risk assessment method based on multi-source data fusion, comprising the following steps: Step 1, data fusion and preprocessing Collecting relevant multi-source data, dividing into time characteristic data, space characteristic data and space-time characteristic data according to the time and space characteristics of the data, and preprocessing each type of characteristic data, wherein the spatial reference coordinates are unified through the ArcGIS platform, and the spatial filtering of the space characteristic data and the space-time characteristic data is carried out; Step 2, space-time unit division Divide the urban area into square space grids with edge lengths of 1.5km, 3km and 5km through the ArcGIS platform, and map the space characteristic data and the space-time characteristic data using the unique grid unit number UID, set three time scales of week, month and quarter, and construct a multi-dimensional space-time data matrix; Step 3, feature extraction and normalization Extract time features, space features and space-time features from multi-source data, process categorical features using one-hot encoding, and normalize numerical features using Z-score standardization or Min-Max normalization; Step 4, model construction and training Construct a risk space-time graph neural network Risk-STGNet, which includes a time feature extraction module, a space feature extraction module, a space-time feature extraction module and a feature fusion module, the feature fusion module is composed of two fully connected networks, which are used to cascade the time, space and space-time features output by the time feature extraction module, the space feature extraction module and the space-time feature extraction module after dimension expansion and alignment; Step 5, model verification and result output Verify the model through historical data backtesting, and output the risk prediction value of each unit grid and the visual risk map under the specified space-time scale.
[0006] Further, in step 1, the multi-source data includes historical accident data, weather data, road data, POI data and population data; The historical accident data includes occurrence time, latitude and longitude, number of vehicles involved and casualty information; The weather data includes temperature, weather type, wind direction and wind power, and air quality index; The road data comes from the vector road network of OpenStreetMap, including the length and density characteristics of main roads, secondary roads, connecting roads, branch roads and other roads; The POI data includes scenic spots, public services, companies and businesses, transportation services, life services, consumer entertainment, and their density characteristics; The population data is high-resolution population raster data generated based on WorldPop correction.
[0007] Further, in step 3, according to the attributes of the multi-source data, the data with time characteristics is meteorological data, the data with space characteristics is road data, POI data and population data, and the data with space-time characteristics is historical accident data (the grid traffic accident risk value Risk is generated by the calculation unit as the space-time feature).
[0008] Further, in step 3, the normalization processing includes: The numerical features in the time features adopt Min-Max normalization. The number and density of consumer entertainment POIs, the length and density of connecting roads, the length and density of branch roads, the length and density of other roads, and the population density in the space features adopt Z-score standardization, and the rest of the space numerical features adopt Min-Max normalization.
[0009] Further, in step 4, the time feature extraction module includes a 2-layer LSTM network followed by a fully connected layer to expand the global time features to each grid. The space feature extraction module includes a 3-layer GCN network, each of which is followed by a ReLU activation function and a Dropout layer with a dropout rate of 0.2. The space-time feature extraction module includes a 2-layer ConvLSTM network. Among them, the hidden layer dimensions of the time feature extraction module, the space feature extraction module, and the space-time feature extraction module are all 128.
[0010] Further, in step 4, a weighted mean square error loss function is used for model training (the weighted mean square error loss function gives more weight to high-risk areas, so that the model optimization pays more attention to high-risk areas), and the weighted mean square error loss function is as follows: Among them, is the sample weight, N is the number of samples, and represent the true risk value and the model prediction value of the i-th sample, respectively.
[0011] It should be noted that the output features of the three modules (time feature extraction module, space feature extraction module and space-time feature extraction module) are concatenated after dimension expansion and alignment, specifically: the space features are increased in batch dimension by unsqueeze (0), and then aligned with the time features in batch dimension by expand method, and finally input into the feature fusion module composed of two fully connected networks, and output the grid risk prediction value.
[0012] The model of the application realizes effective integration of multi-source heterogeneous data through a space-time decoupling-fusion architecture. In particular, a weighted loss function is designed for high-risk areas (a risk division threshold value is more than 95% of the quantile of the target value): for samples with a risk prediction value > 1.0, the loss function will impose a 5-fold weight penalty on them; and for low-risk samples, a smaller weight is set to prevent overfitting. Through this adjustment, the model can effectively alleviate the problem of data imbalance, while taking into account the prediction of low-risk areas, it significantly improves the prediction effect of high-risk areas.
[0013] In step 5, the indicators for model verification include mean absolute error (MAE), mean square error (MSE), and F1 score for binary classification. 1) Mean absolute error (MAE) is used to measure the average absolute deviation between the model prediction value and the true risk value, and the calculation formula is: wherein, is the true risk value, is the model prediction value, and N is the sample size; MAE is not sensitive to outliers and can directly reflect the absolute error level of the model, and the smaller the value, the higher the prediction accuracy. 2) Mean square error (MSE) is used to reflect the robustness of the model, and the calculation formula is: wherein, is the true risk value, is the model prediction value, and N is the sample size; due to the amplification effect of the square term, MSE is more sensitive to large errors and can reflect the robustness of the model, and the smaller the value, the better the model's ability to predict extreme risk values. 3) F1 score is a harmonic mean of precision and recall, which is used to reflect the imbalance of the classification task, and the calculation formula is: wherein, Precision is the proportion of actual positive samples in the predicted positive samples, and Recall is the proportion of correctly predicted positive samples in the actual positive samples.
[0014] Precision (Precision), calculation formula: Recall (Recall), calculation formula: The confusion matrix is a core tool for classification tasks, and its structure is shown in the following table: Confusion matrix for regional accident risk prediction Based on the historical risk data of the validation set, the risk threshold is determined as the 95% quantile of the target value; the predicted results and the actual results are divided into two categories of samples of high risk (>= threshold) and low risk (< threshold) according to the threshold, wherein the high risk is a positive class; Based on the binary classification result, the precision, recall and F1 score are calculated to evaluate the prediction ability of the model for the high-risk area.
[0015] Compared with the prior art, the present application has the following beneficial effects: 1. Comprehensive multi-source data fusion: The feature weight of the multi-source data (historical accident data, meteorological data, road data, POI data and population data) in the present application is obtained through the implicit learning of the neural network structure of the model, rather than manually setting a fixed weight. The multi-source heterogeneous data (time series data, spatial graph structure data, spatio-temporal sequence data) are respectively input into independent feature extraction modules; the outputs of different modules are integrated into a unified feature vector through a feature concatenation operation; the importance weight of each feature is automatically learned by using a fully connected neural network, and the weight is dynamically optimized based on the loss function during the training process, covering multi-dimensional influence factors of space-time and urban function, breaking through the limitation of a single data source, and enabling the urban traffic environment to be comprehensively described from a multi-dimensional perspective, so that the evaluation result is more comprehensive.
[0016] 2. Strong model prediction ability: The Risk-STGNet model of the present application is the first to use the advanced LSTM-GCN-ConvLSTM fusion structure, which effectively extracts the time series dynamics, spatial correlation and spatio-temporal coupling rules through the synergistic effect of LSTM, GCN and ConvLSTM, and simultaneously uses a weighted loss function to effectively solve the data imbalance problem, thereby enhancing the stability and generalization ability of the model and significantly improving the accuracy of risk prediction.
[0017] 3. Strong multi-scale adaptability: The present application supports flexible selection of 1.5km-5km spatial grid and week-quarter time scale, meeting the needs of short-term high-precision prediction and medium-and long-term trend analysis.
[0018] 4. Result interpretability and practicality: The risk level can be output through grid cell visualization, providing quantifiable and locatable decision-making basis for traffic management departments, and the model has high expandability and is suitable for different cities or regions.
[0019] 5. High expandability: The evaluation method of the present application is suitable for traffic accident risk prediction in other cities or regions, and only the data set needs to be updated. BRIEF DESCRIPTION OF DRAWINGS
[0020] The drawings herein are incorporated into the specification and form part of the specification, which together with the specification serve to explain the principles of the present application.
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, for those skilled in the field, other drawings can also be obtained from these drawings without any creative effort.
[0022] Figure 1 is the flow chart of the urban area traffic accident risk assessment method based on multi-source data fusion proposed by the present application; Figure 2 is the road network distribution characteristic map of a certain city in northwest China in the embodiment of the present application; Figure 3 is the population distribution characteristic map of a certain city in northwest China in the embodiment of the present application; Figure 4 is the traffic accident distribution characteristic map of a certain city in northwest China in January 2018 in the embodiment of the present application; Figure 5 is the traffic accident distribution map of a certain city in northwest China in January 2018 after dividing three types of urban space grids in ArcGIS in the embodiment of the present application, and UID is generated; wherein, (a) is a 1.5km grid scale, (b) is a 3km grid scale, and (c) is a 5km grid scale; Figure 6 is the multi-source data processing flow chart of the present application; Figure 7 is the schematic diagram of the Risk-STGNet model structure proposed by the present application; Figure 8 is the comparison diagram of the prediction result and the actual result in the embodiment of the present application; wherein, (a) is the comparison of the real value and the predicted value of the regional grid accident risk under the 5km grid, (b) is the comparison of the real value and the predicted value of the regional grid accident risk under the 3km grid, and (c) is the comparison of the real value and the predicted value of the regional grid accident risk under the 1.5km grid. DETAILED DESCRIPTION
[0023] The exemplary embodiments will be described in detail herein with reference to the accompanying drawings. Unless otherwise indicated, the same numbers on different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of apparatuses consistent with some aspects of the present application as detailed in the appended claims.
[0024] In order to make those skilled in the art better understand the technical solutions of the present application, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. EMBODIMENT
[0025] Please refer toFigure 1 The application provides a city area traffic accident risk assessment method based on multi-source data fusion. Taking a city in northwest China as an example, the method specifically comprises the following steps: Step 1, data fusion and preprocessing Five types of multi-source data, including historical accident data, meteorological data, road data, POI data and population data, are collected. The data is unified in spatial reference coordinates through the ArcGIS platform and spatially aligned using unique grid cell numbers (UID) to ensure consistency and integrity of the data.
[0026] Among them, the historical accident data includes accident occurrence time, longitude and latitude, number of motor vehicles involved and casualty situation; the meteorological data includes air temperature, weather type, wind direction and wind power, air quality index (AQI); the road data comes from the vector road network information of OpenStreetMap, including the length and density characteristics of main roads, secondary roads, connecting roads, branch roads and other roads; the POI data can be extracted through the Gaode map API, which is divided into six categories: scenic spots, public services, companies and businesses, transportation services, life services and consumer entertainment, and their density characteristics; the population data is based on high-resolution population raster data generated after WorldPop correction.
[0027] (1) Meteorological data The time characteristic variable selected in this embodiment that has a significant impact on regional accident risk is meteorological related data, which specifically includes many factors such as air temperature, weather conditions, wind power level, wind direction, air quality index, etc. which have been proven to have a close relationship with the severity of road traffic accidents.
[0028] The meteorological information of a typical city in northwest China from 2018 to 2023 is derived from the historical weather data opened by the 2345 weather forecast website. There are some missing information such as AQI (air quality index), wind power and direction, weather conditions, etc. in the data. Through the supplementary data of China's online air quality monitoring and analysis platform, the sample data is finally obtained as shown in Table 1.1.
[0029] Table 1.1 Meteorological data sample (2) POI data The POI data used in this embodiment is obtained through the API interface of Gaode Map. Considering the limitations of the API interface, it can only obtain current POI data, which makes it impossible to obtain historical POI data of a typical city in northwest China. Therefore, the POI data finally used by the embodiment corresponds to June 2023. According to the obtained raw data information, it contains detailed contents such as point of interest type, latitude and longitude coordinates, industry category, industry subcategory, industry small category, and street address. The basic situation is shown in Table 1.2.
[0030] Table 1.2 POI data classification statistics The original POI data provided by the Gaode Map open platform includes 25 POI categories, including life services, entertainment, medical care, and many other fields, providing rich spatial data support for accident risk assessment. In order to more conveniently input these data into subsequent research models, the embodiment reclassifies and integrates POI according to their different functions, and divides them into 6 categories, namely life service category, consumer entertainment category, company business category, scenic spot category, public service category, and transportation service category. After obtaining the data, the data preprocessing process is carried out to eliminate the data contents that do not meet the requirements such as repeated items and blank items. Finally, 581970 valid data are retained.
[0031] (3) Road data The road data used in this embodiment is collected from the Open Street Map (OSM) open source map download platform. The vector road network data of a city in northwest China is obtained using the ArcGIS clipping function. The data contains many important fields of road segments such as the geometric position of the road segment itself, road segment level, road segment length, and other information. The road segment level contained in OSM is stored in the "fclass" field of the road network data. According to the original classification method, there are 27 small categories. In the embodiment, the road network data is extracted to obtain 25 categories of road segment levels in a city in northwest China. The road network distribution characteristics of the city are shown in Figure 2
[0032] The road network structure of the city radiates from the city center and develops outward in a ring and radial pattern from the urban area to the suburbs. The city center has a high-density road network that is densely interwoven, providing high road accessibility and meeting the traffic demand. As it goes outward, especially from the southwest of the city center, the road network becomes sparse after entering the mountainous area in the southwest region.
[0033] In order to facilitate subsequent analysis and research, these road types are further classified into 5 categories based on their different functions: main roads, secondary roads, connecting roads, branch roads, and other roads. The specific classification of these 5 categories of roads is shown in Table 1.3.
[0034] Table 1.3 Road data classification statistics (4) Population distribution data The population data used in this embodiment is derived from the WorldPop dataset and the 2020 population census data of a certain city in Northwest China. The WorldPop dataset is provided by the WorldPop project launched by the University of Southampton in 2013. The resolution of this dataset is 100 meters, and it is mainly used to generate geographic spatial data including population number, population density, population flow, age and gender structure, etc.
[0035] Due to the differences between the data released by WorldPop and the population census data published by the government of a certain city in Northwest China, in this embodiment, the 2020 population census data published by the government is used as a reference to correct the population distribution data of a certain city in Northwest China released by WorldPop in 2020. After correction, the population distribution is as shown in Table 1.4. Figure 3
[0036] From the population distribution map of a certain city in Northwest China, it can be seen that the population distribution presents a pattern of high density in the center and scattered outward. The high value area of population density is mainly distributed in the urban area, and the economic, public resources and transportation conditions are relatively mature, so more population is distributed in this area. The population density in the suburban county area is not high, and it presents the characteristics of mainly agricultural, low urbanization level, weak location characteristics, etc. The overall population distribution presents a trend of "high in the center and low in the outer circle".
[0037] (5) Accident data The spatio-temporal characteristics data used in this embodiment is derived from the traffic accident dataset of a certain city in Northwest China from January 1, 2018 to October 20, 2023, totaling 931,690 records. This dataset covers various types of traffic accidents, including single vehicle accidents, collisions between motor vehicles, and accidents between motor vehicles and non-motor vehicles. The original data contains accident date, accident number, accident longitude and latitude, number of motor vehicles involved, number of deaths, number of injuries, and accident briefings.
[0038] There are a large number of latitude and longitude missing in the original data, so the accident location information needs to be extracted from the accident brief and rely on the API interface of Gaode map to calibrate the accident records with missing latitude and longitude information. Select January 2018 to September 2023 as the limited time period of accident data, eliminate invalid columns, and eliminate accidents that cannot be calibrated coordinates, resulting in 917791 valid accident information. The valid traffic accident data is shown in Table 1.4: Table 1.4 Example of valid traffic accident data Take the data of January 2018 as an example, use ArcMAP to visualize the traffic accident distribution, as shown in Figure 4 According to the accident distribution map in January 2018, accidents occur mainly in urban areas, and the accident distribution in suburban areas is much less than in urban areas. The accident density of the accident location is much smaller than that in urban areas. Therefore, the accident frequency in urban areas is higher, which may be related to the large vehicle flow, complex road conditions, and vehicle congestion in urban areas.
[0039] Step 2, division of space-time unit Spatial division: Referring to Figure 5 , 6 , through the ArcGIS platform, the city in the embodiment is divided into different grid division scales, and the square grid division method is selected to divide the city into 1.5km, 3km, 5km three square grids, the regularity of the division makes the subsequent data fusion, calculation and model construction more convenient and reliable, and provides a solid foundation for subsequent space-time analysis; Time division: divide into three time scales of week, month and quarter, combined with short-term risk and long-term trend division, which can meet the needs of analyzing short-term risk fluctuations, and also provide analysis of long-term trend changes, providing analysis basis for the universality and reliability of model prediction.
[0040] Step 3, feature extraction and normalization Through feature classification and preprocessing of multi-source heterogeneous data obtained from various channels, it is divided into three categories: one is time characteristic variable, one is space characteristic variable, and the other is space-time characteristic variable. The time characteristic variable is the meteorological related data, the space characteristic variable is the POI data, the road data and the population data, and the space-time characteristic variable is the accident data. After preprocessing the original data, feature extraction needs to be performed on various data, and the feature extraction results are shown in Table 3.1.
[0041] Table 3.1 Feature extraction results wherein the unit grid traffic accident risk value Risk is assigned weights for different types of accidents, and the regional traffic risk is calculated in combination with the accident frequency. Risk = å (accident type weight x accident frequency); wherein the accident type weight is set according to the casualty situation (death accident = 3, serious injury accident = 2, no injury accident = 1). For example, if a certain grid unit has 1 death accident and 2 no injury accidents in the study period, the traffic safety risk of the grid unit is: 1 x 3 + 2 x 1 = 5.
[0042] After feature extraction, normalization needs to be performed according to the feature situation. In the time feature, the One-Hot Encoding form is adopted for the categorical feature, and the temperature, wind power, and AQI level are taken as numerical variables, and the Min-Max normalization is selected. In the spatial feature data, the features using the Z-Score standardization processing method include xfyl, xfyl_per, ljdl, ljdl_per, zldl, zldl_per, qtdl, qtdl_per, and popdensity, and the rest of the numerical features are processed by the Min-Max normalization method.
[0043] Step 4, model construction and training Reference Figure 7 , a risk prediction model Risk-STGNet is constructed, including a time feature extraction module: an LSTM network is used to extract meteorological and AQI time series features; a spatial feature extraction module: a GCN is used to extract road, POI, and population spatial features based on grid adjacency relationships; a spatio-temporal feature extraction module: a ConvLSTM is used to fuse spatio-temporal features; a fusion layer and an output module: a full connection layer is used to output the accident risk value Risk of the grid unit; a weighted loss function is introduced to alleviate the zero inflation problem in the accident data. The constructed model is trained, and after selecting the optimal model parameters, the future period accident risk is predicted; taking this embodiment as an example, the data set division and sliding window parameters are shown in Table 4.1, and the grid division results are shown in Table 4.2: Table 4.1 Data set division and sliding window parameters In this embodiment, based on the full spatio-temporal continuous observation data from 2018 to September 2023, the data is divided into a training set and a validation set in a ratio of 8:2, which is applicable to monthly, quarterly, and weekly time scales, and the sliding window method for constructing input samples remains unchanged. To avoid the uneven number of weeks that may occur when dividing the weekly time scale, the continuous week count method (Week1-Week273) is used to prevent confusion caused by week number jumps. The monthly time scale construction method of the sliding window has a time overlap phenomenon, which is a normal phenomenon determined by the step (1 month) being smaller than the window length (12 months).
[0044] Table 4.2 Grid division results Step 5, model verification and result output The high-risk areas of the city are marked in combination with the prediction results, and a visual risk map is output. The prediction results of the present embodiment are shown in Table 5.1, and the visual risk map is shown in Figure 8 .
[0045] Table 5.1 Prediction results of the present embodiment The results of the present embodiment show that as the spatiotemporal resolution increases, the prediction effect of the model is better, and the MAE and MSE of the 1.5km grid are optimal, being 0.1842 and 0.1088 respectively, but the F1 score is only 74.06%. Therefore, when performing high-precision prediction in the short term, the 1.5km spatial scale and the weekly scale should be selected; and for medium and long-term analysis, the 3km spatial scale and the monthly scale or the quarterly scale have better stability.
[0046] In summary, the present application adopts the form of a combined model, solves the problem of the conflict between recall rate and precision rate of a single model, so that the overall evaluation effect of the model is improved; in addition, the combined model can divide the risk areas into risk levels according to the evaluation results, meeting the demand of the traffic management department for risk area management.
[0047] The above description is merely a specific implementation of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application.
[0048] It should be understood that the present application is not limited to the above-described embodiments, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for urban area traffic accident risk assessment based on multi-source data fusion, characterized in that, Comprising the following steps: Step 1, data fusion and preprocessing Collect relevant multi-source data, divide into time characteristic data, space characteristic data and space-time characteristic data according to the space-time characteristics of the data, and preprocess each type of characteristic data, wherein the space characteristic data and space-time characteristic data are filtered through the ArcGIS platform unified spatial reference coordinates; Step 2, space-time unit division Divide the city area into square space grids with edge lengths of 1.5km, 3km and 5km through the ArcGIS platform, and map the space characteristic data and space-time characteristic data using the unique grid unit number UID, set three time scales of week, month and quarter, and construct a multi-dimensional space-time data matrix; Step 3, feature extraction and normalization Extract time features, space features and space-time features from multi-source data, use one-hot encoding to process categorical features, and use Z-score standardization or Min-Max normalization to process the remaining numerical type features; Step 4, model construction and training Build a risk space-time graph neural network Risk-STGNet, which includes a time feature extraction module, a space feature extraction module, a space-time feature extraction module, and a feature fusion module composed of two fully connected networks, which are used to concatenate the time, space and space-time features output by the time feature extraction module, the space feature extraction module and the space-time feature extraction module after dimension expansion and alignment; Step 5, model verification and result output Verify the model through historical data backtesting, and output the risk prediction value of each unit grid and the visual risk map under the specified space-time scale.
2. The method of claim 1, wherein, In step 1, the multi-source data includes historical accident data, weather data, road data, POI data and population data; The historical accident data includes occurrence time, latitude and longitude, number of vehicles involved and casualty information; The weather data includes temperature, weather type, wind direction and wind power, and air quality index; The road data comes from the vector road network of OpenStreetMap, including the length and density characteristics of main roads, secondary roads, connecting roads, branch roads and other roads; The POI data includes scenic spots, public services, companies and businesses, transportation services, life services, and consumer entertainment, as well as their density characteristics; The population data is high-resolution population raster data generated based on WorldPop correction.
3. The method of claim 2, wherein, In step 3, the data extracted from the multi-source data as time features is weather data, and the data extracted as space features is road data, POI data and population data. The historical accident data used to describe the unit grid traffic accident risk value is used as the data source for extracting space-time features.
4. The urban area traffic accident risk assessment method based on multi-source data fusion according to claim 2, characterized in that, In step 3, the normalization process includes: The numerical features in the time features are normalized by Min-Max; The number and density of consumer entertainment POIs, the length and density of connecting roads, the length and density of branch roads, the length and density of other roads, and the population density in the space features are standardized by Z-score, and the remaining space numerical features are normalized by Min-Max.
5. The method of claim 1, wherein, In step 4, the time feature extraction module comprises a 2-layer LSTM network followed by a fully connected layer to expand the global time feature to each grid; The spatial feature extraction module comprises a 3-layer GCN network, each followed by a ReLU activation function and a Dropout layer with a dropout rate of 0.2; The spatio-temporal feature extraction module comprises a 2-layer ConvLSTM network. The hidden layer dimension of the time feature extraction module, the spatial feature extraction module and the spatio-temporal feature extraction module is 128.
6. The method of claim 1, wherein, In step 4, the weighted mean squared error loss function is used for model training, and the weighted mean squared error loss function is as follows: where ω i is the sample weight, N is the number of samples, y i and represent the true risk value and the model predicted value of the i-th sample, respectively.
7. The urban area traffic accident risk assessment method based on multi-source data fusion according to claim 1, characterized in that, In step 5, the model verification indicators include mean absolute error (MAE), mean squared error (MSE) and F1 score of binary classification. 1) The mean absolute error (MAE) is used to measure the average absolute deviation between the predicted value and the true risk value, and the calculation formula is as follows: wherein y i is the true risk value, is the model predicted value, and N is the number of samples; 2) The mean squared error (MSE) is used to reflect the robustness of the model, and the calculation formula is as follows: wherein y i is the true risk value, is the model predicted value, and N is the number of samples. 3) The F1 score is a harmonic mean of precision and recall, which is used to reflect the imbalance of the classification task, and the calculation formula is as follows: where Precision is the proportion of samples actually positive in the predicted positive samples, and Recall is the proportion of correctly predicted samples in the actual positive samples.