Highway accident prediction system based on big data

By using a big data-based highway accident prediction system to generate feature vectors through spatiotemporal grid partitioning and data mapping, and combining them with risk assessment logic, the system solves the problem of inconsistent spatiotemporal benchmarks for multi-source data, achieves efficient highway accident risk identification and dynamic assessment, and improves the spatiotemporal resolution and robustness of prediction.

CN120932453AInactive Publication Date: 2025-11-11深圳市睿拓新科技有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511179416.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-11-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing highway accident prediction schemes suffer from low spatiotemporal alignment efficiency and ambiguous feature correlations due to inconsistent spatiotemporal benchmarks and heterogeneous data structures from multiple sources, making it impossible to achieve high-precision accident risk identification and dynamic assessment.

Method used

A highway accident prediction system based on big data is adopted. Through spatiotemporal grid division, data mapping and correlation processing, feature vectors with spatiotemporal labels are generated. Combined with risk assessment logic of multiple time windows, efficient spatiotemporal alignment and fusion of multi-source data are achieved.

Benefits of technology

It significantly improves the spatiotemporal resolution and robustness of highway accident prediction, provides accurate risk assessment results, and supports real-time scenario-based risk assessment and efficient accident early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932453A_ABST
    Figure CN120932453A_ABST
Patent Text Reader

Abstract

The invention discloses a road accident prediction system based on big data, and relates to the technical field of big data analysis and prediction, the road accident prediction system comprises an acquisition module, a fusion module and a prediction module, the acquisition module obtains multi-source data of road accident prediction, transmits the multi-source data to the fusion module, carries out space-time alignment processing on the multi-source data through a processing unit, and carries out prediction on the multi-source data; generating feature vectors, transmitting the feature vectors to a prediction module, generating an accident risk prediction result through a prediction unit according to the real-time input feature vectors, constructing a three-dimensional data set by integrating multi-source data, establishing a unified coordinate system by means of space-time grid division, and realizing data differential fusion, static parameter positioning according to road section ID, and dynamic data space-time two-dimensional association. Historical accident space-time backtracking matching is carried out, a feature set with a space-time label is generated, the problem that traditional data benchmarks are not uniform is solved, statistical features are extracted through multiple windows in the prediction stage, multi-factor coupling is analyzed in combination with risk assessment logic, prediction space-time resolution and robustness are improved, and accurate safety management is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data analysis and prediction technology, and in particular to a highway accident prediction system based on big data. Background Technology

[0002] In recent years, highway accident prediction has become a core component of proactive safety control, playing a crucial role in improving road safety management efficiency and reducing accident losses. With the deep penetration of IoT, sensor, and big data technologies, the ability to collect multi-source heterogeneous data in highway scenarios has achieved leapfrog development: Geographic Information Systems accurately characterize highway alignment parameters and the spatial attributes of bridge and tunnel infrastructure; roadside sensing devices capture dynamic traffic flow data such as traffic flow and vehicle lane change frequency in real time; meteorological monitoring networks continuously transmit environmental state parameters such as rainfall and visibility; and historical accident archives fully record the spatiotemporal coordinates and type characteristics of accidents. These data cover everything from micro-level road segment features to macro-level road network layout in the spatial dimension, and form a continuous sequence from second-level real-time data to annual statistical data in the temporal dimension, providing a rich data foundation for building high-precision accident prediction.

[0003] Existing highway accident prediction schemes face significant technical bottlenecks: On the one hand, due to inconsistent spatiotemporal benchmarks and heterogeneous data structures, the fusion of basic geographic data, dynamic traffic data, environmental perception data, and historical accident data results in low spatiotemporal alignment efficiency and ambiguous feature correlations, making it difficult to form effective feature expressions including spatiotemporal coupling relationships. On the other hand, traditional prediction lacks multi-dimensional correlation analysis of dynamic traffic flow with environmental factors and road conditions, and cannot capture the accident causal patterns under the intertwined effects of spatial grid units, time windows, and multi-source parameters in real time, resulting in insufficient spatiotemporal resolution and coarse risk level classification in the prediction results. Therefore, there is an urgent need for a prediction system that can efficiently perform spatiotemporal alignment processing on multi-source heterogeneous data and integrate spatiotemporal grid features, dynamic traffic features, and environmental features to achieve refined identification and dynamic assessment of highway accident risks. Summary of the Invention

[0004] The technical problem addressed by this invention is that existing highway accident prediction schemes face significant technical bottlenecks. On the one hand, due to inconsistent spatiotemporal benchmarks and heterogeneous data structures, the fusion of basic geographic data, dynamic traffic data, environmental perception data, and historical accident data results in low spatiotemporal alignment efficiency and ambiguous feature correlations, making it difficult to form effective feature expressions including spatiotemporal coupling relationships. On the other hand, traditional prediction lacks multi-dimensional correlation analysis of dynamic traffic flow with environmental factors and road conditions, and cannot capture the accident causal patterns under the intertwined effects of spatial grid units, time windows, and multi-source parameters in real time. This leads to insufficient spatiotemporal resolution and coarse risk level classification in the prediction results. Therefore, there is an urgent need for a prediction system that can efficiently perform spatiotemporal alignment processing on multi-source heterogeneous data and integrate spatiotemporal grid features, dynamic traffic features, and environmental features to achieve refined identification and dynamic assessment of highway accident risks.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a highway accident prediction system based on big data, comprising a data acquisition module, a fusion module, and a prediction module; The acquisition module is used to acquire multi-source data for highway accident prediction and transmit it to the fusion module; The fusion module is used to perform spatiotemporal alignment processing on the multi-source data through the processing unit, generate feature vectors, and transmit them to the prediction module; The prediction module is used to generate accident risk prediction results through the prediction unit based on the real-time input feature vector.

[0006] As a preferred embodiment of the highway accident prediction system based on big data described in this invention, the multi-source data includes basic geographic data, dynamic traffic data, environmental perception data, and historical accident data. The basic geographic data includes highway alignment parameters and road infrastructure information; The dynamic traffic data includes real-time traffic flow, vehicle speed distribution, and lane change frequency. The environmental perception data includes real-time meteorological parameters, including rainfall, visibility, and road surface temperature. The historical accident data includes the spatiotemporal coordinates of the accident, whether the accident occurred, and the type of accident. The accident types include rear-end collisions, rollovers, and collisions.

[0007] As a preferred embodiment of the highway accident prediction system based on big data described in this invention, the processing unit includes a spatiotemporal grid partitioning unit, a data mapping unit, and a data association processing unit. The spatiotemporal grid partitioning unit includes: Obtain the geographical scope and time dimension requirements of the highway, and set the preset spatial interval and preset time interval; The highway is divided into several continuous highway grid units based on the preset spatial interval, and a road segment ID is assigned to each highway grid unit. The time axis is divided into several consecutive time period units based on the preset time interval, and a timestamp is assigned to each time period unit. Integrating the aforementioned highway grid units and time period units to generate a spatiotemporal grid system, specifically including: Establish the association mapping relationship between the highway grid unit and the time period unit, combine each highway grid unit with all the time period units in terms of spatiotemporal dimensions to form a spatiotemporal grid unit, and assign a spatiotemporal identifier to each spatiotemporal grid unit; The spatiotemporal identifier is associated with the road segment ID of the highway grid unit corresponding to the spatiotemporal grid unit and the timestamp of the corresponding time period unit; By binding the spatial attributes of the highway grid unit to the temporal attributes of the time period unit through the spatiotemporal identifier, a structured set including all the spatiotemporal grid units is formed; The spatiotemporal grid system is transmitted to the data mapping unit. The spatiotemporal grid system includes road segment IDs and timestamps.

[0008] As a preferred embodiment of the big data-based highway accident prediction system described in this invention, the data mapping unit includes: Extract the spatiotemporal attribute information from the multi-source data; The spatiotemporal attribute information includes the spatial coordinate range and road segment association information of basic geographic data, the collection timestamp and monitoring road segment range of dynamic traffic data, the monitoring time and spatial monitoring point location of environmental perception data, and the spatiotemporal coordinates of accident occurrence of historical accident data. A unified spatiotemporal coordinate system is established based on the road segment IDs and timestamps in the aforementioned spatiotemporal grid system; A mapping relationship is constructed based on the multi-source data and the spatiotemporal grid system.

[0009] As a preferred embodiment of the highway accident prediction system based on big data described in this invention, the mapping relationship includes a first mapping relationship, a second mapping relationship, and a third mapping relationship; The first mapping relationship, which associates the spatiotemporal attribute information of the multi-source data with the road segment ID and timestamp in the spatiotemporal grid system, includes: The basic geographic data is matched to the corresponding highway grid unit through spatial coordinate range and road segment association information; The dynamic traffic data is matched to the corresponding time period unit and highway grid unit by collecting timestamps and monitoring road segment ranges; The environmental perception data is matched to the corresponding time period unit and highway grid unit by monitoring time and spatial monitoring point location; The historical accident data is matched to the corresponding time period unit and highway grid unit by the spatiotemporal coordinates of the accident occurrence.

[0010] As a preferred embodiment of the big data-based highway accident prediction system of the present invention, the second mapping relationship, for the types of basic geographic data, dynamic traffic data, environmental perception data, and historical accident data, respectively formulates matching logic, specifically including: The basic geographic data is statically matched according to road segment ID; The dynamic traffic data and environmental perception data are dynamically matched in both spatiotemporal dimensions. The historical accident data is backtracked and matched according to the spatiotemporal coordinates of the accident occurrence.

[0011] As a preferred embodiment of the highway accident prediction system based on big data described in this invention, the third mapping relationship binds the matched multi-source data with the corresponding highway grid unit and time period unit to generate a feature set with spatiotemporal labels. The feature set includes road segment ID, timestamp, multi-source data type identifier, and the data value corresponding to the multi-source data type identifier; The feature set is transmitted to the data association processing unit.

[0012] As a preferred embodiment of the big data-based highway accident prediction system of the present invention, the data association processing unit includes: Extract the spatiotemporal coordinates and accident tags of the historical accident data, where the spatiotemporal coordinates of the accident correspond to the road segment ID and timestamp; The accident label includes whether an accident has occurred and the type of accident. Performing spatiotemporal matching of the feature set with historical accident data using the road segment ID and timestamp includes: Spatial dimension matching is performed based on the road segment IDs in the feature set and the road segment IDs corresponding to historical accident data. If the road segment IDs match, the spatial dimension matching is determined to be successful. Based on the timestamps in the feature set and the timestamps corresponding to historical accident data, time dimension matching is performed. If the timestamps are in the same time period unit, the time dimension matching is determined to be successful. When both the spatial and temporal dimensions are successfully matched, it is determined that the feature set is successfully matched spatiotemporally with the historical accident data; The successfully matched accident labels are associated with the corresponding features to generate a spatiotemporal feature vector with the accident labels, and the spatiotemporal feature vector is transmitted to the prediction module. The spatiotemporal feature vector includes spatiotemporal features, data features, and corresponding accident labels.

[0013] As a preferred embodiment of the big data-based highway accident prediction system of the present invention, the prediction unit includes: The spatiotemporal feature vector is used to generate a scrolling feature vector based on a preset time window; The preset time window includes a short-term window, a medium-term window, and a long-term window, which respectively generate short-term feature vectors, medium-term feature vectors, and long-term feature vectors. For the indicators of the multi-source data, the mean characteristics, variance characteristics and extreme value characteristics are calculated within the corresponding time period of the preset time window; The indicators of the multi-source data include real-time traffic flow, vehicle speed distribution, lane change frequency, rainfall, visibility, and road surface temperature. The rolling feature vector is calculated based on a preset risk assessment engine to obtain the accident risk prediction result; The risk assessment engine includes a preset set of assessment logics, which includes: When the rainfall exceeds the first preset threshold and the curve radius is less than a meters, the combined cause assessment logic for increasing the accident probability threshold is used. The dynamic traffic assessment logic raises the risk level by one level when the real-time traffic flow is greater than b vehicles / hour and the lane change frequency is greater than c times / minute.

[0014] As a preferred embodiment of the highway accident prediction system based on big data described in this invention, the accident risk prediction result includes the probability value of accident occurrence and the corresponding risk level. The probability value of the accident occurrence is based on the correlation analysis between historical accident data and rolling feature vectors, including: When the conditions of rainfall exceeding the first preset threshold and curve radius being less than d meters are triggered, a real-time probability value is generated by superimposing the probability increase amount on the baseline probability value obtained from the accident occurrence rate statistics of historical spatiotemporal grid units of the same type. The risk levels are divided into multiple threshold ranges based on the assessment logic, and are associated with real-time probability values ​​and the triggering status of dynamic traffic assessment logic, specifically including: The multi-level threshold range includes a low-risk threshold range, a medium-risk threshold range, and a high-risk threshold range; The boundary values ​​of the low-risk, medium-risk, and high-risk threshold intervals are associated with the parameter thresholds in the evaluation logic set, including: In the initial state, when the probability value of an accident falls into the low-risk threshold range, it corresponds to a low-risk level; when the probability value of an accident falls into the medium-risk threshold range, it corresponds to a medium-risk level; and when the probability value of an accident falls into the high-risk threshold range, it corresponds to a high-risk level. When the level adjustment condition in the assessment logic is triggered, the threshold interval boundary value corresponding to the current risk level is adjusted so that the risk level enters the corresponding level interval according to the adjustment condition.

[0015] The beneficial effects of this invention are as follows: By integrating multi-dimensional data and performing spatiotemporal coupling analysis, the predictive efficiency of highway accidents is significantly improved. It integrates multi-source data from basic geography, dynamic traffic, environmental perception, and historical accidents to construct a three-dimensional dataset including road facilities, traffic flow, environmental parameters, and historical accident characteristics. A unified spatiotemporal coordinate system is established through spatiotemporal grid division, enabling differentiated data fusion. Static road parameters are accurately located based on road segment IDs, dynamic traffic and environmental data are correlated through spatiotemporal dimensions, and historical accident data is back-matched according to spatiotemporal coordinates, generating a structured feature set including spatiotemporal labels. This solves the problems of inconsistent spatiotemporal benchmarks and inefficient data fusion in traditional solutions. During the prediction stage, statistical features of mean and variance are extracted through multiple time windows. Combined with risk assessment logic, the co-causal relationship between meteorological conditions, road alignment, and traffic flow parameters is analyzed to achieve accurate calculation of accident probability and dynamic adjustment of risk levels. The invention finely characterizes the accident risk features of each spatiotemporal grid unit, improving the spatiotemporal resolution and robustness of predictions. It provides real-time, scenario-based risk assessment results for traffic management departments, supporting efficient accident early warning, lane scheduling, and resource allocation, and promoting the intelligent and precise advancement of highway safety management. Attached Figure Description

[0016] Figure 1 This is a basic flowchart of a highway accident prediction system based on big data, provided as an embodiment of the present invention. Detailed Implementation

[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0018] Example, refer to Figure 1 As an embodiment of the present invention, a highway accident prediction system based on big data is provided, including a data acquisition module, a fusion module and a prediction module; The acquisition module is used to acquire multi-source data for highway accident prediction and transmit it to the fusion module; The fusion module is used to perform spatiotemporal alignment processing on multi-source data through the processing unit, generate feature vectors, and transmit them to the prediction module. The prediction module is used to generate accident risk prediction results through the prediction unit based on the real-time input feature vector.

[0019] In one embodiment, highway accident prediction is achieved through the collaborative efforts of a data acquisition module, a fusion module, and a prediction module. The data acquisition module integrates roadside sensors, a GIS system, and a historical database, and collects multi-source data in real time, including basic geographic data, dynamic traffic data, environmental perception data, and historical accident data, and transmits it to the fusion module. The fusion module constructs a unified coordinate system including road segment IDs and timestamps by dividing the data into spatiotemporal grids (e.g., 500-meter spatial intervals and 15-minute time intervals). It integrates the data according to three logical categories: static road segment matching, spatiotemporal dynamic correlation, and accident coordinate backtracking, and generates feature vectors with spatiotemporal labels. The prediction module extracts statistical features of the data based on short, medium, and long-term time windows, and calculates the probability of accidents and risk levels by combining preset risk assessment logic (e.g., coupling of meteorological and road conditions and linkage of traffic flow parameters). Finally, it outputs real-time prediction results, providing accurate risk warning support for traffic management.

[0020] Multi-source data includes basic geographic data, dynamic traffic data, environmental sensing data, and historical accident data; Basic geographic data includes highway alignment parameters and road infrastructure information; Dynamic traffic data includes real-time traffic flow, vehicle speed distribution, and lane change frequency; Environmental sensing data includes real-time meteorological parameters, such as rainfall, visibility, and road surface temperature. Historical accident data includes the spatiotemporal coordinates of the accident, whether the accident occurred, and the type of accident; Accident types include rear-end collisions, rollovers, and crashes.

[0021] The processing unit includes a spatiotemporal grid partitioning unit, a data mapping unit, and a data association processing unit; The spatiotemporal grid division unit includes: Obtain the geographical scope and time dimension requirements of the highway, and set the preset spatial interval and preset time interval; The highway is divided into several continuous highway grid units based on a preset spatial interval, and a road segment ID is assigned to each highway grid unit. The timeline is divided into several consecutive time periods based on a preset time interval, and a timestamp is assigned to each time period. Integrating highway grid units and time period units to generate a spatiotemporal grid system, specifically including: Establish the association mapping relationship between highway grid units and time period units, combine each highway grid unit with all time period units in spatiotemporal dimensions to form a spatiotemporal grid unit, and assign a spatiotemporal identifier to each spatiotemporal grid unit; The spatiotemporal identifier is associated with the road segment ID of the highway grid unit corresponding to the spatiotemporal grid unit and the timestamp of the corresponding time period unit; By binding the spatial attributes of highway grid units with the temporal attributes of time period units through spatiotemporal identifiers, a structured set including all spatiotemporal grid units is formed; The spatiotemporal grid system is transmitted to the data mapping unit. The spatiotemporal grid system includes road segment IDs and timestamps.

[0022] In one embodiment, the multi-source data includes four core types of data: basic geographic data including highway alignment parameters (such as curve radius and longitudinal slope) and road facility information (such as guardrail type and lighting conditions); dynamic traffic data including real-time traffic flow, vehicle speed distribution (such as average speed of each lane) and real-time monitoring parameters of lane change frequency; environmental perception data collecting real-time meteorological parameters such as rainfall, visibility and road surface temperature; and historical accident data recording the spatiotemporal coordinates (latitude and longitude and corresponding road segment ID), accident status (whether it occurred) and type (rear-end collision, rollover and collision) of the accident.

[0023] In the processing unit, the spatiotemporal grid division unit first sets a preset spatial interval (e.g., 500 meters) and time interval (e.g., 15 minutes) based on the geographical range of the highway (e.g., the K0-K100 section of a certain expressway) and the accuracy requirements of time analysis (e.g., real-time updates every 15 minutes). The highway is divided into continuous grid units with unique road segment IDs (e.g., ID-001 for the K0-K0.5 section), and the time axis is divided into time period units with timestamps (e.g., T001 for 00:00-00:15). By associating and mapping the highway grid with the time period units, spatiotemporal grid units are formed (e.g., ID-001 and T001 combined to form ST-001). Spatiotemporal identifiers including road segment IDs and timestamps are assigned, spatial locations are bound to time period ranges, a structured spatiotemporal grid system (e.g., two-dimensional table storage) is generated, and transmitted to the data mapping unit, providing a unified benchmark for spatiotemporal alignment of multi-source data.

[0024] The data mapping unit includes: Extracting spatiotemporal attribute information from multi-source data; Spatiotemporal attribute information includes the spatial coordinate range and road segment association information of basic geographic data, the collection timestamp and monitoring road segment range of dynamic traffic data, the monitoring time and spatial monitoring point location of environmental perception data, and the spatiotemporal coordinates of accident occurrence of historical accident data. A unified spatiotemporal coordinate system is established based on road segment IDs and timestamps in the spatiotemporal grid system; A mapping relationship is constructed based on multi-source data and a spatiotemporal grid system.

[0025] In one embodiment, the data mapping unit first extracts core spatiotemporal attributes from multi-source data, including the spatial coordinate range of basic geographic data (such as the latitude and longitude of the road segment's starting and ending points), road segment association information, and the collection timestamp of dynamic traffic data (such as 2025-08-14). The system collects data including the monitoring time and spatial location of monitoring points (such as the coordinates of the meteorological station at K8+500) for the monitored road segment range and environmental perception data, as well as the spatiotemporal coordinates of historical accident data. Then, based on the road segment ID and timestamp of the spatiotemporal grid system, a unified spatiotemporal coordinate system including spatial and temporal dimensions is constructed to provide a benchmark for all data. Finally, a mapping relationship is established by matching spatiotemporal attributes with the grid system. Basic geographic data is matched to the corresponding highway grid unit according to spatial coordinate range; dynamic traffic data achieves spatiotemporal dual-dimensional association through collection timestamps and monitored road segment ranges; environmental perception data is matched to the corresponding time period and grid unit through monitoring time and location; and historical accident data is directly located to the spatiotemporal grid unit through spatiotemporal coordinates. Ultimately, a feature set with spatiotemporal labels, including road segment ID, timestamp, data type, and corresponding data values, is generated, providing standardized input for subsequent processing and ensuring accurate alignment of multi-source data within a unified spatiotemporal framework.

[0026] The mapping relationships include the first mapping relationship, the second mapping relationship, and the third mapping relationship; The first mapping relationship associates the spatiotemporal attribute information of multi-source data with the road segment ID and timestamp in the spatiotemporal grid system, including: Basic geographic data is matched to the corresponding highway grid cells through spatial coordinate range and road segment association information; Dynamic traffic data is matched to the corresponding time period unit and highway grid unit by collecting timestamps and monitoring road segment ranges; Environmental perception data is matched to the corresponding time period units and highway grid units by monitoring time and spatial monitoring point locations; Historical accident data is matched to the corresponding time period and highway grid unit by the spatiotemporal coordinates of the accident occurrence.

[0027] In one embodiment, the mapping relationship includes a first mapping relationship, a second mapping relationship, and a third mapping relationship. The first mapping relationship achieves precise data alignment by associating the spatiotemporal attributes of multi-source data with the road segment ID and timestamp of the spatiotemporal grid system. Basic geographic data extracts spatial coordinate ranges (e.g., starting point latitude and longitude 116.3°E / 39.9°N, ending point latitude and longitude 116.5°E / 39.8°N) and road segment association information. Based on a preset spatial interval (e.g., 500 meters), GIS spatial overlay analysis is used to match it to a highway grid unit (e.g., road segment ID-020) that includes its coordinate range, and binds static spatial attributes. Dynamic traffic data (e.g., 2025-08-14)... Traffic flow monitored at 10:00:00 (K15-K20 section) was extracted, along with the timestamp and monitoring range. The timestamp was matched to a time period unit (T040) at a preset time interval (e.g., 15 minutes), and the road segment range was matched to a highway grid unit (ID-030 to ID-040), forming a spatiotemporal two-dimensional correlation. Environmental sensing data (e.g., weather station data at K8+300, dated August 14, 2025) was also included. Rainfall data collected at 09:30 was used to extract monitoring time and location. Time was matched to time period unit (T038), and location was matched to highway grid unit (ID-017) to achieve spatiotemporal positioning. Historical accident data (such as a rear-end collision that occurred at K25+100 at 14:20 on May 10, 2023) was used to extract spatiotemporal coordinates and directly matched to the corresponding spatiotemporal grid unit (ST-050-056 unit, which is a combination of ID-050 and T056). Through the above matching logic, the spatiotemporal attributes of various data were mapped to road segment ID and timestamp, generating standardized records including road segment ID, timestamp, data type and value, providing a unified benchmark for subsequent processing.

[0028] The second mapping relationship specifies matching logic for different types of data, including basic geographic data, dynamic traffic data, environmental perception data, and historical accident data. Basic geographic data is statically matched based on road segment ID; Dynamic traffic data and environmental perception data are dynamically matched in both spatiotemporal dimensions; Historical accident data is backtracked and matched according to the spatiotemporal coordinates of the accident occurrence.

[0029] In one embodiment, the second mapping relationship implements differentiated matching logic based on the characteristics of multi-source data types to achieve accurate alignment and efficient fusion: Basic geographic data (such as highway alignment parameters and road facility information) is stable over a long period. During the initialization phase, it establishes a one-to-one correspondence between road segment association information (such as road number and mileage marker) and the road segment ID in the spatiotemporal grid system (e.g., the curve radius of segment K10-K15 corresponds to ID-020), forming a static mapping table of "road segment ID → basic geographic data," thus completing the fixed binding of road infrastructure attributes and spatial grids. Dynamic traffic data (such as real-time traffic flow) and environmental perception data (such as rainfall) change in real time, employing spatiotemporal dual-dimensional dynamic matching. Taking vehicle speed distribution data as an example, the monitoring time (e.g., 2025-08-14) is extracted. At 10:15:00, data is mapped to the time period unit (T041) at 15-minute intervals. Simultaneously, the monitored road segment range (K20-K25) is matched with highway grid units (ID-040 to ID-050) to achieve dynamic association of "timestamp + road segment ID." Environmental data (such as road surface temperature at K5+200) is spatiotemporally bound using monitoring time (09:45→T039) and location (ID-010 grid). Historical accident data (such as the rollover at K30+800 on June 5, 2024) is also included. The spatiotemporal coordinates (road segment ID-060, timestamp T065) of the accident are extracted. The corresponding grid cell (ST-060-065) is located back through the spatiotemporal identifier. The accident label (whether the accident occurred and the accident type) is associated with historical time period data to form a retrospective mapping of annotation information. Through the above logic, basic geographic data achieves static spatial binding, dynamic and environmental data achieves spatiotemporal dynamic association, and historical accident data achieves annotation retrospective, ultimately generating a structured feature set including data type, spatiotemporal label and attribute value.

[0030] The third mapping relationship binds the matched multi-source data with the corresponding highway grid units and time period units to generate a feature set with spatiotemporal labels; The feature set includes road segment ID, timestamp, multi-source data type identifier, and the corresponding data value of the multi-source data type identifier; The feature set is transmitted to the data association processing unit.

[0031] In one embodiment, the third mapping relationship generates a standardized feature set through precise binding of multi-source data with spatiotemporal grid units: First, based on the matching results of the first and second mapping relationships (e.g., basic geographic data matched to road segment ID-020, dynamic traffic data matched to road segment ID-040 and time segment unit T041), the basic geographic data, dynamic traffic data, environmental perception data, and historical accident data are uniquely bound to their corresponding highway grid units (road segment IDs) and time segment units (time stamps). For example, the curve radius of the K10-K15 section is bound to ID-020 and 2025-08-14. Traffic flow at K20-K25 at 10:15:00 is associated with ID-040 and T041; road surface temperature at K5+200 at 09:45 is associated with ID-010 and T039; and the rollover accident at K30+800 on June 5, 2024 is associated with ID-060 and T065. Next, a feature set with spatiotemporal labels is generated according to a unified structure, including the road segment ID identifying the highway grid unit (e.g., ID-020), the timestamp identifying the time period unit (e.g., T041), and a multi-source data type identifier distinguishing data categories. For example, "dynamic traffic data" and data values ​​storing specific parameters (such as traffic flow of 1800 vehicles / hour) are ultimately transmitted in real time to the data association processing unit. This provides standardized input for subsequent spatiotemporal matching and feature vector generation by combining historical accident data, thus realizing the efficient fusion and analysis of multi-source data under a unified spatiotemporal framework.

[0032] The data association processing unit includes: Extract the spatiotemporal coordinates and accident tags of historical accident data, as well as the road segment ID and timestamp corresponding to the spatiotemporal coordinates of the accident. Accident labels include whether an accident has occurred and the type of accident; Spatiotemporal matching of feature sets with historical accident data using road segment IDs and timestamps includes: Spatial dimension matching is performed based on the road segment ID in the feature set and the road segment ID corresponding to the historical accident data. If the road segment IDs match, the spatial dimension matching is considered successful. The time dimension is matched based on the timestamps in the feature set and the timestamps corresponding to the historical accident data. If the timestamps are in the same time period, the time dimension is considered to be matched successfully. When both the spatial and temporal dimensions are successfully matched, the feature set is determined to be a successful spatiotemporal match with the historical accident data. The successfully matched accident labels are associated with the corresponding features, generating a spatiotemporal feature vector with accident labels, and the spatiotemporal feature vector is transmitted to the prediction module; The spatiotemporal feature vector includes spatiotemporal features, data features, and corresponding accident labels.

[0033] In one embodiment, the data association processing unit generates labeled feature vectors by spatiotemporally matching historical accident data with a feature set. First, it extracts the spatiotemporal coordinates of the accident occurrence from the historical accident data (corresponding to the road segment ID and timestamp, e.g., the accident point corresponds to road segment ID-040 and timestamp T041, which is based on a preset 15-minute time interval, corresponding to the 10:15-10:30 time period unit) and accident tags (including whether an accident occurred and its type, such as rear-end collision). Then, it performs spatiotemporal matching in two stages: spatially, it compares the feature set with the road segment IDs of the historical accident data; if they are completely identical (e.g., both are ID-040), the spatial matching is successful; temporally, it determines whether the timestamps of the two belong to the same time period unit based on the 15-minute interval (e.g., T041 corresponds to the same time period). If the timestamps within the time frame are all considered to match, and the corresponding 15-minute intervals are consistent, then the time matching is successful. When both the spatial and temporal dimensions are successfully matched (e.g., the road segment ID-040 and timestamp T041 in the feature set completely correspond to the accident data), the accident label (e.g., "Accident occurred, type: rear-end collision") is associated with the feature set, generating a spatiotemporal feature vector including spatiotemporal features (road segment ID and timestamp), data features (e.g., traffic flow of 1800 vehicles / hour and road surface temperature of 25℃), and the accident label (e.g., structured data {"road segment ID":"ID-040","timestamp":"T041","traffic flow":1800,"road surface temperature":25,"accident occurred":true,"accident type":"rear-end collision"}), and transmitted to the prediction module in real time.

[0034] The prediction unit includes: Spatiotemporal feature vectors generate rolling feature vectors based on a preset time window; The preset time window includes a short-term window, a medium-term window, and a long-term window, which respectively generate short-term feature vectors, medium-term feature vectors, and long-term feature vectors. For indicators of multi-source data, the mean, variance and extreme value characteristics are calculated within the corresponding time period of the preset time window; Indicators from multi-source data include real-time traffic flow, vehicle speed distribution, lane change frequency, rainfall, visibility, and road surface temperature; The rolling feature vector is calculated based on the preset risk assessment engine to obtain the accident risk prediction result; The risk assessment engine includes a pre-defined set of assessment logics, which includes: When the rainfall exceeds the first preset threshold and the curve radius is less than a meters, the combined cause assessment logic for increasing the accident probability threshold is used. The dynamic traffic assessment logic raises the risk level by one level when the real-time traffic flow is greater than b vehicles / hour and the lane change frequency is greater than c times / minute.

[0035] In one embodiment, a preset time window is first set according to the needs of highway safety analysis. The short-term window is set to 15 minutes (corresponding to the real-time traffic data update frequency), the medium-term window to 1 hour (covering typical traffic flow cycles), and the long-term window to 12 hours (meeting the needs of day-night environmental change analysis). Short-term, medium-term, and long-term feature vectors are generated respectively. For six types of multi-source data indicators—real-time traffic flow, vehicle speed distribution, lane change frequency, rainfall, visibility, and road surface temperature—three statistical features are calculated within each time window: mean (e.g., average traffic flow within 15 minutes), variance (e.g., the degree of fluctuation in lane change frequency), and extreme values ​​(e.g., the highest hourly rainfall). This forms a rolling feature vector including spatiotemporal labels (e.g., including the average traffic flow of 1800 vehicles / hour and the extreme rainfall of 30mm for time period T041). The risk assessment engine has a built-in preset set of assessment logic, including combined causal assessment logic and dynamic traffic assessment logic. The evaluation logic of Hezhiyin is as follows: when the rainfall in the environmental perception data is greater than 50 mm / h (the first preset threshold is set at 50 mm / h, based on the highway meteorological disaster level standard), and the curve radius in the basic geographic data is less than 500 meters (a=500, the general curve radius safety threshold in highway design specifications), the accident probability threshold enhancement mechanism is triggered, increasing the basic probability of accident risk of the current spatiotemporal grid unit by 40%. The dynamic traffic evaluation logic is as follows: if the real-time traffic flow in the dynamic traffic data is greater than 2000 vehicles / hour (b=2000, the saturation flow threshold of the one-way lane of the road section) and the lane change frequency is greater than 50 times / minute (c=50, the high-risk lane change frequency threshold based on historical accident data statistics), the risk level is increased by one level from the basic level (e.g., from "low risk" to "medium risk"). Based on the above feature vectors and evaluation logic, prediction results including spatiotemporal features, multi-dimensional statistical features, and risk levels are generated.

[0036] Accident risk prediction results include the probability value of accident occurrence and the corresponding risk level; The correlation analysis between historical accident data and rolling feature vectors for accident occurrence probability values ​​includes: When the conditions of rainfall exceeding the first preset threshold and curve radius being less than d meters are triggered, a real-time probability value is generated by superimposing the probability increase amount on the baseline probability value obtained from the accident occurrence rate statistics of historical spatiotemporal grid units of the same type. The risk level is divided into multiple threshold ranges based on the assessment logic, and is associated with the real-time probability value and the triggering status of the dynamic traffic assessment logic, specifically including: The multi-level threshold ranges include low-risk threshold ranges, medium-risk threshold ranges, and high-risk threshold ranges. The boundary values ​​of the low-risk, medium-risk, and high-risk threshold ranges are correlated with the parameter thresholds in the evaluation logic set, including: In the initial state, when the probability value of an accident falls into the low-risk threshold range, it corresponds to a low-risk level; when the probability value of an accident falls into the medium-risk threshold range, it corresponds to a medium-risk level; and when the probability value of an accident falls into the high-risk threshold range, it corresponds to a high-risk level. When the level adjustment condition in the assessment logic is triggered, the threshold interval boundary value corresponding to the current risk level is adjusted so that the risk level enters the corresponding level interval according to the adjustment condition.

[0037] In one embodiment, the accident risk prediction result generated by the prediction unit includes an accident occurrence probability value and a risk level. The accident occurrence probability value is generated based on historical accident data statistics and real-time feature vector correlation analysis. First, historical grid cells of the same type that match the current spatiotemporal grid cell features (such as a curve radius less than 500 meters and rainfall greater than 50 mm / h) are extracted, and their accident occurrence rate over the past 12 months is calculated as a baseline probability value (for example, the accident occurrence rate of this type of grid cell is 5% according to statistics). When real-time data triggers the combination of "rainfall greater than 50 mm / h and curve radius less than 500 meters", the prediction is further refined. Because the evaluation logic (the first preset threshold is set at 50 mm / h according to the highway meteorological disaster level standard, d=500, which is the general curve radius safety critical value in highway design specifications) is based on a 40% probability increase fitted from historical accident data (the accident probability under this combination of conditions is 40% higher than that in ordinary scenarios), a real-time probability value is generated by superimposing it on the baseline probability (e.g., 5%×(1+40%)=7%). The risk level is divided according to a multi-level threshold range and associated with the evaluation logic. Based on the historical accident probability distribution and highway safety management needs, a three-level threshold range is defined: the low-risk range is 0%~10%. % (corresponding to conventional scenarios with no significant risk factors), medium risk range is 10%–30% (corresponding to scenarios triggering a single risk factor or a low-intensity combination of factors), and high risk range is 30% and above (corresponding to scenarios triggering a high-intensity combination of factors or multiple dynamic risk factors); if the real-time feature vector simultaneously triggers dynamic traffic assessment logic (e.g., real-time traffic flow greater than 2000 vehicles / hour and lane change frequency greater than 50 times / minute, where 2000 vehicles / hour is the saturation flow threshold for a one-way lane on the road segment, and 50 times / minute is the critical value for high-risk lane change frequency based on historical accident data statistics), the current risk level is... The corresponding threshold range boundary value will be adjusted upward by one level (for example, the upper limit of the low-risk range will be increased from 10% to 20%, and the original 15% probability value will be dynamically adjusted from "medium risk" to "low risk" due to the adjustment of the lower limit of the medium-risk range to 20%). By combining the real-time probability value and the assessment logic trigger status, a prediction result including spatiotemporal coordinates (such as road segment ID-040, timestamp T041), accident occurrence probability (such as 7%) and risk level (such as "low risk") will be generated to provide differentiated safety warnings for highway management departments (such as real-time sound and light alarms triggered by high-risk levels, and video surveillance patrols initiated by medium-risk levels).

[0038] This invention significantly improves the effectiveness of highway accident prediction through multi-dimensional data integration and spatiotemporal coupling analysis. It integrates multi-source data from basic geography, dynamic traffic, environmental perception, and historical accidents to construct a three-dimensional dataset including road facilities, traffic flow, environmental parameters, and historical accident characteristics. By establishing a unified spatiotemporal coordinate system through spatiotemporal grid division, it achieves differentiated data fusion. Static road parameters are accurately located based on road segment IDs, dynamic traffic and environmental data are correlated through spatiotemporal dimensions, and historical accident data is back-matched according to spatiotemporal coordinates, generating a structured feature set including spatiotemporal labels. This solves the problems of inconsistent spatiotemporal benchmarks and inefficient data fusion in traditional solutions. In the prediction stage, statistical features of mean and variance are extracted through multiple time windows. Combined with risk assessment logic, it analyzes the multi-factor coupling of meteorological conditions, road alignment, and traffic flow parameters to achieve accurate calculation of accident probability and dynamic adjustment of risk levels. It finely characterizes the accident risk features of each spatiotemporal grid unit, improving the spatiotemporal resolution and robustness of predictions. This provides traffic management departments with real-time, scenario-based risk assessment results, supporting efficient accident early warning, lane scheduling, and resource allocation, and promoting the intelligent and precise advancement of highway safety management.

[0039] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0040] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A highway accident prediction system based on big data, characterized in that... It includes a data acquisition module, a data fusion module, and a prediction module; The acquisition module is used to acquire multi-source data for highway accident prediction and transmit it to the fusion module; The fusion module is used to perform spatiotemporal alignment processing on the multi-source data through the processing unit, generate feature vectors, and transmit them to the prediction module; The prediction module is used to generate accident risk prediction results through the prediction unit based on the real-time input feature vector.

2. The highway accident prediction system based on big data as described in claim 1, characterized in that: The multi-source data includes basic geographic data, dynamic traffic data, environmental perception data, and historical accident data; The basic geographic data includes highway alignment parameters and road infrastructure information; The dynamic traffic data includes real-time traffic flow, vehicle speed distribution, and lane change frequency. The environmental perception data includes real-time meteorological parameters, including rainfall, visibility, and road surface temperature. The historical accident data includes the spatiotemporal coordinates of the accident, whether the accident occurred, and the type of accident. The accident types include rear-end collisions, rollovers, and collisions.

3. The highway accident prediction system based on big data as described in claim 2, characterized in that: The processing unit includes a spatiotemporal grid partitioning unit, a data mapping unit, and a data association processing unit; The spatiotemporal grid partitioning unit includes: Obtain the geographical scope and time dimension requirements of the highway, and set the preset spatial interval and preset time interval; The highway is divided into several continuous highway grid units based on the preset spatial interval, and a road segment ID is assigned to each highway grid unit. The time axis is divided into several consecutive time period units based on the preset time interval, and a timestamp is assigned to each time period unit. Integrating the aforementioned highway grid units and time period units to generate a spatiotemporal grid system, specifically including: Establish the association mapping relationship between the highway grid unit and the time period unit, combine each highway grid unit with all the time period units in terms of spatiotemporal dimensions to form a spatiotemporal grid unit, and assign a spatiotemporal identifier to each spatiotemporal grid unit; The spatiotemporal identifier is associated with the road segment ID of the highway grid unit corresponding to the spatiotemporal grid unit and the timestamp of the corresponding time period unit; By binding the spatial attributes of the highway grid unit to the temporal attributes of the time period unit through the spatiotemporal identifier, a structured set including all the spatiotemporal grid units is formed; The spatiotemporal grid system is transmitted to the data mapping unit. The spatiotemporal grid system includes road segment IDs and timestamps.

4. The highway accident prediction system based on big data as described in claim 3, characterized in that: The data mapping unit includes: Extract the spatiotemporal attribute information from the multi-source data; The spatiotemporal attribute information includes the spatial coordinate range and road segment association information of basic geographic data, the collection timestamp and monitoring road segment range of dynamic traffic data, the monitoring time and spatial monitoring point location of environmental perception data, and the spatiotemporal coordinates of accident occurrence of historical accident data. A unified spatiotemporal coordinate system is established based on the road segment IDs and timestamps in the aforementioned spatiotemporal grid system; A mapping relationship is constructed based on the multi-source data and the spatiotemporal grid system.

5. The highway accident prediction system based on big data as described in claim 4, characterized in that: The mapping relationship includes a first mapping relationship, a second mapping relationship, and a third mapping relationship; The first mapping relationship, which associates the spatiotemporal attribute information of the multi-source data with the road segment ID and timestamp in the spatiotemporal grid system, includes: The basic geographic data is matched to the corresponding highway grid unit through spatial coordinate range and road segment association information; The dynamic traffic data is matched to the corresponding time period unit and highway grid unit by collecting timestamps and monitoring road segment ranges; The environmental perception data is matched to the corresponding time period unit and highway grid unit by monitoring time and spatial monitoring point location; The historical accident data is matched to the corresponding time period unit and highway grid unit by the spatiotemporal coordinates of the accident occurrence.

6. The highway accident prediction system based on big data as described in claim 5, characterized in that: The second mapping relationship, for the types of basic geographic data, dynamic traffic data, environmental perception data, and historical accident data, specifies matching logic, specifically including: The basic geographic data is statically matched according to road segment ID; The dynamic traffic data and environmental perception data are dynamically matched in both spatiotemporal dimensions. The historical accident data is backtracked and matched according to the spatiotemporal coordinates of the accident occurrence.

7. The highway accident prediction system based on big data as described in claim 6, characterized in that: The third mapping relationship binds the matched multi-source data with the corresponding highway grid units and time period units to generate a feature set with spatiotemporal labels; The feature set includes road segment ID, timestamp, multi-source data type identifier, and the data value corresponding to the multi-source data type identifier; The feature set is transmitted to the data association processing unit.

8. The highway accident prediction system based on big data as described in claim 7, characterized in that: The data association processing unit includes: Extract the spatiotemporal coordinates and accident tags of the historical accident data, where the spatiotemporal coordinates of the accident correspond to the road segment ID and timestamp; The accident label includes whether an accident has occurred and the type of accident. Performing spatiotemporal matching of the feature set with historical accident data using the road segment ID and timestamp includes: Spatial dimension matching is performed based on the road segment IDs in the feature set and the road segment IDs corresponding to historical accident data. If the road segment IDs match, the spatial dimension matching is determined to be successful. Based on the timestamps in the feature set and the timestamps corresponding to historical accident data, time dimension matching is performed. If the timestamps are in the same time period unit, the time dimension matching is determined to be successful. When both the spatial and temporal dimensions are successfully matched, it is determined that the feature set is successfully matched spatiotemporally with the historical accident data; The successfully matched accident labels are associated with the corresponding features to generate a spatiotemporal feature vector with the accident labels, and the spatiotemporal feature vector is transmitted to the prediction module. The spatiotemporal feature vector includes spatiotemporal features, data features, and corresponding accident labels.

9. The highway accident prediction system based on big data as described in claim 8, characterized in that: The prediction unit includes: The spatiotemporal feature vector is used to generate a scrolling feature vector based on a preset time window; The preset time window includes a short-term window, a medium-term window, and a long-term window, which respectively generate short-term feature vectors, medium-term feature vectors, and long-term feature vectors. For the indicators of the multi-source data, the mean characteristics, variance characteristics and extreme value characteristics are calculated within the corresponding time period of the preset time window; The indicators of the multi-source data include real-time traffic flow, vehicle speed distribution, lane change frequency, rainfall, visibility, and road surface temperature. The rolling feature vector is calculated based on a preset risk assessment engine to obtain the accident risk prediction result; The risk assessment engine includes a preset set of assessment logics, which includes: When the rainfall exceeds the first preset threshold and the curve radius is less than a meters, the combined cause assessment logic for increasing the accident probability threshold is used. The dynamic traffic assessment logic raises the risk level by one level when the real-time traffic flow is greater than b vehicles / hour and the lane change frequency is greater than c times / minute.

10. The highway accident prediction system based on big data as described in claim 9, characterized in that: The accident risk prediction results include the probability value of accident occurrence and the corresponding risk level; The probability value of the accident occurrence is based on the correlation analysis between historical accident data and rolling feature vectors, including: When the conditions of rainfall exceeding the first preset threshold and curve radius being less than d meters are triggered, a real-time probability value is generated by superimposing the probability increase amount on the baseline probability value obtained from the accident occurrence rate statistics of historical spatiotemporal grid units of the same type. The risk levels are divided into multiple threshold ranges based on the assessment logic, and are associated with real-time probability values ​​and the triggering status of dynamic traffic assessment logic, specifically including: The multi-level threshold range includes a low-risk threshold range, a medium-risk threshold range, and a high-risk threshold range; The boundary values ​​of the low-risk, medium-risk, and high-risk threshold intervals are associated with the parameter thresholds in the evaluation logic set, including: In the initial state, when the probability value of an accident falls into the low-risk threshold range, it corresponds to a low-risk level; when the probability value of an accident falls into the medium-risk threshold range, it corresponds to a medium-risk level; and when the probability value of an accident falls into the high-risk threshold range, it corresponds to a high-risk level. When the level adjustment condition in the assessment logic is triggered, the threshold interval boundary value corresponding to the current risk level is adjusted so that the risk level enters the corresponding level interval according to the adjustment condition.

Citation Information

Cited By

  • Tunnel traffic safety early warning system and method based on traffic internet of things

    CN121564942A