A highway weather data processing method and early warning system
Patent Information
- Application Number
- CN202610802059.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-15
Smart Images

Figure CN122761569A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of highway meteorological service technology, and in particular to a highway meteorological data processing method and early warning system. Background Technology
[0002] With the intensification of global climate change and the increasing frequency of extreme weather events, highway networks, due to their linear distribution and complex terrain, are particularly sensitive to severe weather events such as heavy rainfall, snow, and fog, posing a serious challenge to their safe operation. Traditional meteorological early warning services are mostly based on administrative divisions or macro-meteorological grids, which are difficult to match the segment-level precision required for highway management. Currently, the industry generally faces several bottlenecks in promoting the integration of meteorology and transportation applications: First, there is a significant inconsistency in data formats. Meteorological grid forecast data, such as NetCDF format, differs from the vector road network and risk point data of transportation departments in terms of format and spatiotemporal resolution. Relying on manual conversion and alignment is inefficient and difficult to guarantee accuracy, and the lack of data quality control further affects the reliability of subsequent analysis. Second, most current early warning models are general-purpose meteorological products that fail to deeply integrate highway-specific risk factors, such as historical disaster sites, leading to a disconnect between early warning results and actual road segment risks, resulting in false alarms and missed warnings. Furthermore, the generation and dissemination chain of early warning information is long, making it difficult to support rapid response to disasters.
[0003] More importantly, the existing technology system not only restricts the leap from administrative region forecasting to precise handling of highway sections in highway weather warnings, but also fails to provide intelligent driving safety guarantees for logistics companies and drivers. It is difficult to recommend the safest and most efficient routes during the route planning stage, and it is also impossible to achieve optimal route matching that balances safety and efficiency. This makes it difficult to avoid safety risks in the logistics transportation process in advance, and may also increase transportation time and costs due to improper route selection, further highlighting the urgency of upgrading and improving the existing technology.
[0004] Therefore, there is an urgent need in this field for a new generation of technical solutions that can achieve deep fusion of multi-source data, have the ability to make intelligent judgments based on scenarios, and support open collaboration. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a highway meteorological data processing method and an early warning system to solve the technical problems of low efficiency of multi-source data fusion and disconnect between early warning models and actual road segment scenarios in the prior art.
[0006] One aspect of the present invention provides a method for highway meteorological data processing and early warning, the method comprising the following steps: Acquire multi-source heterogeneous raw data for the target highway segment. The multi-source heterogeneous raw data includes at least meteorological grid data from external meteorological departments, traffic and highway vector data containing road segment risk attribute characteristics from traffic management departments, and historical meteorological and disaster case data stored locally and associated with historical early warning analysis periods.
[0007] The multi-source heterogeneous raw data is standardized and fused to obtain standardized fused data for the target highway segment. The standardized fusion process includes: spatially matching the meteorological grid data with the target highway segment to generate segment-level time-series meteorological data; performing geocoding-based spatial location matching and time window aggregation on the traffic highway vector data and the historical meteorological and disaster case data to form segment-enhanced attribute data; and performing quality verification and missing value imputation on the segment-level time-series meteorological data and the segment-enhanced attribute data.
[0008] Based on the standardized fusion data, for the early warning analysis period, a multi-dimensional feature vector associated with the target highway segment is constructed, including: selecting key meteorological elements within the early warning analysis period from the segment-level time-series meteorological data and calculating their statistics to construct a meteorological feature sub-vector; extracting attribute features representing the inherent risk level of the segment from the segment enhanced attribute data to construct a segment risk attribute feature sub-vector; extracting regular features representing historical disaster patterns from the segment enhanced attribute data to construct a historical disaster feature sub-vector; and concatenating and normalizing the meteorological feature sub-vector, the segment risk attribute feature sub-vector, and the historical disaster feature sub-vector to form the multi-dimensional feature vector.
[0009] The multidimensional feature vector is input into a pre-trained random forest early warning model to obtain the initial early warning level probability distribution corresponding to the target highway segment.
[0010] The initial warning level of the target highway section is determined based on the probability distribution of the initial warning level and the preset probability threshold rules.
[0011] Based on the feature subvector of the road segment risk attribute, if the target highway segment is determined to meet the preset high-risk attribute conditions, the initial warning level will be increased by at least one level to generate and output the final warning result.
[0012] In some embodiments of the present invention, the meteorological grid data includes surface forecast element data and precipitation data; the surface forecast element data includes at least air pressure, humidity, wind direction and wind speed, and its original transmission format is the Network Common Data Format (NetCDF); The traffic and highway vector data includes highway risk point data, historical disaster damage data, and basic road network data; The historical meteorological and disaster case data includes historical meteorological element files stored in a common network data format.
[0013] In some embodiments of the present invention, the meteorological grid data is spatially matched with the target highway segment to generate segment-level time-series meteorological data, including: Using nearest neighbor matching, bilinear interpolation, or inverse distance weighted interpolation algorithms, the corresponding meteorological element values are extracted and calculated from the meteorological grid data based on the geographical coordinates of the center point or representative point of the target highway segment.
[0014] In some embodiments of the present invention, the meteorological grid data is spatially matched with the target highway segment to generate segment-level time-series meteorological data, including: The target road segment in the traffic and highway vector data is sampled at equal intervals according to a preset sampling interval to generate a sequence of multiple sampling points continuously distributed along the target road segment; Based on the spatial matching relationship between multiple sampling points and the corresponding grid cells of the meteorological grid data, the meteorological element values of each sampling point at the same time or within the same time window are obtained. According to a preset aggregation strategy, the meteorological element values corresponding to each sampling point within the same target highway segment are aggregated to generate the segment-level time-series meteorological data; the aggregation strategy includes any one of interval mean, interval extreme value or length-weighted average.
[0015] In some embodiments of the present invention, the spatial matching relationship is established and reused in the following manner: A spatial index is constructed based on the latitude and longitude coordinates of the meteorological grid data. Nearest neighbor matching is performed on each sampling point to determine the meteorological grid index corresponding to each sampling point. Store the mapping relationship between sampling points and meteorological grid index; When processing subsequent meteorological grid data, consistency checks are performed based on the dimensional dimensions, latitude and longitude boundary ranges, and representative coordinate sequence summary information of the meteorological grid. The mapping relationship is reused when the verification is consistent, and the mapping relationship is re-established when the verification is inconsistent.
[0016] In some embodiments of the present invention, the step of performing geocoding-based spatial location matching and time window aggregation between the traffic highway vector data and the historical meteorological and disaster case data to form road segment enhanced attribute data includes: Based on the nearest neighbor spatial matching algorithm, the geographical location of the historical disaster case data is associated with the nearest target highway segment defined in the traffic highway vector data; and based on the preset warning time window, the associated historical disaster case data is collected into the corresponding warning time window according to the relative relationship between its occurrence time and the warning analysis time.
[0017] In some embodiments of the present invention, the step of performing quality verification and missing value imputation on the road segment-level time-series meteorological data and road segment enhanced attribute data includes: Outliers in the road segment-level time-series meteorological data are identified and removed based on the three Sigma principle; and a linear interpolation algorithm is used to fill in the data points that have been removed or are missing.
[0018] In some embodiments of the present invention, the pre-training steps of the random forest early warning model include: Obtain a training sample set, where each sample corresponds to a historical early warning analysis period and contains the multidimensional feature vector obtained based on the standardized fusion data of the historical early warning analysis period; For each sample, an early warning level label is added to the sample based on the severity, frequency, and number of associated risk points of the disaster event that occurred on the target highway section during the corresponding historical early warning analysis period. The multidimensional feature vector is input into the random forest early warning model for training, and the predicted probability distribution of the early warning level corresponding to the sample is output. Based on the deviation between the predicted probability distribution of the warning level and the warning level label, a weighted cross-entropy loss function is constructed; wherein, the weights assigned to different warning levels in the weighted cross-entropy loss function are inversely proportional to the frequency of occurrence of each warning level in the training sample set; the training results of the random forest model are evaluated based on the weighted cross-entropy loss function, and the training parameters of the random forest model are optimized. When the weighted cross-entropy loss function converges to a stable interval on the validation set, or when the macro F1 score of the random forest early warning model on the validation set reaches its optimum, training is stopped, and the final parameters of the random forest early warning model are fixed to obtain the pre-trained random forest early warning model.
[0019] On the other hand, the present invention also provides a highway meteorological data processing and early warning system, the system comprising: The multi-source data access terminal is used to obtain multi-source heterogeneous raw data of the target highway section from external meteorological departments, traffic management departments and local storage. A data fusion processing server is communicatively connected to the multi-source data access terminal and is used to perform standardized fusion processing on the multi-source heterogeneous raw data to generate standardized fusion data corresponding to the target highway segment. An early warning analysis server, which is communicatively connected to the data fusion processing server, is used to perform early warning analysis calculations based on the standardized fusion data and generate early warning analysis results. The early warning service publishing terminal communicates with the early warning analysis server to receive the early warning analysis results and output the final early warning result; The data fusion processing server, the early warning analysis server, and the early warning service publishing terminal operate in concert and are configured to execute the steps of the method described above.
[0020] This invention provides a highway meteorological data processing method and early warning system. The method includes: acquiring multi-source heterogeneous raw data of a target highway segment, including meteorological grid data, traffic highway vector data, and historical meteorological and disaster case data; performing standardized fusion processing on the raw data, including format conversion, spatiotemporal alignment, and intelligent cleaning, to generate standardized fused data; constructing a multi-dimensional feature vector based on the fused data, integrating meteorological features, road segment risk attribute features, and historical disaster features, and inputting it into a pre-trained random forest early warning model to obtain an initial early warning level probability distribution; determining the initial early warning level according to a preset probability threshold rule; and, based on the road segment risk attribute features, upgrading the initial early warning level of corresponding road segments that meet preset high-risk attribute conditions, generating and outputting the final early warning result. This invention enables deep fusion and collaborative analysis of meteorological, road network, and historical disaster data, improving the accuracy, timeliness, and scenario adaptability of road segment-level meteorological disaster early warnings.
[0021] Furthermore, when constructing the pre-trained random forest early warning model, a weighted cross-entropy loss function is used for training, where the weights assigned to different early warning levels are inversely proportional to their frequency of occurrence in the training samples. This method addresses the problem of scarce high-risk samples and uneven distribution of different categories in highway early warning scenarios. It enables the model to pay more attention to minority class samples during training, thereby effectively improving the ability to identify high-risk meteorological disasters, reducing the false negative rate, and enhancing the model's practicality and reliability.
[0022] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.
[0023] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description
[0024] The accompanying drawings, which are provided to further illustrate the invention and form part of this application, are not intended to limit the scope of the invention.
[0025] Figure 1 This is a flowchart illustrating the highway meteorological data processing and early warning method according to an embodiment of the present invention.
[0026] Figure 2 This is a schematic diagram of the module functions and data flow of the data layer in the architecture of the highway meteorological data processing and early warning system according to another embodiment of the present invention.
[0027] Figure 3 This is a schematic diagram of the module functions and data flow of the processing layer in the architecture of the highway meteorological data processing and early warning system according to another embodiment of the present invention.
[0028] Figure 4 This is a schematic diagram of the module functions and data flow in the application layer of the highway meteorological data processing and early warning system according to another embodiment of the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0030] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0031] In the field of highway meteorological services, current highway meteorological early warning technology mainly relies on traditional meteorological monitoring and simple data analysis methods. Faced with increasingly frequent, localized, and sudden highway meteorological disasters, these traditional methods are insufficient in terms of early warning accuracy, timeliness, and specificity. Their limitations are mainly reflected in the following three aspects: 1. Low data fusion efficiency and weak multi-source data collaboration capabilities. In existing technical solutions, there are significant issues with inconsistent data formats across multiple sources, including meteorological data, road network risk point data, and historical disaster data. On one hand, data formats and standards are not uniform; for example, there is a lack of standardized interface between meteorological data and structured data from transportation departments, requiring manual conversion and resulting in low processing efficiency. On the other hand, spatiotemporal dimensions are difficult to align. Meteorological data is mostly at the regional hourly granularity, while road network data requires minute-by-minute monitoring down to the road segment level, leading to data matching errors of hundreds of meters or tens of minutes. Furthermore, data quality control lacks automated mechanisms, and handling missing and outlier values relies on manual calibration, resulting in data integrity of less than 90%, directly impacting the accuracy of subsequent analyses.
[0032] 2. The early warning model has poor adaptability, and its accuracy and timeliness are insufficient. Traditional early warning models are mostly general-purpose meteorological warning tools that do not deeply integrate with the characteristics of highway scenarios. First, they fail to consider the differentiated risks of different road sections, lacking targeted modeling for the wind resistance and confluence requirements of road sections with high traffic volume and severe weather impacts, leading to over-warning of ordinary road sections and insufficient response to high-risk road sections. Second, their prediction accuracy is low, relying on empirical formulas or simple statistical models, with an accuracy rate of less than 70% for identifying small- to medium-scale disasters such as rainstorms and road icing, and a false alarm rate exceeding 20%. Third, their response is delayed, with a long average time from data access to early warning issuance, often causing grassroots units to miss the optimal disaster prevention window.
[0033] In view of this, one aspect of the present invention provides a method for highway meteorological data processing and early warning, the method comprising the following steps S101-S106: S101: Obtain multi-source heterogeneous raw data of the target highway segment. This multi-source heterogeneous raw data includes at least meteorological grid data from external meteorological departments, traffic and highway vector data containing road segment risk attribute characteristics from traffic management departments, and historical meteorological and disaster case data stored locally and associated with historical early warning analysis periods.
[0034] S102: Standardize and fuse multi-source heterogeneous raw data to generate standardized fused data for the target highway segment. The standardization and fusion process includes: spatially matching meteorological grid data with the target highway segment to generate segment-level time-series meteorological data; performing geocoding-based spatial location matching and time window aggregation on traffic and highway vector data with historical meteorological and disaster case data to form segment-enhanced attribute data; and performing quality verification and missing value imputation on the segment-level time-series meteorological data and the segment-enhanced attribute data.
[0035] S103: Based on standardized fused data, construct a multi-dimensional feature vector associated with the target highway segment for the early warning analysis period. This includes: selecting key meteorological elements within the early warning analysis period from the segment-level time-series meteorological data and calculating their statistics to construct a meteorological feature sub-vector; extracting attribute features representing the inherent risk level of the segment from the segment enhanced attribute data to construct a segment risk attribute feature sub-vector; extracting regular features representing historical disaster patterns from the segment enhanced attribute data to construct a historical disaster feature sub-vector; and concatenating and normalizing the meteorological feature sub-vector, the segment risk attribute feature sub-vector, and the historical disaster feature sub-vector to form a multi-dimensional feature vector.
[0036] S104: Input the multidimensional feature vector into the pre-trained random forest early warning model to obtain the initial early warning level probability distribution corresponding to the target highway segment.
[0037] S105: Determine the initial warning level of the target highway section based on the probability distribution of the initial warning level and the preset probability threshold rules.
[0038] S106: Based on the feature sub-vector of road segment risk attributes, if the target highway segment is determined to meet the preset high-risk attribute conditions, the initial warning level will be increased by at least one level to generate and output the final warning result.
[0039] In step S101, multi-source heterogeneous raw data of the target highway segment is acquired. Specifically, this step requires extracting raw information related to the target highway segment from three types of data sources, including: meteorological grid data, traffic and highway vector data, and historical meteorological and disaster case data.
[0040] Among them, meteorological grid data consists of meteorological element values on regular grid points covering a geographical area, which are obtained by establishing a communication link with an external meteorological data service system.
[0041] In some embodiments, meteorological grid data is automatically retrieved periodically from a designated meteorological server via File Transfer Protocol (FTP). The retrieved data files are encoded and stored using Network Common Data Format (NetCDF) or Grid Binary Format (GRIB).
[0042] In some embodiments, meteorological grid data includes: Surface forecast data: contains surface meteorological data fields for multiple forecast periods in the future. Key elements include at least sea level pressure, snowfall, 2-meter air temperature, 2-meter relative humidity, and 10-meter U / V wind components.
[0043] Precipitation data: Gridded precipitation forecasts or fused real-time products with a spatial resolution of 2.5 kilometers and a temporal resolution of hourly.
[0044] The aforementioned traffic and highway vector data is vector graphic data with geographic coordinates and attribute tables, which is obtained by calling the open data application programming interface (API) provided by the traffic management department.
[0045] In some embodiments, traffic highway vector data includes: Highway risk point data: High-risk locations in the road network are identified by point or line features. The attribute table includes risk point number, road name, risk type, risk level, and latitude and longitude of the starting and ending points.
[0046] Historical disaster data: Historical disaster information is stored in the form of event records. Its attribute table includes disaster event number, associated risk point number, disaster type, occurrence time, loss description and handling status.
[0047] Road network basic data: The digital map of the highway network is composed of linear elements, and its attribute table includes route number, name, technical grade, number of lanes and management unit information.
[0048] The aforementioned historical meteorological and disaster case data are structured association datasets stored locally and used for training the random forest early warning model.
[0049] In some embodiments, historical meteorological data is stored in the form of NetCDF files, such as time series files including PRATEsfc.nc (surface precipitation rate), TMP2m.nc (2-meter temperature), and ws10m.nc (10-meter wind speed).
[0050] The corresponding historical disaster case data is obtained from the aforementioned historical disaster case data by filtering and defining the time range from the "historical disaster damage data", and is then correlated and aligned with historical meteorological data in time and space to form a case sample library.
[0051] In step S102, the multi-source heterogeneous raw data obtained in step S101 is standardized and fused to generate standardized fused data for the target highway segment, specifically including sub-steps S1021~S1023: Step S1021: Spatial matching of meteorological grid data with target highway segments to generate segment-level time-series meteorological data.
[0052] Specifically, since meteorological grid data is organized using a regular grid format covering a geographical area, while the target highway segment is a continuous linear spatial element, there are differences in their spatial representation. Therefore, it is necessary to convert the meteorological element values in the meteorological grid data into segment-level meteorological data corresponding to the target highway segment through spatial matching. This invention provides two selectable matching methods.
[0053] In some embodiments, a spatial interpolation matching method is employed. Specifically, a meteorological data parsing engine (e.g., the xarray framework based on Python combined with netcdf4 or cfgrib underlying libraries) reads meteorological grid data files in the Network Common Data Format (NetCDF) or Grid Binary Format (GRIB) and converts them into a processable multidimensional array data structure. Subsequently, the latitude and longitude coordinates of the target road segment are extracted based on the geometric centerline or a pre-defined sequence of representative points. For each target point, a bilinear interpolation algorithm is used to calculate the meteorological element value corresponding to the target point based on the meteorological element values of the four corner points of the meteorological grid cell where the target point is located and the positional weight of the target point relative to each corner point; alternatively, an inverse distance weighted interpolation algorithm is used to select neighboring grid points within a preset search radius centered on the target point, and the weights are determined based on the reciprocal of the distance between each neighboring grid point and the target point, obtaining the meteorological element value corresponding to the target point through a weighted average. By repeating the above processing over multiple consecutive time intervals in the time dimension, the hourly updated regional grid meteorological data can be converted into road segment-level time-series meteorological data with corresponding time resolution, based on road segments.
[0054] In other embodiments, a spatial matching method based on road segment discretization and mapping multiplexing is employed. This method further includes the following steps: (1) Road segment discretization and sampling point generation: For the vector line segments of the target highway segment, the total spherical length of the segment is calculated using the Havesing formula in the WGS84 coordinate system; the number of sampling points is determined according to the preset sampling interval, and sampling parameters are generated proportionally in the normalized parameter domain of the road segment. The latitude and longitude coordinates of multiple sampling points evenly distributed along the road segment are obtained by spherical linear interpolation. A globally unique identifier is generated for each sampling point. The above globally unique identifier is composed of the road segment number, sub-segment number and sampling point number.
[0055] (2) Establishing the spatial matching relationship between sampling points and meteorological grids: Extract the latitude and longitude dimension array of meteorological grid data and construct a set of coordinates of two-dimensional regular grid nodes; establish a KD-Tree spatial index based on the set of grid node coordinates, and perform a nearest neighbor query on each sampling point to determine the meteorological grid node with the smallest spherical distance from each sampling point, thereby obtaining the corresponding grid two-dimensional row and column index; establish and store the mapping relationship between sampling point identifiers and grid two-dimensional row and column indexes.
[0056] (3) Reuse and consistency verification of mapping relationships: When establishing a mapping relationship for the first time, the core metadata of the meteorological grid is extracted, including the grid dimension size, latitude and longitude spatial boundaries, and latitude and longitude array header feature sequences. A unique fingerprint of the grid is generated by the SHA-256 hash algorithm and stored in association with the mapping relationship. When processing new meteorological grid data in subsequent batches, the grid fingerprint of the new data is first calculated and compared with the stored fingerprint: if they match, the mapping relationship is directly reused; if they do not match, the mapping relationship is re-established and the fingerprint is updated.
[0057] (4) Generation of road segment-level time-series meteorological data: Meteorological element values of each sampling point are read from the meteorological grid according to the mapping relationship. The element values of all sampling points within the same road segment are merged into a representative value of the road segment using a preset aggregation strategy. The preset aggregation strategy includes mean, maximum value, or length-weighted average. The above processing is applied to multiple consecutive time periods in the time series to finally generate road segment-level time-series meteorological data with road segments as the unit and hourly time granularity.
[0058] Step S1022: Perform geocoding-based spatial location matching and time window aggregation on traffic and highway vector data with historical meteorological and disaster case data to form road segment enhanced attribute data.
[0059] Specifically, this sub-step is used to spatiotemporally correlate discrete historical event information with static road network structure, thereby adding attribute information reflecting the historical risk patterns of each target highway segment.
[0060] In some embodiments, spatial location matching is achieved by constructing geocoded mapping relationships. First, the road network infrastructure data in the traffic and highway vector data is topologically constructed to ensure that each road segment has a unique identifier and accurate geometric information. For historical disaster case data, the coordinates of the recorded occurrence locations are extracted. Subsequently, a spatial index-based nearest neighbor matching algorithm is used to determine the spatially closest target highway segment for each disaster case point, and the corresponding disaster record is associated with that road segment. Simultaneously, the highway risk point data defined in the traffic and highway vector data are associated with the corresponding road segments based on their origin and destination latitude and longitude information.
[0061] Building upon spatial correlation, further time window aggregation is performed. Specifically, time windows related to early warning operations can be set for each target highway segment, such as monthly, seasonal, or time windows divided by the early warning lead time (e.g., 6 hours or 12 hours ahead). Statistical data on historical disaster types, frequency, and average losses occurring within each time window are collected for that segment, and these statistics are added as new attribute fields to the segment attribute table. Furthermore, historical meteorological data can be correlated with corresponding road segments using its timestamp and the spatial matching method from step S1021, forming a chain of related cases: "Road Segment—Time—Meteorological Conditions—Disaster Consequences." Ultimately, each target highway segment not only includes its inherent static attributes such as technical grade and number of lanes, but also its disaster pattern characteristics across historical periods, thus forming multi-dimensional enhanced attribute data for the road segment.
[0062] Step S1023: Perform quality verification and missing value filling on the road segment-level time-series meteorological data and road segment enhanced attribute data.
[0063] Specifically, this step involves automated quality control to address potential outliers and missing values in the data, thereby improving data integrity and availability.
[0064] In some embodiments, the three Sigma principle based on statistical distribution is first used to detect and remove outliers in the road segment-level time-series meteorological data: for each meteorological element (such as hourly precipitation, wind speed) forming a time series, its mean is calculated. with standard deviation The value falls within the interval Data points outside the range are identified as statistical outliers and removed. This method effectively filters out physically unreasonable values caused by sensor false alarms or data transmission errors.
[0065] For data gaps resulting from outlier removal, as well as potential missing data in the original data, this step uses a linear interpolation algorithm to fill in the gaps. Linear interpolation is based on the assumption of local continuity in the time series and uses adjacent valid data values before and after the missing location for estimation.
[0066] Specifically, for an isolated missing point, the estimated value of the missing point is calculated proportionally based on the data values at the preceding and following valid time points, according to its temporal position. This method is also applicable to data segments missing at multiple consecutive time points. Using the latest valid value before the segment begins and the earliest valid value after the segment ends as benchmarks, and based on a linear temporal relationship, the values at each time point within the entire missing segment are estimated sequentially, thereby reconstructing a continuous and reasonable data sequence.
[0067] Through the above quality verification and missing value imputation processes, data with insufficient integrity due to anomalies and missing values can be transformed into standardized data that meets the requirements for subsequent use in random forest early warning models.
[0068] In step S103, based on the standardized fusion data generated in step S102, a multi-dimensional feature vector associated with the target highway segment is constructed for the predetermined early warning analysis period.
[0069] Specifically, the standardized fusion data generated in step S102 includes at least road segment-level time-series meteorological data obtained through spatial matching and road segment enhanced attribute data obtained through spatial location matching and time window aggregation. Based on this, features representing meteorological processes, inherent risk attributes of road segments, and historical disaster patterns are extracted and combined and standardized to form a multi-dimensional feature vector for input to the random forest early warning model. Step S103 may further include sub-steps S1031 to S1034: Step S1031: Select key meteorological elements within the early warning analysis period from the road segment-level time-series meteorological data and calculate their statistics to construct meteorological feature sub-vectors.
[0070] In some embodiments, key meteorological elements include cumulative precipitation, average wind speed, maximum temperature, minimum temperature, and average relative humidity. For the time-series values of these key meteorological elements during the warning analysis period, their corresponding statistical characteristics are calculated. These statistical characteristics include, but are not limited to, the cumulative value of precipitation, the maximum and average wind speed, the range of temperature, and the average humidity during the warning analysis period. The numerical set composed of the above statistical characteristics can serve as a meteorological feature sub-vector to characterize the intensity and variation characteristics of the disaster-causing meteorological factors acting on the target highway section.
[0071] Step S1032: Extract attribute features that characterize the inherent risk level of the road segment from the enhanced attribute data of the road segment to construct a sub-vector of road segment risk attribute features.
[0072] In some embodiments, the attribute features include features extracted from the static attributes of the road segment, including but not limited to seismic fortification level, flood control standard, design traffic flow, number of lanes, and basic risk level. These attribute features are encoded to form a road segment risk attribute feature sub-vector, used to characterize the physical attributes of the road infrastructure itself and its inherent vulnerability level.
[0073] Step S1033: Extract regular features representing historical disaster patterns from the enhanced attribute data of the same road segment to construct a feature sub-vector of historical disasters.
[0074] In some embodiments, statistical analysis is performed on historical disaster cases associated with the target highway segment to extract historical disaster pattern characteristics. These pattern characteristics include, but are not limited to: the cumulative number of disaster events occurring on the target highway segment under similar weather conditions within a preset historical period, the average loss level of historical disaster events, and the time interval between the most recent disaster event and the current early warning analysis. These pattern characteristics form a historical disaster feature sub-vector, used to characterize the exposure degree and consequence patterns of the target highway segment when historically facing disaster threats.
[0075] Step S1034: Concatenate the meteorological feature sub-vector, the road section risk attribute feature sub-vector, and the historical disaster feature sub-vector, and normalize the concatenated feature vector to form a multidimensional feature vector.
[0076] In some embodiments, meteorological feature sub-vectors, road section risk attribute feature sub-vectors, and historical disaster feature sub-vectors are concatenated in a preset order to form a composite feature vector; subsequently, the composite feature vector is normalized. Optionally, a max-min normalization method is used to linearly scale the values of each feature to the [0,1] interval to reduce the impact of differences in the dimensions and value ranges of different features on the model training and inference process. The normalized composite feature vector constitutes a multidimensional feature vector.
[0077] In step S104, the multidimensional feature vector constructed in step S103 is input into a pre-trained random forest early warning model to obtain the initial early warning level probability distribution corresponding to the target highway segment.
[0078] In some embodiments of the present invention, the pre-training steps of the random forest early warning model include steps S1041-S1046: Step S1041: Obtain a training sample set. Each sample corresponds to a historical early warning analysis period and contains a multi-dimensional feature vector obtained based on the standardized fusion data of the historical early warning analysis period. Its construction method is the same as that in step S103.
[0079] Step S1042: For each sample, add a warning level label to the sample based on the severity, frequency and number of associated risk points of the disaster event that occurred on the target highway section during the corresponding historical warning analysis period.
[0080] The warning levels include at least low risk, medium risk, and high risk, and their generation rules are based on the severity, frequency, and number of risk points of the disaster event.
[0081] Step S1043: Input the multidimensional feature vector into the random forest model for training, and output the predicted probability distribution of the warning level corresponding to the sample; Step S1044: Based on the deviation between the predicted probability distribution of warning levels and the warning level labels, a weighted cross-entropy loss function is constructed to guide the optimization of the random forest warning model. The weights assigned to different warning levels in the weighted cross-entropy loss function are inversely proportional to the frequency of each warning level in the training sample set. This design addresses the class imbalance problem in training data caused by the low actual occurrence frequency of high-risk events, allowing the random forest warning model to focus more on scarce but crucial high-risk samples during the learning process.
[0082] Step S1045: Evaluate the training results of the random forest model based on the weighted cross-entropy loss function, and optimize the training parameters of the random forest model. The model training is monitored using a validation set. Training is stopped when the value of the weighted cross-entropy loss function on the validation set converges to a stable interval, or when the model's macro F1 score on the validation set reaches its optimum.
[0083] Step S1046: Save the trained model parameters to obtain the pre-trained random forest early warning model.
[0084] In step S105, the initial warning level of the target highway segment is determined based on the initial warning level probability distribution obtained in step S104 and the preset probability threshold rules.
[0085] Specifically, the preset probability threshold rule defines the decision boundary for mapping continuous probability values to discrete warning levels.
[0086] In some embodiments, the rule is pre-configured in the form of a list of thresholds. For example, the rule might be set as follows: if the probability value of a "high-risk" level is not less than 0.6, it is determined to be high-risk; if the probability value of a "high-risk" level is less than 0.6, but the probability value of a "medium-risk" level is not less than 0.4, it is determined to be medium-risk; otherwise, it is determined to be low-risk. When applying this rule, each probability value in the initial warning level probability distribution is compared with the corresponding threshold set in the rule to output a definite initial warning level.
[0087] The determination of probability threshold rules is based on the trade-off between false positive rate and false negative rate according to business needs.
[0088] In some embodiments, the optimal threshold operating point for a specific business scenario is selected by analyzing the model’s performance on a historical validation set, such as by plotting the receiver operating characteristic (ROC) curve, and then solidified into a preset rule.
[0089] In step S106, the initial warning level determined in step S105 is corrected based on the feature sub-vector of road segment risk attributes in order to generate and output the final warning result.
[0090] Specifically, the first step is to determine whether the target highway segment meets the preset high-risk attribute conditions. This determination is made by examining the feature values representing the inherent attributes of the road segment in the feature sub-vector of the road segment's risk attributes.
[0091] In some embodiments, high-risk attribute conditions include at least one of the following: seismic fortification level below standard, insufficient flood control standard, real-time or predicted traffic flow exceeding saturation, or belonging to a historically high-frequency road segment. If the value of the corresponding feature in the risk attribute feature sub-vector of a road segment satisfies any of the above conditions, the road segment is determined to meet the high-risk attribute conditions.
[0092] When a target road segment is determined to meet the high-risk attribute conditions, a level upgrade rule is triggered. This rule upgrades the initial warning level determined in step S105 by at least one level. For example, if the initial warning level is "medium risk," it is revised to "high risk"; if the initial warning level is already "high risk," it can be maintained or upgraded to a higher "extremely high risk" level according to the rule. If the target road segment does not meet any preset high-risk attribute conditions, its initial warning level will be directly used as the final warning level without modification.
[0093] Finally, the revised or unrevised warning levels are combined with the target road segment identification, the warning analysis period, and key disaster-causing factor information to generate the final warning result in a formatted format, and then output.
[0094] In some embodiments, the output includes writing the warning results to a database and pushing them to a visualization interface for display.
[0095] On the other hand, the present invention also provides a highway meteorological data processing and early warning system, the system comprising: The multi-source data access terminal is used to obtain multi-source heterogeneous raw data of the target highway section from external meteorological departments, traffic management departments and local storage. The data fusion processing server communicates with the multi-source data access terminal and is used to perform standardized fusion processing on multi-source heterogeneous raw data to generate standardized fusion data corresponding to the target highway segment. The early warning analysis server communicates with the data fusion processing server to perform early warning analysis calculations based on standardized fused data and generate early warning analysis results. The early warning service publishing terminal communicates with the early warning analysis server to receive early warning analysis results and output the final early warning result; Among them, the data fusion processing server, the early warning analysis server, and the early warning service publishing terminal work together and are configured to execute the steps of any of the above methods.
[0096] The present invention will now be described with reference to a specific embodiment: This embodiment will explain the overall design principles, layered architecture, and collaborative workflow of the internal modules of each layer of the highway meteorological data processing and early warning method from the perspective of system implementation, so as to fully present the specific implementation form of the method in engineering practice.
[0097] 1. Basic Principles This design addresses the core issues in existing technologies by constructing a three-layer collaborative architecture: unified data integration, scenario adaptation in the processing layer, and visual display in the application layer.
[0098] To address the issue of "low data fusion efficiency," a multi-source data standardization processing engine was constructed. Through automatic format conversion, spatiotemporal alignment, and intelligent cleaning, scattered meteorological, road network, and disaster data were transformed into usable data in a unified format, eliminating data format inconsistencies.
[0099] To address the issue of "poor adaptability of early warning models," a two-dimensional early warning model based on "meteorological elements + road segment attributes" is established by incorporating road segment risk characteristics and historical disaster characteristics into the random forest ensemble learning algorithm. Furthermore, the early warning logic of "universal calculation + precise correction" is achieved through differentiated threshold adjustments.
[0100] 2. Detailed technical solution The aforementioned three-layer collaborative architecture consists of a data layer, a processing layer, and an application layer, with each layer having clearly defined functions and operating collaboratively.
[0101] The first layer: The data layer serves as the system's data foundation, responsible for aggregating all core data sources. It is divided into three types of data modules, such as... Figure 2 As shown, the functions and data flow of each module are clearly defined: (1) Meteorological Forecast Data Module: This module includes two sub-modules: "Surface Forecast Element Data" and "Precipitation Data". The "Surface Forecast Element Data" is transmitted via FTP protocol in NetCDF format and covers elements such as air pressure, 2-meter relative humidity, and 10-meter wind direction and speed. The "Precipitation Data" has a spatial resolution of 2.5 kilometers and a temporal resolution of hourly, and is also transmitted via FTP. Both types of data are directly sent to the "Data Fusion Processing Engine" in the processing layer.
[0102] (2) Highway Vector Data Module: This module includes three sub-modules: "Risk Point Data," "Disaster Damage Data," and "Road Network Basic Data." "Risk Point Data" stores fields such as road name, risk level, and latitude and longitude of start and end points. "Disaster Damage Data" records disaster type, occurrence time, damage details, and associated risk point information. "Road Network Basic Data" encompasses highway network map vector data and maintenance unit information. All three types of data are accessed via API and uniformly transmitted to the "Data Fusion Processing Engine."
[0103] (3) Historical meteorological data module: It contains 8 core NetCDF file sub-modules (such as PRATEsfc.nc ground precipitation rate, TMP2m.nc 2-meter temperature, ws10m.nc 10-meter wind speed, etc.), which are accessed through the local reading component and sent to the "data fusion processing engine".
[0104] The second layer, the processing layer, acts as the "central hub" of the system. It is responsible for transforming the raw data input from the data layer into usable information and for completing the calculations of the early warning model, such as... Figure 3 As shown, it contains two core engine modules: (1) Data Fusion Processing Engine: This engine comprises three sub-modules: "Format Conversion Module," "Spatiotemporal Alignment Module," and "Intelligent Cleaning Module." The "Format Conversion Module" uses Python's cfgrib and xarray libraries to convert GRIB2 format data to NetCDF format; the "Spatiotemporal Alignment Module" uses the pyproj library to unify all data into the WGS84 coordinate system, with the time granularity unified to hourly; the "Intelligent Cleaning Module" uses 3... Outliers are removed in principle, and missing values are filled in using linear interpolation. The processed data is then sent to the "early warning model engine".
[0105] (2) Early warning model engine: It includes three sub-modules: "feature extraction module", "random forest training module" and "differentiated early warning module".
[0106] The "feature extraction module" extracts meteorological elements (precipitation, wind speed, etc.), road segment attributes (such as seismic fortification level and traffic flow), and historical disaster frequency from the processed data.
[0107] The "Random Forest Training Module" uses historical data to train the model. The historical data contains two types of core information: first, input feature data, including historical meteorological element data (such as time series of precipitation, wind speed, temperature, humidity, etc.) and road segment attribute data (such as seismic fortification level and traffic flow); second, risk points and disaster data.
[0108] Specifically, when a road section has experienced disaster or has risk points under specific weather conditions, the data for that period is labeled as "positive sample". The model's label is an ordered multi-category variable, corresponding to the warning level (e.g., low risk, medium risk, high risk). Its generation rules are based on the severity, frequency, and number of risk points of the disaster event. For example, a severe or above disaster event or property loss greater than 5 million yuan is labeled as "high risk"; a relatively serious disaster event or property loss between 1 million and 5 million yuan is labeled as "medium risk"; and a general disaster event or property loss less than 1 million yuan is labeled as "low risk".
[0109] To ensure the forward-looking nature of the warnings, the labels correspond to meteorological and road condition data prior to the event, for example, data from 48 hours before the event is used as training samples.
[0110] During model training, the weighted cross-entropy loss function is used to address the class imbalance problem (such as the scarcity of high-risk samples). The calculation formula is as follows: ; in, This indicates the total number of samples in a training batch or the entire training set. Indicates the first The input feature vector of each sample; Indicates the first The true warning level label for each sample; Represents the loss function; This represents the category weight, which is inversely proportional to the number of samples. This represents the predicted probability.
[0111] Hyperparameter optimization is achieved through cross-validation and grid search, as shown in steps 1 to 3: Step 1: Divide the training set and validation set. For example, the ratio of the training set to the validation set can be set to 0.7:0.3. Set the parameters for the search space, such as the number of trees: 50~500; maximum depth: 5~20.
[0112] Step 2: Use grid search or Bayesian optimization to optimize parameters using the macro F1 score of the validation set as the evaluation metric. Step 3: Select the parameter combination with the highest F1 score and test the test set to confirm the generalization ability to obtain the pre-trained random forest early warning model.
[0113] The "differentiated early warning module" then logically corrects the early warning level based on whether the road segment meets the high-risk attribute conditions (such as insufficient seismic fortification level or traffic flow oversaturation) after model inference. For example, it raises the early warning level of the road segment that meets the preset high-risk attribute conditions by 1 level.
[0114] The third layer: The application layer acts as the system's "terminal," responsible for providing services to users, such as... Figure 4 As shown, it contains one core service module: (1) PC-side system modules: including three sub-modules: “Data Visualization Module”, “Report Generation Module”, and “System Management Module”. The “Data Visualization Module” uses ECharts to draw meteorological element trend charts (such as temperature and precipitation changes) and early warning maps (using red / orange / yellow colors to mark the road section early warning level); the “Report Generation Module” automatically generates structured early warning and response reports (including road condition information, disaster risk, and defense guidelines) according to preset templates; the “System Management Module” realizes functions such as user management, permission allocation, and password modification.
[0115] Through the collaborative operation of the above three-layer architecture, this design achieves full-process automation and intelligence from multi-source data access, fusion processing, intelligent analysis to service deployment. The initial output of the random forest early warning model is the probability distribution for each risk category (e.g., [low risk: 0.2, medium risk: 0.5, high risk: 0.3]). When generating an early warning, the initial warning level is first determined according to preset probability threshold rules (e.g., a high-risk probability ≥ 0.6 is considered high-risk). Then, the "differentiated early warning module," as an independent post-processing logic correction layer, performs rule-based adjustments to the initial level based on high-risk attribute conditions (e.g., raising the level by 1 for road sections with oversaturated traffic flow). Finally, it generates and outputs early warning results with high interpretability and business relevance, effectively improving the accuracy, timeliness, and system collaboration capabilities of highway meteorological early warnings.
[0116] The present invention will now be described with reference to another specific embodiment: This embodiment describes a standardized processing method from the perspective of data fusion processing, which unifies highway linear spatial data, regular meteorological grid data, disaster event data, and risk object attribute data onto a standard unit at the road segment level. This method can be used as a specific implementation of the data fusion stage in the aforementioned early warning method, or it can be used independently to construct the standard sample set required for model training.
[0117] I. Highway Vector Discretization and Generation of Standard Road Segment Units In this embodiment, the original highway vector road network spatial data is first discretized. Specifically, the latitude and longitude coordinates of all ordered vertices of the vector line segment are extracted, and the actual distances between vertices on the Earth's surface between them are calculated segment by segment based on the Havesing spherical distance formula. These distances are then summed to obtain the total spherical length of the entire road segment. The Havesing formula is as follows: ; ; ; ; in, Indicates the latitude of the starting point of the line segment; Indicates the latitude of the endpoint of the line segment; Indicates the difference in latitude between the starting and ending points; Indicates the longitude of the starting point of the line segment; Indicates the longitude of the endpoint of the line segment; Indicates the difference in longitude between the starting and ending points; Indicates intermediate variables; This represents the central angle between two points on a sphere. This represents the Earth's average radius, with a value of 6,371,000 meters. This represents the spherical distance of a single line segment. The total length of the entire road segment. The length of each sub-segment The sum of .
[0118] According to the preset fixed sampling interval The formula for calculating the number of equally spaced sampling points is: ; in, This indicates the total number of equidistant sampling points along the road segment; This indicates the total spherical length of the road segment, expressed in meters. This indicates the preset fixed sampling interval, in meters; This represents the floor function, used to ensure that sampling points are uniformly covered along the entire road segment without missing any endpoints.
[0119] In the normalized parameter domain of road segments The sampling parameters are generated proportionally, and the latitude and longitude coordinates of the sampling points are generated by spherical linear interpolation.
[0120] Each sampling point generates a globally unique identifier, with the following encoding rule:<road_id><line_index><point_index> The coding meaning is: the road segment number, the sub-segment number within the road segment, and the sampling point number within the sub-segment, which realizes the unique identity definition of the sampling points in the entire road network and eliminates the problems of duplication and misalignment.
[0121] The spatial interval between adjacent sampling points is defined as the minimum standard unit of road segment.
[0122] The formula for calculating the normalized sampling parameters is: ; in, Indicates the first The normalized position parameters of each sampling point on the road segment are distributed at equal intervals along the parameter domain of the line segment, and finally the latitude and longitude coordinates of the sampling points are generated that perfectly match the preset spherical spacing.
[0123] II. Meteorological Grid Spatial Matching and Mapping Reuse After generating sampling points, the latitude and longitude dimension array of the regular grid in the meteorological NC file is extracted to form a set of coordinates of the two-dimensional regular grid nodes that cover the entire area. A KD-Tree high-dimensional spatial index is then constructed based on the coordinates of all grid nodes.
[0124] Input the coordinates of each sampling point obtained above through equidistant scattered sampling and spherical linear interpolation into the spatial index. Let the first... The coordinates of the sampling points are ,in The longitude of the sampling point. This refers to the latitude of the sampling point. Execute. The nearest neighbor query finds the meteorological grid node with the smallest spherical distance to the sampling point. The determination formula is as follows: ; in, Indicates the first One sampling point; This represents the set of coordinates of all nodes in the meteorological grid. Indicates the position in the meteorological grid. line, number Column grid nodes; This represents the Havelsing spherical distance between the sampling point and the grid node, in meters; The minimum value operator; subscript Represents a set The search is performed by traversing all the grid nodes.
[0125] Based on the matching results above, record the two-dimensional row and column index of the grid node corresponding to each sampling point. It also generates a fixed mapping dictionary to permanently store the correspondence between each highway sampling point and the row and column indices of the meteorological grid.
[0126] To address changes in meteorological data source formats, a grid consistency verification mechanism is established. Upon initial establishment of the mapping relationship, core metadata of the meteorological grid is extracted, including grid dimensions, latitude and longitude spatial boundaries, and a fixed-length header feature sequence for the latitude and longitude arrays. After converting the latitude and longitude header feature sequence into a standardized byte stream, a unique fingerprint for each grid is generated using the SHA-256 cryptographic hash algorithm. The calculation formula is as follows: ; in, This represents the head feature sequence of the longitude array; Represents the head feature sequence of the dimensional array; This represents a unique hash fingerprint for the grid. The generated fingerprint is then associated with and stored as a mapping relationship.
[0127] When processing new NC meteorological files in each subsequent batch, the grid metadata and hash features of the current file are first extracted and compared with the pre-stored fingerprints in all fields: if they match, the mapping relationship is directly reused; if they do not match, the data processing is terminated and a remapping is triggered.
[0128] III. Construction of Meteorological Time Series and Event Attribution at the Road Section Level After mapping the sampling points to the grid, the meteorological sequence of the grid unit over continuous time is assigned to the corresponding road segment sampling points, and then aggregated into a road segment-level meteorological time series. For a certain meteorological element value of a certain road segment unit at a certain time, any of the following aggregation methods can be used: 1) Interval mean; 2) Interval maximum value; 3) Interval minimum value; 4) Average value of representative points at both ends; 5) Length-weighted average; 6) Joint expression of extreme values and mean.
[0129] For short road sections, representative point values can be used first; for longer road sections, a combination of endpoint values and interval averages can be used first. Through this step, the original grid-level meteorological process data is uniformly upgraded to road section-level meteorological process data.
[0130] The location of a disaster event is resolved into spatial points, and based on spatial distance or linear projection, the event is assigned to the nearest standard road segment unit. The static attributes of risk objects also employ a consistent spatial assignment method, creating a unified association between disaster tags, risk objects, and road segment units. The specific assignment rules are as follows: 1) Prioritize selecting road segment units with the closest spatial distance; 2) When the spatial distance is less than a preset threshold, the attribution is completed directly; 3) When there are multiple candidate road segments, select the road segment with the smallest projection distance or the highest directional consistency; 4) When the event point falls in a complex intersection area, the determination can be made by combining the adjacent road grade, continuous direction or historical attribution results.
[0131] This step can transform the original point-level disaster records into road segment-level event tags.
[0132] The present invention will now be described with reference to yet another specific embodiment: This embodiment is a further development based on the above embodiments, mainly illustrating how to extract multi-window meteorological features around the event time, and on this basis, form standardized fusion samples for model training, risk assessment, or statistical analysis. I. Multi-window meteorological feature extraction oriented towards the event time In this embodiment, statistical features describing the pre-disaster environmental state and evolution process are extracted from road segment-level meteorological time series data, focusing on the time of the disaster event. The extracted features include at least: the most recent observation at the time of the event, the most recent observation before the event, the cumulative value, the maximum value, and the average value within several time windows before the event.
[0133] For cumulative factors such as precipitation, snowfall, and freezing, multiple time windows are set, including 24 hours, 48 hours, 72 hours, 96 hours, 120 hours, 144 hours, and 168 hours. For state-related factors such as temperature, humidity, air pressure, and wind, shorter base windows are set, and expansion is allowed as needed for business operations.
[0134] II. Standardized Fusion Sample Generation After obtaining the above meteorological statistical characteristics, the following information is integrated into a single road segment-level sample: basic spatial information of the road segment, static attributes of risk objects, disaster event labels, road segment-level meteorological statistical characteristics, and necessary derived structural attributes. Each sample corresponds to a combination of a road segment unit and an event time.
[0135] In some embodiments, if it is necessary to construct disaster-free samples, i.e. negative samples, the event time can be reset while keeping the spatial location and static attributes unchanged, so that it does not overlap with the real disaster event in the time dimension, thereby generating negative samples for model training.
[0136] III. Quality Control and Standardized Data Objects To improve the consistency and usability of the samples, this embodiment also sets the following quality control strategies: validating the spatial coordinates and removing points that are out of range or cannot be resolved; uniformly formatting the time field; verifying the consistency of the meteorological grid system; deduplicating or resolving conflicts for duplicate events, duplicate samples, and duplicate road segment attributions; and marking, interpolating, or retaining missing values.
[0137] This embodiment recommends forming at least the following four types of standardized data objects: 1) Basic objects of road segment: including road segment sign, road sign, start and end points, center point, road segment length, and adjacency relationship.
[0138] 2) Road segment meteorological time series objects: including road segment identification, time, meteorological element category, element value, data source identification, and grid consistency identification.
[0139] 3) Disaster event objects: including event identifier, road segment identifier, event location, event time, disaster type, and disaster level.
[0140] 4) Road segment-level fusion sample objects: including event identifiers, road segment identifiers, static attribute sets of risk objects, meteorological statistical feature sets, disaster damage type labels, and disaster damage level labels.
[0141] Through the processing described in the above embodiments, the original multi-source heterogeneous data can be transformed into a road segment-level standard sample set with unified structure, consistent caliber, and repeatable generation, which can be used for subsequent disaster identification, risk assessment, statistical analysis, and early warning model training.
[0142] IV. Sample Data To facilitate understanding of the specific form of each data object in this embodiment, typical examples are provided below.
[0143] (I) Example of basic road segment objects (II) Examples of raw meteorological data (III) Examples of meteorological time series objects for road sections (iv) Examples of objects affected by disasters (V) Examples of road segment-level fusion samples In summary, this invention provides a highway meteorological data processing method and early warning system. The method includes: acquiring multi-source heterogeneous raw data of a target highway segment, including meteorological grid data, traffic and highway vector data, and historical meteorological and disaster case data; performing standardized fusion processing on the raw data, including format conversion, spatiotemporal alignment, and intelligent cleaning, to generate standardized fused data; constructing a multi-dimensional feature vector based on this fused data, integrating meteorological features, road segment risk attribute features, and historical disaster features, and inputting it into a pre-trained random forest early warning model to obtain an initial early warning level probability distribution; determining the initial early warning level according to a preset probability threshold rule; and, based on the road segment risk attribute features, upgrading the initial early warning level of road segments that meet preset high-risk attribute conditions, generating and outputting the final early warning result. This invention enables deep fusion and collaborative analysis of meteorological, road network, and historical disaster data, improving the accuracy, timeliness, and scenario adaptability of road segment-level meteorological disaster early warnings.
[0144] Furthermore, when constructing the pre-trained random forest early warning model, a weighted cross-entropy loss function is used for training, where the weights assigned to different early warning levels are inversely proportional to their frequency of occurrence in the training samples. This method addresses the problem of scarce high-risk samples and uneven distribution of different categories in highway early warning scenarios. It enables the model to pay more attention to minority class samples during training, thereby effectively improving the ability to identify high-risk meteorological disasters, reducing the false negative rate, and enhancing the model's practicality and reliability.
[0145] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.
[0146] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0147] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0148] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A highway weather data processing and early warning method, characterized in that, The method includes the following steps: Acquire multi-source heterogeneous raw data of the target highway segment. The multi-source heterogeneous raw data includes at least meteorological grid data from external meteorological departments, traffic and highway vector data containing road segment risk attribute characteristics from traffic management departments, and historical meteorological and disaster case data stored locally and associated with historical early warning analysis periods. The multi-source heterogeneous raw data is standardized and fused to generate standardized fused data for the target highway segment. The standardized fusion process includes: spatially matching the meteorological grid data with the target highway segment to generate segment-level time-series meteorological data; performing geocoding-based spatial location matching and time window aggregation on the traffic and highway vector data with the historical meteorological and disaster case data to form segment-enhanced attribute data; and performing quality verification and missing value imputation on the segment-level time-series meteorological data and the segment-enhanced attribute data. Based on the standardized fusion data, for the early warning analysis period, a multi-dimensional feature vector associated with the target highway segment is constructed, including: selecting key meteorological elements within the early warning analysis period from the segment-level time-series meteorological data and calculating their statistics to construct a meteorological feature sub-vector; extracting attribute features representing the inherent risk level of the segment from the segment enhanced attribute data to construct a segment risk attribute feature sub-vector; extracting regular features representing historical disaster patterns from the segment enhanced attribute data to construct a historical disaster feature sub-vector; and concatenating and normalizing the meteorological feature sub-vector, the segment risk attribute feature sub-vector, and the historical disaster feature sub-vector to form the multi-dimensional feature vector. The multidimensional feature vector is input into a pre-trained random forest early warning model to obtain the initial early warning level probability distribution corresponding to the target highway segment; The initial warning level of the target highway segment is determined based on the initial warning level probability distribution and the preset probability threshold rules. Based on the feature subvector of the road segment risk attribute, if the target highway segment is determined to meet the preset high-risk attribute conditions, the initial warning level will be increased by at least one level to generate and output the final warning result.
2. The road weather data processing and early warning method according to claim 1, characterized in that, The meteorological grid data includes surface forecast element data; the surface forecast element data includes at least one or more of precipitation, snowfall, air pressure, humidity, wind direction, and wind speed; Its data storage and transmission formats include Network Common Data Format (NetCDF) or Grid Binary Format (GRIB). The traffic and highway vector data includes highway risk point data, historical disaster damage data, and basic road network data; The historical meteorological and disaster case data includes historical meteorological element files stored in a common network data format.
3. The road weather data processing and warning method according to claim 1, characterized in that, Spatially match the meteorological grid data with the target highway segment to generate segment-level time-series meteorological data, including: Using nearest neighbor matching, bilinear interpolation, or inverse distance weighted interpolation algorithms, the corresponding meteorological element values are extracted and calculated from the meteorological grid data based on the geographical coordinates of the center point or representative point of the target highway segment.
4. The road weather data processing and warning method according to claim 1, characterized in that, Spatially match the meteorological grid data with the target highway segment to generate segment-level time-series meteorological data, including: The target road segment in the traffic and highway vector data is sampled at equal intervals according to a preset sampling interval to generate a sequence of multiple sampling points continuously distributed along the target road segment; Based on the spatial matching relationship between multiple sampling points and the corresponding grid cells of the meteorological grid data, the meteorological element values of each sampling point at the same time or within the same time window are obtained. According to a preset aggregation strategy, the meteorological element values corresponding to each sampling point within the same target highway segment are aggregated to generate the segment-level time-series meteorological data; the aggregation strategy includes any one of interval mean, interval extreme value or length-weighted average.
5. The road weather data processing and early warning method according to claim 4, characterized in that, The spatial matching relationship is established and reused in the following ways: A spatial index is constructed based on the latitude and longitude coordinates of the meteorological grid data. Nearest neighbor matching is performed on each sampling point to determine the meteorological grid index corresponding to each sampling point. Store the mapping relationship between sampling points and meteorological grid index; When processing subsequent meteorological grid data, consistency checks are performed based on the dimensional dimensions, latitude and longitude boundary ranges, and representative coordinate sequence summary information of the meteorological grid. The mapping relationship is reused when the verification is consistent, and the mapping relationship is re-established when the verification is inconsistent.
6. The road weather data processing and warning method according to claim 1, characterized in that, The step of performing geocoding-based spatial location matching and time window aggregation between the traffic and highway vector data and the historical meteorological and disaster case data to form road segment enhanced attribute data includes: Based on the nearest neighbor spatial matching algorithm, the geographical location of the historical disaster case data is associated with the nearest target highway segment defined in the traffic highway vector data; and based on the preset warning time window, the associated historical disaster case data is collected into the corresponding warning time window according to the relative relationship between its occurrence time and the warning analysis time.
7. The road weather data processing and warning method according to claim 1, characterized in that, The process of performing quality verification and missing value imputation on the road segment-level time-series meteorological data and road segment enhanced attribute data includes: Outliers in the road segment-level time-series meteorological data are identified and removed based on the three Sigma principle; and a linear interpolation algorithm is used to fill in the data points that have been removed or are missing.
8. The road weather data processing and warning method according to claim 1, characterized in that, The steps for pre-training the random forest early warning model include: Obtain a training sample set, where each sample corresponds to a historical early warning analysis period and contains the multidimensional feature vector obtained based on the standardized fusion data of the historical early warning analysis period; For each sample, an early warning level label is added to the sample based on the severity, frequency, and number of associated risk points of the disaster event that occurred on the target highway section during the corresponding historical early warning analysis period. The multidimensional feature vector is input into the random forest model for training, and the predicted probability distribution of the warning level corresponding to the sample is output. Based on the deviation between the predicted probability distribution of the warning level and the warning level label, a weighted cross-entropy loss function is constructed; wherein, the weights assigned to different warning levels in the weighted cross-entropy loss function are inversely proportional to the frequency of occurrence of each warning level in the training sample set; the training results of the random forest model are evaluated based on the weighted cross-entropy loss function, and the training parameters of the random forest model are optimized. When the weighted cross-entropy loss function converges to a stable interval on the validation set, or when the macro F1 score of the model on the validation set reaches its optimum, training is stopped, the final model parameters are fixed, and the pre-trained random forest early warning model is obtained.
9. A highway weather data processing and early warning system, characterized in that, The system includes: The multi-source data access terminal is used to obtain multi-source heterogeneous raw data of the target highway section from external meteorological departments, traffic management departments and local storage. A data fusion processing server is communicatively connected to the multi-source data access terminal and is used to perform standardized fusion processing on the multi-source heterogeneous raw data to generate standardized fusion data corresponding to the target highway segment. An early warning analysis server, which is communicatively connected to the data fusion processing server, is used to perform early warning analysis calculations based on the standardized fusion data and generate early warning analysis results. The early warning service publishing terminal communicates with the early warning analysis server to receive the early warning analysis results and output the final early warning result; The data fusion processing server, the early warning analysis server, and the early warning service publishing terminal operate in concert and are configured to perform the steps of the method as described in any one of claims 1 to 8.