Safety risk assessment system based on railway yard information
By constructing a railway safety knowledge graph and a graph-enhanced time series prediction model, the problem of insufficient training data for newly built and low-risk stations has been solved, efficient and accurate risk prediction and early warning have been achieved, and the efficiency of the safety risk assessment system has been improved.
Patent Information
- Application Number
- CN202510930382.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-07
AI Technical Summary
When constructing risk prediction models for newly built and low-risk stations, existing technologies suffer from insufficient training data, resulting in time-consuming and inefficient deployment of safety risk assessment systems.
By constructing a railway safety knowledge graph, adopting a cross-station data enhancement strategy under the transfer learning framework, using multi-source heterogeneous data for unified analysis, and combining the graph-enhanced time series prediction model and the graph similarity estimation model, a risk prediction model for the target station is trained.
It improves the deployment efficiency and accuracy of risk prediction models for newly built stations, realizes unified analysis of multi-source heterogeneous data, and improves the application efficiency and accuracy of safety risk assessment systems.
Smart Images

Figure CN120822830A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of railway transportation safety technology, and specifically is a safety risk assessment system based on railway station information. Background Art
[0002] Railway yards are a core component of the railway transportation system, typically referring to the comprehensive railway facilities used for train arrival and departure, marshaling, maintenance, cargo loading and unloading, and passenger boarding and alighting. They can be categorized by purpose as passenger, freight, and mixed-use. Safety risk assessment is a systematic analytical approach designed to identify, quantify, and manage potential risks to prevent accidents and improve system safety. Safety risk assessment is particularly important in railway yard scenarios, as it involves complex, multi-dimensional factors such as train operations, equipment status, and personnel operations.
[0003] In existing railway stations, data such as equipment status, operating procedures, and environmental monitoring are scattered across different systems, making it difficult to achieve unified analysis of multi-source heterogeneous data, resulting in low risk assessment efficiency for the entire station. At the same time, the construction of station risk prediction models in newly built and low-risk stations often suffers from a lack of training data, resulting in a long time required to deploy a safety risk assessment system, and the accuracy of the safety risk assessment system is poor. Therefore, the safety risk assessment system for railway station information still needs further improvement. Summary of the Invention
[0004] The present application aims to solve at least one of the technical problems existing in the prior art; to this end, the present application proposes a safety risk assessment system based on railway station information, which is used to solve the technical problem that the prior art constructs risk prediction models for newly built stations and low-risk stations, and there is often a small amount of training data, which makes the deployed safety risk assessment system time-consuming and inefficient.
[0005] To achieve the above-mentioned purpose, the first aspect of the present application provides a safety risk assessment system based on railway station information, comprising: a data acquisition module, a data analysis module, an early warning module and a database;
[0006] The data acquisition module acquires station data through data acquisition equipment; the station data includes station ID, station purpose, equipment data, environmental data and operation data;
[0007] The data analysis module: constructs a railway safety knowledge graph based on station data; generates a risk prediction model based on the railway safety knowledge graph; generates predicted risk events and predicted risk levels based on the risk prediction model; and generates an alarm signal based on the predicted risk level;
[0008] The early warning module: makes prompts according to the alarm signal and contacts the management personnel;
[0009] The database is used to store data of each module and store historical data required for training the model.
[0010] Through the above steps, this application simultaneously constructs a railway safety knowledge graph and a risk prediction model, which not only improves the accuracy of risk identification, but also enables advance warning and proactive prevention and control of risks; by uniformly time-coding data of different frequencies, unified analysis of multi-source heterogeneous data is achieved; and a cross-station data enhancement strategy under the transfer learning framework is adopted to explore the risk transmission patterns between similar sites, so that newly built sites can still build high-precision prediction models when the amount of training data is insufficient, thereby improving the scenario adaptability of the safety assessment system and the efficiency of engineering implementation.
[0011] Furthermore, the railway safety knowledge graph is constructed based on the station data, including:
[0012] Acquire station data; the station data includes station ID, equipment data, environmental data and operation data;
[0013] Generate specification data based on equipment data, environmental data, and operation data. The specification data refers to standard data used to construct a railway safety knowledge graph. The specification data includes several entities, several attributes, and several relationships.
[0014] Use a graph database to build a safety knowledge graph, and embed spatiotemporal data storage and rule-driven systems into the safety knowledge graph. The spatiotemporal data storage is used to store the spatiotemporal data corresponding to the site data. The rule-driven systems are constructed by experts based on the site data and its corresponding risk level.
[0015] Input the specification data into the safety knowledge graph to obtain the railway safety knowledge graph.
[0016] Furthermore, generating specification data based on equipment data, environment data, and operation data includes:
[0017] Acquire equipment data, environmental data, and operation data; the equipment data includes equipment ID, equipment status, and equipment parameters; the environmental data includes meteorological data and geological data; the operation data includes operation ID, operation instructions, maintenance records, and surveillance videos;
[0018] Data classification is performed on the device data, environment data, and operation data to obtain text data, view data, time series data, and structured data; the structured data refers to data with a fixed format and clear fields; the view data refers to video data and image data;
[0019] Standardized data is obtained by using a hierarchical fusion strategy for text data, view data, time series data and structured data; the hierarchical fusion strategy operates in a multimodal fusion manner.
[0020] Through the above steps, the equipment data, environmental data and operation data in the station can be unified into multi-source heterogeneous data, providing strong data support for the subsequent risk assessment of the railway safety knowledge graph and the risk prediction model for future risk predictions.
[0021] Furthermore, the layered fusion strategy is operated in a multimodal fusion manner, including the following steps:
[0022] Step 1: Obtain text data, view data, time series data, and structured data;
[0023] Step 2: Extract features from text data, view data, time series data, and structured data to obtain a number of semantic vectors and spatial vectors; the semantic vectors and spatial vectors include a number of entities, attributes, and relationships;
[0024] Step 3: Add a unified time code to several semantic vectors and align data of different frequencies through sliding window aggregation to obtain aligned semantic vectors;
[0025] Step 4: Obtain normalized data by concatenating the aligned semantic vector and the spatial vector.
[0026] Furthermore, generating a risk prediction model based on the railway safety knowledge graph includes:
[0027] Obtain railway safety knowledge graphs and matching priorities PYi corresponding to several station IDs and railway safety knowledge graphs corresponding to target station IDs; the matching priorities are calculated using historical station data corresponding to the several station IDs; the target station ID is the railway station ID for which a risk prediction model needs to be constructed;
[0028] Sort several matching priorities from largest to smallest to obtain a matching sequence PL;
[0029] Select the station IDs and railway safety knowledge graphs corresponding to the top N matching priorities in PL, and name the railway safety knowledge graphs as source railway safety knowledge graphs; where N is an integer, N>1;
[0030] Extract several common subgraph data from the railway safety knowledge graph corresponding to the target station ID and the source railway safety knowledge graph;
[0031] Merging a plurality of common sub-graph data and the station data corresponding to the target station ID to obtain target comprehensive data; the target comprehensive data includes a plurality of sub-graph data;
[0032] The risk prediction model is obtained by training the graph-enhanced time series prediction model on the target comprehensive data.
[0033] Furthermore, the matching priority is calculated using historical station data corresponding to a number of station IDs, including:
[0034] Obtain historical station data corresponding to a number of station IDs; the historical station data includes historical risk levels, historical risk counts, and occurrence times;
[0035] Calculate the comprehensive risk scores corresponding to several station IDs based on historical risk levels, historical risk times and occurrence time;
[0036] Obtain several source railway safety knowledge graphs and railway safety knowledge graphs corresponding to the target station ID;
[0037] Input several source railway safety knowledge graphs and the railway safety knowledge graphs corresponding to the target station ID into the graph similarity prediction model to obtain several similarity levels XD i ; The graph similarity prediction model is constructed through an artificial intelligence model;
[0038] By formula PY i =g×tan -1 (α1×ZFP i +α2×XD i ) Calculate the matching priority PY i ; Where i represents the number of station IDs, g is the proportional coefficient, g∈(0,π / 2); α1 and α2 are weight coefficients, α1 and α2∈(0,1); and the sum of α1 and α2 is 1; ZFP i It is expressed as the comprehensive risk score corresponding to the i-th station ID.
[0039] Furthermore, the comprehensive risk scores corresponding to several station IDs are calculated based on the historical risk level, the number of historical risks, and the time of occurrence, including:
[0040] Obtain historical risk level FD, historical risk times and occurrence time Δt;
[0041] By formula Calculation of comprehensive risk score ZFP i ; where j represents the number of historical risk events, max() represents the maximum value operation, max(j) represents the number of historical risk events, and β represents the time attenuation coefficient, β∈(0,1); Δt i,j It is expressed as the time from the occurrence of the jth risk at the i-th station ID to the present; DT is expressed as unit time; the sensitivity to the total number of risks is retained by using the log() function, but extreme values are suppressed by the logarithmic function.
[0042] Furthermore, the graph similarity prediction model is constructed through an artificial intelligence model, including:
[0043] Obtaining several historical source railway safety knowledge graphs and railway safety knowledge graphs corresponding to historical target station IDs and their corresponding historical similarity levels;
[0044] The railway safety knowledge graphs corresponding to several historical source railway safety knowledge graphs and historical target station IDs and their corresponding historical similarity levels are divided into training data, verification data, and test data; the training data, verification data, and test data are preprocessed to obtain a training set, a verification set, and a test set;
[0045] Select an artificial intelligence model as the base model;
[0046] Train the basic model using the training set, and adjust the learning rate and hyperparameters on the validation set to obtain the pre-trained model;
[0047] By verifying the pre-trained model on the test set, we finally obtained a railway safety knowledge graph whose input is several source railway safety knowledge graphs and corresponding to the target station ID, and the output is a graph similarity prediction model with several similarity levels.
[0048] Furthermore, the risk prediction model is obtained by training the graph-enhanced time series prediction model on the target comprehensive data, including:
[0049] Acquire target comprehensive data; the target comprehensive data includes a plurality of sub-graph data;
[0050] Extract several historical subgraph data and their corresponding historical time series data, historical risk levels, and risk events;
[0051] Integrate several historical subgraph data and their corresponding several historical time series data into several historical prediction data;
[0052] Divide a number of historical forecast data and their corresponding historical risk levels and risk events into training data, verification data, and test data; perform data preprocessing on the training data, verification data, and test data to obtain a training set, a verification set, and a test set;
[0053] Select the graph-enhanced time series prediction model as the base model;
[0054] Train the basic model using the training set, and adjust the learning rate and hyperparameters on the validation set to obtain the pre-trained model;
[0055] By verifying the pre-trained model on the test set, we finally obtain a risk prediction model whose input is prediction data and output is predicted risk level and predicted risk event.
[0056] Furthermore, generating predicted risk events and predicted risk levels based on the risk prediction model includes:
[0057] Real-time acquisition of time series data and risk prediction models corresponding to the target station ID and a railway safety knowledge graph; the railway safety knowledge graph includes several subgraph data;
[0058] Extract several historical time series data and their corresponding historical subgraph data and integrate them with the current time series data and subgraph data to obtain prediction data;
[0059] The predicted data is input into the risk prediction model to obtain the predicted risk level and predicted risk events.
[0060] Furthermore, generating an alarm signal according to the predicted risk level includes:
[0061] Obtaining predicted risk levels and level intervals; the level intervals include low risk intervals, medium risk intervals, and high risk intervals; the level intervals are set by experts based on historical risk levels and corresponding historical risk events;
[0062] When the predicted risk level falls within the low-risk range, a warning signal for a low-risk event is generated;
[0063] When the predicted risk level falls within the medium risk range, a medium risk event alarm signal is generated;
[0064] When the predicted risk level belongs to the high risk range, an emergency risk alarm signal is generated.
[0065] Through the above steps, this application can respond in a timely manner when a predicted risk level occurs, notify relevant departments to make advance preventive arrangements, reduce safety hazards, and improve the security of the safety risk assessment system.
[0066] Compared with the prior art, the present invention has the following advantages:
[0067] 1. This application constructs a railway safety knowledge graph based on station data; generates a risk prediction model based on the railway safety knowledge graph; generates predicted risk events and predicted risk levels based on the risk prediction model; generates an alarm signal based on the predicted risk level, and constructs a railway safety knowledge graph and a risk prediction model at the same time, which can not only improve the accuracy of prediction when risks occur, but also accurately predict risks and avoid them in advance; by uniformly time-coding data of different frequencies, unified analysis of multi-source heterogeneous data is achieved; and the risk prediction model of the target station is trained using risk data from other source stations, so that the target station can still obtain a risk prediction model with high accuracy when using a small amount of training data, thereby improving the application efficiency of the safety risk assessment system.
[0068] 2. This application obtains the matching priority between other railway stations and the target station, and uses the data of several stations with higher matching priority and similar to the target station to train the risk prediction model of the target station. In the case of less training data for the target station, the deployment efficiency of the target station risk prediction model and the accuracy of the model prediction are improved, thereby improving the application efficiency of the safety risk assessment system.
[0069] 3. This application uses multi-source data to calculate the comprehensive risk score of the source station by considering the historical risk level, historical risk frequency and occurrence time of the source station, thereby improving the representativeness of the comprehensive risk score and providing data support for finding suitable source station data. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0071] Figure 1 This is a schematic diagram of the principle of a safety risk assessment system based on railway station information in this application;
[0072] Figure 2 This is a flow chart of a safety risk assessment method based on railway station information in this application;
[0073] Figure 3 This is the flow chart of the layered fusion strategy of this application. DETAILED DESCRIPTION
[0074] The following will clearly and completely describe the technical solutions of this application in conjunction with the embodiments. Obviously, the embodiments described are only a part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0075] See also Figure 1-Figure 2 , the first embodiment of the present application provides a safety risk assessment system based on railway station information, including: a data acquisition module, a data analysis module, an early warning module and a database;
[0076] Data acquisition module: acquires station data through data acquisition equipment; station data includes station ID, station purpose, equipment data, environmental data and operation data; data acquisition equipment includes several sensors, etc.
[0077] Data Analysis Module: Builds a railway safety knowledge graph based on station data. The railway safety knowledge graph is a graphical representation of the entire station information. Generates a risk prediction model based on the railway safety knowledge graph. The risk prediction model is a model for predicting risks. Generates predicted risk events and predicted risk levels based on the risk prediction model. Generates alarm signals based on the predicted risk levels.
[0078] Early warning module: issues prompts based on alarm signals and contacts management personnel; alarm signals include low-risk event warning signals and emergency risk warning signals;
[0079] The database is used to store data for each module and store historical data required for training the model.
[0080] In this embodiment, the railway safety knowledge graph is constructed based on station data, including:
[0081] Obtain station data; station data includes station ID, equipment data, environmental data and operation data;
[0082] Generate normative data based on equipment data, environmental data, and operational data. Normative data refers to standard data used to construct a railway safety knowledge graph. Normative data includes several entities, several attributes, and several relationships.
[0083] A graph database is used to build a safety knowledge graph, and spatiotemporal data storage and rule-driven development are embedded in the security knowledge graph. The spatiotemporal data storage is used to store the spatiotemporal data corresponding to the site data. Experts construct rule-driven development based on the site data and its corresponding risk level. Neo4j is used as the graph database, and PostGIS extension is used as the spatiotemporal data storage.
[0084] Input the specification data into the safety knowledge graph to obtain the railway safety knowledge graph.
[0085] In this embodiment, the specification data is generated based on the device data, environment data, and operation data, including:
[0086] Acquire equipment data, environmental data, and operation data; equipment data includes equipment ID, equipment status, and equipment parameters; environmental data includes meteorological data and geological data; operation data includes operation ID, operation instructions, maintenance records, and surveillance videos;
[0087] Data classification is performed on device data, environment data, and operation data to obtain text data, view data, time series data, and structured data. Structured data refers to data with a fixed format and clear fields; view data refers to video data and image data.
[0088] Standardized data is obtained by using a layered fusion strategy for text data, view data, time series data, and structured data; the layered fusion strategy operates through multimodal fusion. Since station data has various forms and the collection frequency of each type of data is also different, it is crucial to unify the multimodal data in order to obtain the risk assessment status of the entire station through a single system.
[0089] See also Figure 3 The layered fusion strategy in this embodiment operates in a multimodal fusion manner, including the following steps:
[0090] Step 1: Obtain text data, view data, time series data, and structured data;
[0091] Step 2: Extract features from text data, view data, time series data, and structured data to obtain semantic vectors and spatial vectors. Semantic vectors and spatial vectors include entities, attributes, and relationships. Semantic vectors are extracted from text data using a pre-trained language model, such as the BERT model. For view data, features are extracted using a ResNet network, and spatial location is determined using object detection and SLAM technology.
[0092] Step 3: Add a unified time code to several semantic vectors. The unified time code includes Time2Vec and other methods. Align data of different frequencies through sliding window aggregation to obtain aligned semantic vectors. The sliding window aggregation method can be set to fuse multimodal data within every M-minute window. M can be set based on experience. It can be set to fuse multimodal data within every 5-minute window or every 10-minute window. When aligning semantic vectors at the same timestamp, the features of each modality can be directly spliced together, or the cross-modal attention mechanism can be used to dynamically assign weights to each modal feature for direct splicing.
[0093] Step 4: obtain standard data by splicing the aligned semantic vector and the spatial vector. When splicing the aligned semantic vector and the spatial vector, use a direct splicing method to splice them.
[0094] In this embodiment, the risk prediction model is generated based on the railway safety knowledge graph, including:
[0095] Obtain railway safety knowledge graphs and matching priorities corresponding to several station IDs i The railway safety knowledge graph corresponding to the target station ID; the matching priority is calculated based on the historical station data corresponding to several station IDs; the target station ID refers to the railway station ID for which the risk prediction model needs to be built;
[0096] Sort several matching priorities from largest to smallest to obtain a matching sequence PL;
[0097] Select the station IDs and railway safety knowledge graphs corresponding to the top N matching priorities in PL, and name the top N railway safety knowledge graphs as source railway safety knowledge graphs; where N is an integer, N>1, and the specific value is set based on experience. In this embodiment, N is set to 3;
[0098] Extract several common subgraph data from the railway safety knowledge graph corresponding to the target station ID and the source railway safety knowledge graph. This step mainly searches for common subgraph data in the source railway safety knowledge graph that are related to the entities, attributes, and relationships in the railway safety knowledge graph of the target ID. Using these common subgraph data can help train the risk prediction model for the target station ID and provide it with a large amount of relevant training data.
[0099] Merge several common subgraph data and the station data corresponding to the target station ID to obtain target comprehensive data; the target comprehensive data includes several subgraph data, that is, all the data required to train the risk prediction model for the target station;
[0100] The risk prediction model is obtained by training the graph-enhanced time series prediction model on the target comprehensive data.
[0101] This embodiment obtains the matching priorities between other railway stations and the target station, selects several stations with higher matching priorities, and then applies the data of these stations that are similar to the target station to train the risk prediction model of the target station; when the training data of the target station is limited, this method can significantly improve the deployment efficiency of the target station risk prediction model, while greatly improving the accuracy of the model prediction, and ultimately effectively enhancing the application efficiency of the safety risk assessment system.
[0102] The matching priority in this embodiment is calculated based on the historical station data corresponding to several station IDs, including:
[0103] Obtain historical station data corresponding to several station IDs; historical station data includes historical risk level, historical risk number and occurrence time;
[0104] Calculate the comprehensive risk scores corresponding to several station IDs based on historical risk levels, historical risk times and occurrence time;
[0105] Obtain several source railway safety knowledge graphs and railway safety knowledge graphs corresponding to the target station ID;
[0106] Input several source railway safety knowledge graphs and the railway safety knowledge graphs corresponding to the target station ID into the graph similarity prediction model to obtain several similarity levels XD i; The graph similarity estimation model is constructed through the artificial intelligence model;
[0107] By formula PY i =h×tan -1 (α1×ZFP i +α2×XD i ) Calculate the matching priority PY i ; Where i represents the number of station IDs, g is the proportional coefficient, g∈(0,π / 2), and g is set to make the matching priority PY i ∈(0, 1), the specific value is set according to experience; α1 and α2 are weight coefficients, α1 and α2∈(0, 1), and the sum of α1 and α2 is 1. The specific value is set according to experience. When the importance of the similarity level is considered to be higher than the comprehensive risk score when calculating the matching priority, α1<α2 can be used. When the importance of the comprehensive risk score is considered to be higher than the similarity level, α1>α2 can be used. In this embodiment, the weight coefficient is set to α1<α2; ZFP i It is represented as the comprehensive risk score corresponding to the i-th station ID; as the comprehensive risk score and similarity level corresponding to the station increase, its corresponding matching priority will also increase.
[0108] This embodiment considers other stations from the two dimensions of comprehensive risk score and similarity level, and selects stations with high comprehensive risk scores and similarity levels as source stations for training data. This can obtain more data that meets the needs of the target station and is helpful for training the risk prediction model. This not only improves the accuracy of the prediction results of the risk prediction model of the target station, but also optimizes the deployment effectiveness of the model, allowing the risk prediction model to serve actual scenarios more efficiently and accurately.
[0109] In this embodiment, the comprehensive risk scores corresponding to several station IDs are calculated based on the historical risk level, the number of historical risks, and the time of occurrence, including:
[0110] Obtain historical risk level FD, historical risk times and occurrence time Δt;
[0111] By formula Calculation of comprehensive risk score ZFP i Wherein, j represents the number of historical risk events, max() represents the maximum value operation, max(j) represents the number of historical risk events, β represents the time decay coefficient, β∈(0,1), and the specific value is set based on experience. In this embodiment, β is set to 0.2. The time decay coefficient is set to take into account that the longer the historical risk events occur, the less impact they have on the comprehensive risk score; Δt i,jIt is expressed as the time from the time when the jth risk of the i-th station ID occurred. DT is expressed as the unit time. The specific value is set according to experience. In this embodiment, DT is set to 10 days. The sensitivity to the total number of risks is retained by using the log() function, but the occurrence of extreme values is suppressed by the logarithmic function. The shorter the time of risk event occurrence, the higher the number of risk occurrences and the risk level, the greater the corresponding comprehensive risk score.
[0112] This embodiment fully incorporates multi-source data such as the source station's historical risk level, historical risk frequency, and occurrence time when calculating the source station's comprehensive risk score, thereby improving the representativeness of the comprehensive risk score. This lays a solid foundation for screening source station data that meets the needs and strongly supports the subsequent search for suitable source station data.
[0113] The graph similarity prediction model in this embodiment is constructed using an artificial intelligence model, including:
[0114] Obtaining several historical source railway safety knowledge graphs and railway safety knowledge graphs corresponding to historical target station IDs and their corresponding historical similarity levels;
[0115] The railway safety knowledge graphs corresponding to several historical source railway safety knowledge graphs and historical target station IDs and their corresponding historical similarity levels are divided into training data, verification data, and test data; the training data, verification data, and test data are preprocessed to obtain training sets, verification sets, and test sets; the ratio of the training set, verification set, and test set is set to 7:2:1;
[0116] Select an artificial intelligence model as the basic model; artificial intelligence models include neural network models, etc.
[0117] Train the basic model using the training set, and adjust the learning rate and hyperparameters on the validation set to obtain the pre-trained model;
[0118] By verifying the pre-trained model on the test set, we finally obtained a railway safety knowledge graph whose input is several source railway safety knowledge graphs and corresponding to the target station ID, and the output is a graph similarity prediction model with several similarity levels.
[0119] In this embodiment, the risk prediction model is obtained by training the graph-enhanced time series prediction model on the target comprehensive data, including:
[0120] Obtaining target comprehensive data; the target comprehensive data includes a number of sub-graph data;
[0121] Extract several historical subgraph data and their corresponding historical time series data, historical risk levels, and risk events;
[0122] Integrate several historical subgraph data and their corresponding several historical time series data into several historical prediction data;
[0123] Several historical forecast data and their corresponding historical risk levels and risk events are divided into training data, validation data, and test data; data preprocessing is performed on the training data, validation data, and test data to obtain training sets, validation sets, and test sets; the ratio of the training set, validation set, and test set is set to 7:2:1;
[0124] Select the graph-enhanced time series prediction model as the basic model; the graph-enhanced time series prediction model includes the GETM model, etc.
[0125] Train the basic model using the training set, and adjust the learning rate and hyperparameters on the validation set to obtain the pre-trained model;
[0126] By verifying the pre-trained model on the test set, we finally obtain a risk prediction model whose input is prediction data and output is predicted risk level and predicted risk event.
[0127] In this embodiment, the generation of predicted risk events and predicted risk levels based on the risk prediction model includes:
[0128] Real-time acquisition of time series data and risk prediction models corresponding to the target station ID, as well as the railway safety knowledge graph; the railway safety knowledge graph includes several sub-graph data;
[0129] Extract several historical time series data and their corresponding historical subgraph data and integrate them with the current time series data and subgraph data to obtain prediction data;
[0130] The predicted data is input into the risk prediction model to obtain the predicted risk level and predicted risk events.
[0131] In this embodiment, generating an alarm signal according to the predicted risk level includes:
[0132] Obtain predicted risk levels and level ranges; level ranges include low risk range, medium risk range, and high risk range; level ranges are set by experts based on historical risk levels and corresponding historical risk events;
[0133] When the predicted risk level falls within the low-risk range, a warning signal for a low-risk event is generated;
[0134] When the predicted risk level falls within the medium risk range, a medium risk event alarm signal is generated;
[0135] When the predicted risk level belongs to the high risk range, an emergency risk alarm signal is generated.
[0136] Some of the data in the above formula are calculated by removing the dimensions and taking their numerical values. The formula is a formula that is closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.
[0137] The working principle of this application is as follows: station data is acquired through data acquisition equipment; a railway safety knowledge graph is constructed based on the station data; a risk prediction model is generated based on the railway safety knowledge graph; predicted risk events and predicted risk levels are generated based on the risk prediction model; an alarm signal is generated based on the predicted risk level; prompts are given based on the alarm signal, and management personnel are contacted, and the railway safety knowledge graph and risk prediction model are constructed through integration to enhance the accuracy of risk identification and achieve advance warning; multi-source heterogeneous data time series coding technology is adopted to unify cross-frequency data for comprehensive analysis; multi-source station risk data is used to train the target station prediction model to solve the problem of insufficient training data for the target station, improve the accuracy and application efficiency of the safety risk assessment system, and avoid the problem that the existing technology constructs risk prediction models in newly built stations and low-risk stations, which often has a small amount of training data, making the deployed safety risk assessment system time-consuming and inefficient.
[0138] The above embodiments are only used to illustrate the technical method of the present application and are not intended to limit it. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present application.
Claims
1. A safety risk assessment system based on railway station information, characterized in that: include: Data collection module, data analysis module, early warning module and database; The data acquisition module acquires station data through data acquisition equipment; the station data includes station ID, station purpose, equipment data, environmental data and operation data; The data analysis module is used to construct a railway safety knowledge graph based on station data; Generate a risk prediction model based on the railway safety knowledge graph; generate predicted risk events and predicted risk levels based on the risk prediction model; Generate alert signals based on predicted risk levels.
2. A railway station information-based safety risk assessment system according to claim 1, characterized in that: The construction of a railway safety knowledge graph based on station data includes: Acquire station data; the station data includes station ID, equipment data, environmental data and operation data; Generate specification data based on equipment data, environmental data, and operation data. The specification data refers to standard data used to construct a railway safety knowledge graph. The specification data includes several entities, several attributes, and several relationships. Use a graph database to build a safety knowledge graph, and embed spatiotemporal data storage and rule-driven systems into the safety knowledge graph. The spatiotemporal data storage is used to store the spatiotemporal data corresponding to the site data. The rule-driven systems are constructed by experts based on the site data and its corresponding risk level. Input the specification data into the safety knowledge graph to obtain the railway safety knowledge graph.
3. A railway station information-based safety risk assessment system according to claim 2, characterized in that: Generating specification data based on equipment data, environment data, and operation data includes: Acquire equipment data, environmental data, and operation data; the equipment data includes equipment ID, equipment status, and equipment parameters; the environmental data includes meteorological data and geological data; the operation data includes operation ID, operation instructions, maintenance records, and surveillance videos; Data classification is performed on the device data, environment data, and operation data to obtain text data, view data, time series data, and structured data; the structured data refers to data with a fixed format and clear fields; the view data refers to video data and image data; Standardized data is obtained by using a hierarchical fusion strategy for text data, view data, time series data and structured data; the hierarchical fusion strategy operates in a multimodal fusion manner.
4. A railway station information-based safety risk assessment system according to claim 3, characterized in that: The layered fusion strategy operates through a multimodal fusion approach and includes the following steps: Step 1: Obtain text data, view data, time series data, and structured data; Step 2: Extract features from text data, view data, time series data, and structured data to obtain a number of semantic vectors and spatial vectors; the semantic vectors and spatial vectors include a number of entities, attributes, and relationships; Step 3: Add a unified time code to several semantic vectors and align data of different frequencies through sliding window aggregation to obtain aligned semantic vectors; Step 4: Obtain normalized data by concatenating the aligned semantic vector and the spatial vector.
5. The railway station information-based safety risk assessment system according to claim 1, characterized in that: Generating a risk prediction model based on the railway safety knowledge graph includes: Obtain railway safety knowledge graphs and matching priorities PYi corresponding to several station IDs and railway safety knowledge graphs corresponding to target station IDs; the matching priorities are calculated using historical station data corresponding to the several station IDs; the target station ID is the railway station ID for which a risk prediction model needs to be constructed; Sort several matching priorities from largest to smallest to obtain a matching sequence PL; Select the station IDs and railway safety knowledge graphs corresponding to the top N matching priorities in PL, and name the railway safety knowledge graphs as source railway safety knowledge graphs; where N is an integer, N>1; Extract several common subgraph data from the railway safety knowledge graph corresponding to the target station ID and the source railway safety knowledge graph; Merging a plurality of common sub-graph data and the station data corresponding to the target station ID to obtain target comprehensive data; the target comprehensive data includes a plurality of sub-graph data; The risk prediction model is obtained by training the graph-enhanced time series prediction model on the target comprehensive data.
6. A railway station information-based safety risk assessment system according to claim 5, characterized in that: The matching priority is calculated based on the historical station data corresponding to several station IDs, including: Obtain historical station data corresponding to a number of station IDs; the historical station data includes historical risk levels, historical risk counts, and occurrence times; Calculate the comprehensive risk scores corresponding to several station IDs based on historical risk levels, historical risk times and occurrence time; Obtain several source railway safety knowledge graphs and railway safety knowledge graphs corresponding to the target station ID; Input several source railway safety knowledge graphs and the railway safety knowledge graphs corresponding to the target station ID into the graph similarity prediction model to obtain several similarity levels XD i ; The graph similarity prediction model is constructed through an artificial intelligence model; By formula PY i =g×tan -1 (α1×ZFP i +α2×XD i ) Calculate the matching priority PY i ; Where i represents the number of station IDs, g is the proportional coefficient, g∈(0,π / 2); α1 and α2 are weight coefficients, α1 and α2∈(0,1); and the sum of α1 and α2 is 1; ZFP i It is expressed as the comprehensive risk score corresponding to the i-th station ID.
7. A railway station information-based safety risk assessment system according to claim 6, characterized in that: The comprehensive risk scores corresponding to several station IDs are calculated based on the historical risk level, the number of historical risks and the time of occurrence, including: Obtain historical risk level FD, historical risk times and occurrence time Δt; By formula Calculation of comprehensive risk score ZFP i ; where j represents the number of historical risk events, max() represents the maximum value operation, max(j) represents the number of historical risk events, and β represents the time attenuation coefficient, β∈(0,1); Δt i,j It is expressed as the time from the time when the jth risk of the i-th station ID occurred, and DT is expressed as unit time.
8. The railway station information-based safety risk assessment system according to claim 6, characterized in that: The graph similarity prediction model is constructed using an artificial intelligence model, including: Obtaining several historical source railway safety knowledge graphs and railway safety knowledge graphs corresponding to historical target station IDs and their corresponding historical similarity levels; The railway safety knowledge graphs corresponding to several historical source railway safety knowledge graphs and historical target station IDs and their corresponding historical similarity levels are divided into training data, verification data, and test data; the training data, verification data, and test data are preprocessed to obtain a training set, a verification set, and a test set; Select an artificial intelligence model as the base model; Train the basic model using the training set, and adjust the learning rate and hyperparameters on the validation set to obtain the pre-trained model; By verifying the pre-trained model on the test set, we finally obtained a railway safety knowledge graph whose input is several source railway safety knowledge graphs and corresponding to the target station ID, and the output is a graph similarity prediction model with several similarity levels.
9. The railway station information-based safety risk assessment system according to claim 5, characterized in that: The risk prediction model is obtained by training the graph-enhanced time series prediction model on the target comprehensive data, including: Acquire target comprehensive data; the target comprehensive data includes a plurality of sub-graph data; Extract several historical subgraph data and their corresponding historical time series data, historical risk levels, and risk events; Integrate several historical subgraph data and their corresponding several historical time series data into several historical prediction data; Divide a number of historical forecast data and their corresponding historical risk levels and risk events into training data, verification data, and test data; perform data preprocessing on the training data, verification data, and test data to obtain a training set, a verification set, and a test set; Select the graph-enhanced time series prediction model as the base model; Train the basic model using the training set, and adjust the learning rate and hyperparameters on the validation set to obtain the pre-trained model; By verifying the pre-trained model on the test set, we finally obtain a risk prediction model whose input is prediction data and output is predicted risk level and predicted risk event.
10. The railway station information-based safety risk assessment system according to claim 1, characterized in that: The generating of predicted risk events and predicted risk levels according to the risk prediction model includes: Real-time acquisition of time series data and risk prediction models corresponding to the target station ID and a railway safety knowledge graph; the railway safety knowledge graph includes several subgraph data; Extract several historical time series data and their corresponding historical subgraph data and integrate them with the current time series data and subgraph data to obtain prediction data; The predicted data is input into the risk prediction model to obtain the predicted risk level and predicted risk events.
Citation Information
Patent Citations
Method and apparatus for constructing knowledge map
CN108268581A
Safety risk prediction method and system based on community equipment and facilities
CN115964503A
Road-related engineering traffic safety early warning and protection system
CN118692237A
Smart station monitoring visualization system and method based on multi-source heterogeneous data fusion
CN119047848A
Method for constructing natural disaster risk knowledge graph of photovoltaic project
CN119443218A