Multi-source heterogeneous data knowledge graph construction method for railway disaster prevention monitoring

By building a multi-source heterogeneous data knowledge map for railway disaster prevention monitoring, the data fusion problem in the existing system is solved, efficient and accurate knowledge representation in the field of railway disaster prevention is achieved, and intelligent disaster prevention decision-making capabilities are improved.

CN120492447AInactive Publication Date: 2025-08-15EAST CHINA JIAOTONG UNIVERSITY

Patent Information

Application Number
CN202510991838.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing railway disaster prevention monitoring system is difficult to effectively integrate multi-source heterogeneous data, resulting in low information processing efficiency and lagging decision-making response, lack of ability to identify complex disaster patterns and accurate prediction of risk situations, and it is difficult to meet the needs of intelligent disaster prevention.

Method used

A multi-source heterogeneous data knowledge graph construction method for railway disaster prevention monitoring is adopted. Through the data preparation, ontology construction, knowledge extraction and fusion stages, a high-quality knowledge graph is constructed using the hybrid knowledge extraction engine and domain ontology constraints to achieve accurate extraction and fusion of entities, attributes and relationships.

Benefits of technology

The built knowledge graph can accurately and comprehensively reflect the complex characteristics of railway disaster prevention, improve data processing efficiency and intelligent decision-making capabilities, and support risk prediction and intelligent early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492447A_ABST
    Figure CN120492447A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source heterogeneous data knowledge graph construction method for railway disaster prevention monitoring, and relates to the technical field of knowledge graph construction, and the method comprises the steps: gathering multi-source heterogeneous data related to railway disaster prevention monitoring, and constructing a domain ontology model used for guiding knowledge extraction and fusion; extracting entities, attributes and relationships among the entities from different modal data after standardization preprocessing by using a targeted extraction algorithm; obtaining fused structured knowledge based on a multi-strategy knowledge fusion process of domain ontology constraint and confidence evaluation; the fused structured knowledge is stored in a graph database, and construction of the knowledge graph in the railway disaster prevention monitoring field is completed; through combination of domain ontology construction, a mixed knowledge extraction engine and a multi-strategy knowledge fusion technology, deep semantic fusion of multi-source heterogeneous data in the railway field is realized. The invention aims to construct a knowledge graph capable of comprehensively and accurately reflecting complex characteristics in the railway disaster prevention field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph construction, and in particular to a method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring. Background Art

[0002] Establishing an efficient and intelligent disaster prevention monitoring system is the cornerstone of ensuring transportation safety. Currently, railway disaster prevention monitoring primarily relies on data collection from various sensor networks deployed along the line, combined with manual inspections and preset threshold rules to trigger warnings. However, these traditional monitoring methods have significant limitations when addressing the complex and ever-changing needs of railway disaster prevention. A core challenge lies in data processing and coordination. Monitoring data comes from a wide range of sources, including structured sensor readings, semi-structured maintenance logs, unstructured text reports and images and videos, as well as external meteorological and geological information. This data not only has diverse formats and spatiotemporal scales, but more importantly, it exhibits significant semantic differences and is often stored in disconnected systems, creating "data silos" that hinder information flow and significantly limit the possibility of comprehensive, contextual analysis of railway system status. This fragmented and heterogeneous data directly leads to inefficient information processing and delayed decision-making responses. Due to the difficulty in effectively integrating and deeply mining dispersed data, existing systems are often limited to post-event alerts based on simple thresholds. They lack the ability to identify complex disaster patterns and accurately and proactively assess risk trends, making it difficult to meet the growing demand for refined, intelligent disaster prevention. Furthermore, the vast amount of valuable knowledge contained in historical disaster cases, expert experience, and industry standards, due to its unstructured and implicit nature, is difficult for traditional systems to effectively absorb and utilize, further limiting the level of intelligent monitoring, early warning, and emergency decision-making.

[0003] To address these challenges, knowledge graph technology is considered an effective approach for achieving semantic fusion and knowledge-based management of multi-source, heterogeneous data. By providing a structured, networked representation of key entities and their complex relationships in the field of railway disaster prevention, knowledge graphs can break down data barriers and provide a unified knowledge foundation for in-depth analysis and intelligent reasoning.

[0004] Despite the enormous potential of knowledge graphs, building a high-quality knowledge graph capable of effectively supporting intelligent decision-making in the highly specialized and dynamically changing field of railway disaster prevention and monitoring remains a technical challenge. Existing general knowledge graph construction methods struggle to directly address the challenges of efficiently and accurately extracting valuable knowledge elements (entities, relationships, and attributes) automatically or semi-automatically from massive, real-time, multimodal, heterogeneous data streams, particularly those containing extensive railway terminology and implicit associations, and integrating them into a knowledge graph that comprehensively and accurately reflects the complex nature of the railway disaster prevention domain. This requires targeted approaches to address challenges such as the specificity of domain data, the accuracy of knowledge extraction, and the timeliness of graph updates. Therefore, there is an urgent need to develop a new approach specifically tailored to the needs of railway disaster prevention and monitoring that can effectively integrate and process multi-source heterogeneous data, overcome the limitations of existing technologies, and construct a high-quality domain knowledge graph. This approach will lay a solid data and knowledge foundation for advancing intelligent railway disaster prevention and monitoring. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring, aiming to construct a knowledge graph that can comprehensively and accurately reflect the complex characteristics of the railway disaster prevention field.

[0006] In one aspect, the present invention proposes a method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring, the method comprising: Data preparation phase: Aggregate multi-source heterogeneous data related to railway disaster prevention and monitoring, and perform standardized preprocessing on the aggregated multi-source heterogeneous data. Standardization preprocessing includes data cleaning, format conversion, spatiotemporal coordinate reference system 1, and missing value filling. Ontology construction phase: Analyze domain documents containing railway domain knowledge using a preset topic model, discover the core concepts and potential semantic relationships between them within the railway domain, and screen, revise, supplement, and structure the potential semantic relationships between them to form the initial version of the railway disaster prevention domain ontology. Knowledge extraction phase: A hybrid knowledge extraction engine is launched. The hybrid knowledge extraction engine extracts entities, attributes, and relationships between entities from text data, time series data, and image / video data using corresponding extraction algorithms from multi-source heterogeneous data after standardization and preprocessing, forming knowledge triples. Knowledge fusion stage: Using the railway disaster prevention domain ontology as a semantic reference, the entity alignment algorithm based on vector space embedding is used to align the entities extracted from multi-source heterogeneous data. The knowledge triples are then logically verified for consistency and confidence assessed using preset ontology constraint rules. Knowledge triples that pass the logical consistency check and have a confidence level above the threshold are identified as fused knowledge triples. Graph construction stage: The fused knowledge triples are formatted according to the preset data model and then stored in the preset graph database to complete the construction of the multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring.

[0007] Furthermore, in the above-mentioned method for constructing a knowledge graph of multi-source heterogeneous data for railway disaster prevention monitoring, the multi-source heterogeneous data includes at least one or more of sensor monitoring data, text data, meteorological data, and geographic spatial information; Sensor monitoring data at least includes real-time or historical monitoring data streams generated by various sensors deployed on lines, bridges, and tunnels; Text data includes at least manually completed inspection records, equipment maintenance logs, historical disaster event reports, and accident analysis documents; Meteorological data shall at least include weather forecasts and real-time meteorological data issued by meteorological departments; Geographical spatial information includes at least the geographical location of the railway, the line topology, and the surrounding geological and hydrological environment.

[0008] Furthermore, in the above-mentioned method for constructing a knowledge graph of multi-source heterogeneous data for railway disaster prevention and monitoring, the domain documents include at least railway safety regulations, accident reports, technical standards, and research papers; The steps of screening, correcting, supplementing and structuring the core concepts and the potential semantic relationships between the core concepts include: The core concepts and the potential semantic relationships between them are compared and verified with a pre-built rule base containing the experience and knowledge of railway experts. The core concepts and the potential semantic relationships between them are screened, corrected, supplemented and structured through human-computer collaboration.

[0009] Furthermore, in the above-mentioned method for constructing a knowledge graph of multi-source heterogeneous data for railway disaster prevention monitoring, the hybrid knowledge extraction engine includes: At least one natural language processing model optimized for text data and trained on domain corpus to extract entities, attributes, and relationships from text; At least one spatiotemporal correlation mining algorithm for time series data that integrates graph neural networks and attention mechanisms to identify abnormal patterns of sensor collaboration and infer their association with related entities or events; and At least one recognition model that integrates a computer vision model for image / video data to identify specific objects in the image / video and treat them as entities, while extracting their location and time and associating them with the corresponding geographic location or device entity.

[0010] Furthermore, in the above-mentioned method for constructing a knowledge graph of multi-source heterogeneous data for railway disaster prevention monitoring, the steps of extracting entities, attributes, and relationships between entities from text data, time series data, and image / video data using corresponding extraction algorithms to form knowledge triples include: A natural language processing model based on the Transformer architecture is developed through continuous pre-training and instruction fine-tuning based on railway domain corpus. The natural language processing model is used to extract entities, attributes, and relationships between entities from text data to form knowledge triples. The sensor network in the monitoring area is modeled as a graph structure using a spatiotemporal association mining algorithm. The time series data and graph structure information are fused through a specific GNN model to dynamically learn the state representation of each sensor node. Analyze the changes in the state representation of sensor nodes or the differences with other nodes, identify abnormal readings of single sensors and collaborative abnormal patterns among multiple sensor data, and infer their association with related entities or events to form knowledge triples; Using the recognition model, specific objects in the image are identified and treated as entities. At the same time, their location and time are extracted and associated with the corresponding geographic location or device entity to form a knowledge triple.

[0011] Furthermore, in the above-mentioned method for constructing a knowledge graph of multi-source heterogeneous data for railway disaster prevention monitoring, the steps of using the railway disaster prevention domain ontology as a semantic reference and sequentially adopting an entity alignment algorithm based on vector space embedding to align entities extracted from the multi-source heterogeneous data include: The entity vector space embedding representation based on the constraints of the domain ontology is adopted, and the similarity or distance score between the embedded vectors is calculated to determine whether the entities are aligned. The entities determined to be aligned are mapped to the same instance in the domain ontology.

[0012] Furthermore, in the above-mentioned method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring, the steps of performing logical consistency verification and confidence evaluation on knowledge triples using preset ontology constraint rules, and determining knowledge triples that pass the logical consistency verification and have a confidence level higher than a threshold as fused knowledge triples include: For the extracted knowledge triples, the knowledge that obviously violates the domain ontology definition is eliminated through the preset ontology constraint rules; Based on the reliability score of the knowledge source in the knowledge triple, the confidence output during extraction, and the logical consistency verification results with the domain ontology, a final confidence score is calculated and assigned to the fused knowledge, and the knowledge triple with a final confidence score higher than the threshold is determined as the fused knowledge triple.

[0013] Another object of the present invention is to provide a device for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring, the device comprising: The preparation module is used to aggregate multi-source heterogeneous data related to railway disaster prevention and monitoring during the data preparation phase, and perform standardized preprocessing on the aggregated multi-source heterogeneous data. Standardized preprocessing includes data cleaning, format conversion, spatiotemporal coordinate reference system 1, and missing value filling. The mining module is used in the ontology construction phase to analyze domain documents containing railway domain knowledge using a preset topic model, to mine the core concepts and potential semantic relationships between them in the railway domain, and to screen, amend, supplement, and structure the potential semantic relationships between them to form the initial version of the railway disaster prevention domain ontology. The extraction module is used to start a hybrid knowledge extraction engine in the knowledge extraction stage. The hybrid knowledge extraction engine extracts entities, attributes, and relationships between entities from text data, time series data, and image / video data using corresponding extraction algorithms from multi-source heterogeneous data after standardization preprocessing to form knowledge triples. The fusion module is used in the knowledge fusion stage to: use the railway disaster prevention domain ontology as a semantic reference, sequentially adopt the entity alignment algorithm based on vector space embedding to align the entities extracted from multi-source heterogeneous data, perform logical consistency verification and confidence assessment on the knowledge triples based on preset ontology constraint rules, and determine the knowledge triples that pass the logical consistency verification and have a confidence level above the threshold as the fused knowledge triples; The construction module is used in the graph construction stage to format the fused knowledge triples according to the preset data model and then store them in the preset graph database to complete the construction of the multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring.

[0014] Another object of the present invention is to provide a readable storage medium having a computer program stored thereon, wherein the program implements the steps of the above method when executed by a processor.

[0015] Another object of the present invention is to provide an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the steps of the above method are implemented when the processor executes the program.

[0016] The present invention aggregates and pre-processes multi-source heterogeneous data related to railway disaster prevention monitoring in accordance with preset rules, and builds and dynamically maintains a railway disaster prevention domain ontology model for guiding knowledge extraction and fusion based on an iterative technology combining domain document topic mining with expert rule verification; starts a hybrid knowledge extraction engine, and uses targeted extraction algorithms to extract entities, attributes, and relationships between entities from different modal data after standardized pre-processing; based on a multi-strategy knowledge fusion process of domain ontology constraints and confidence evaluation, the heterogeneous knowledge extracted in the knowledge extraction step is subjected to entity alignment, conflict resolution, and confidence confirmation to obtain fused structured knowledge; and the fused The structured knowledge is stored in a graph database according to a predetermined graph data model, completing the construction of a knowledge graph in the field of railway disaster prevention and monitoring. The clear and specific ontology construction technology, innovative hybrid knowledge extraction engine, and rigorous multi-strategy knowledge fusion process significantly improve the efficiency, accuracy, and depth of automatically constructing high-quality knowledge graphs from complex and heterogeneous data in the railway disaster prevention field, effectively addressing the semantic gap and knowledge representation challenges of multi-source data. The constructed high-fidelity domain knowledge graph can more accurately and comprehensively depict the railway disaster prevention knowledge system, providing a more powerful and reliable knowledge foundation for downstream intelligent applications such as risk prediction, intelligent early warning, and decision support. This solves the problem of existing technologies that make it difficult to construct a knowledge graph that can comprehensively and accurately reflect the complex characteristics of the railway disaster prevention field. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flowchart of a method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring in the first embodiment of the present invention; Figure 2 This is a structural block diagram of the device for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring in the second embodiment of the present invention.

[0018] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION

[0019] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate several embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present invention.

[0020] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is referred to as being "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one skilled in the art to which this invention pertains. The terms used in this specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0022] Example 1 See also Figure 1 , shown is a method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring in the first embodiment of the present invention, and the method includes steps S10 to S14.

[0023] Step S10, data preparation stage: gather multi-source heterogeneous data related to railway disaster prevention monitoring, and perform standardized preprocessing on the gathered multi-source heterogeneous data. The standardized preprocessing includes data cleaning, format conversion, spatiotemporal coordinate reference system 1, and missing value filling.

[0024] This phase involves the aggregation and preprocessing of multi-source heterogeneous data. Through system-preset data interfaces and adapters (e.g., API interfaces, direct database connections, file parsers, etc.), multi-source heterogeneous data related to railway disaster prevention and monitoring are compulsorily aggregated. These data sources include, but are not limited to: real-time or historical monitoring data streams generated by various sensors (such as rainfall, wind speed, displacement, vibration, temperature, and images) deployed on key infrastructure such as lines, bridges, and tunnels; unstructured or semi-structured text data such as manually completed inspection records, equipment maintenance logs, historical disaster event reports, and accident analysis documents; weather forecasts and real-time meteorological data released by meteorological departments; and geospatial information data (such as GIS data) containing information such as the railway's geographical location, line topology, and the surrounding geological and hydrological environment.

[0025] The aggregated raw multi-source heterogeneous data then enters a standardized preprocessing process, which strictly implements operations such as data cleaning (noise removal and outlier processing), format conversion (unified data encoding and structure), spatiotemporal coordinate reference system 1 (for example, the unified use of WGS84 coordinate system and UTC time standard), and missing value filling (using interpolation or prediction models) to ensure data quality and lay the foundation for subsequent processing.

[0026] Step S11, ontology construction phase: Use the preset topic model to analyze the domain documents containing railway domain knowledge, dig out the core concepts in the railway field and the potential semantic relationships between the core concepts, screen, correct, supplement and structure the potential semantic relationships between the core concepts, and form the initial version of the railway disaster prevention domain ontology.

[0027] Among them, this stage applies an iterative ontology construction technology based on the combination of domain document topic mining and expert rule verification to generate and maintain a dynamic domain ontology model (Ontology) that accurately describes the knowledge in the field of railway disaster prevention.

[0028] First, topic models such as LDA (Latent Dirichlet Allocation) are used to analyze large-scale documents in the fields of railway safety regulations, accident reports, technical standards, research papers, etc., to automatically mine core concepts in the field (such as "slope instability", "contact network icing", "rail deformation", etc.) and the potential semantic relationships between them.

[0029] Then, the mining results are compared and verified with a pre-built rule base containing the experience knowledge of railway experts. Through human-computer collaboration, the automatically mined concepts and relationships are screened, corrected, supplemented and structured (defining hierarchical and subordinate relationships, attributes, constraints, etc.) to form the initial version of the railway disaster prevention domain ontology.

[0030] After that, during the operation of the system, the domain ontology model can be iteratively optimized and dynamically expanded based on new entity types and new relationships that appear in newly accessed data or based on the knowledge continuously fed back by experts, so as to maintain its timeliness and accuracy.

[0031] Step S12, start a hybrid knowledge extraction engine, which uses corresponding extraction algorithms to extract entities, attributes and relationships between entities from text data, time series data and image / video data from multi-source heterogeneous data after standardization preprocessing to form knowledge triples.

[0032] Among them, the hybrid knowledge extraction engine includes: At least one natural language processing model optimized for text data and trained on domain corpus to extract entities, attributes, and relationships from text; At least one spatiotemporal correlation mining algorithm for time series data that integrates graph neural networks and attention mechanisms to identify abnormal patterns of sensor collaboration and infer their association with related entities or events; and At least one recognition model that integrates a computer vision model for image / video data to identify specific objects in the image / video and treat them as entities, while extracting their location and time and associating them with the corresponding geographic location or device entity.

[0033] Specifically, for knowledge extraction from text data, a language model based on a Transformer architecture is deployed, which undergoes continuous pre-training and instruction tuning based on railway domain corpus. Leveraging its powerful semantic understanding capabilities, this model is specifically designed to accurately identify predefined railway entities (such as "Tunnel Exit No. 6," "Category HXD1001 Pillar," and "Level III Flood Warning") from text data such as inspection reports and accident records. It then extracts complex semantic relationships between these entities (such as "located at," "affects," "causes," "component belongs to," and "status is") to form knowledge triples. Furthermore, the model's training data encompasses railway industry standards, historical reports, maintenance manuals, and expertly annotated corpus, ensuring a high level of understanding of railway terminology and domain-specific expressions.

[0034] When extracting knowledge from time series data, a spatio-temporal anomaly association mining algorithm that integrates a graph neural network (GNN) and an attention mechanism is applied. The algorithm first models the sensor network in the monitoring area as a graph structure G = (V, E), where V is a set of sensor nodes and E represents the physical proximity or functional association between sensors. Then, a specific GNN model (such as STGAT - Spatio-Temporal Graph Attention Network) is used to fuse time series data (series data) with graph structure information to dynamically learn the state representation of each sensor node. The following exemplary update rule calculates the state representation of node i at time t :

[0035] in, is a nonlinear activation function, 、 、 are model parameters / functions, For nodes i The neighbor set of is the node calculated by the spatiotemporal attention mechanism j For Node i In time t Dynamic impact weight.

[0036] By analyzing the node status representation The algorithm can identify abnormal readings of a single sensor and coordinated abnormal patterns among multiple sensor data (such as regional settlement and cascading failures), and infer their association with specific physical entities or potential events (such as "slope monitoring points P01-P05 have excessive displacements at the same time, which is associated with heavy rainfall event E01") to form knowledge triples.

[0037] When extracting knowledge from image / video data, computer vision (CV) models (such as the YOLO series of target detection models and image segmentation models) are integrated to identify specific objects in the image (such as landslides, foreign objects on the track, and equipment damage status) and treat them as entities. At the same time, their attribute information such as location and time is extracted and associated with the corresponding geographic location or equipment entity to form a knowledge triple.

[0038] Step S13, knowledge fusion stage: using the railway disaster prevention domain ontology as a semantic reference, sequentially adopting the entity alignment algorithm based on vector space embedding to align the entities extracted from the multi-source heterogeneous data, performing logical consistency verification and confidence evaluation on the knowledge triples through the preset ontology constraint rules, and determining the knowledge triples that pass the logical consistency verification and have a confidence level higher than the threshold as the fused knowledge triples.

[0039] This stage implements a multi-strategy knowledge fusion process based on ontology constraints and confidence assessment to integrate heterogeneous knowledge extracted from different sources and modalities. The specific operations are as follows: Entity alignment: Using the defined domain ontology as a semantic reference, an entity alignment algorithm based on vector space embedding is employed. First, a low-dimensional embedding representation is learned for all extracted entities (e.g., using TransE trained on graphs, Rotat, or entity embeddings based on pre-trained language models).

[0040] Then, by calculating the entity embedding vector , embedding vector The similarity score between to identify different entity mentions that refer to the same real-world object. For example, using a distance-based scoring function ,in, represents the norm, p Used to specify the type of norm, q is a power operation on the norm result, bias It is a learnable or adjustable offset parameter. The higher the score calculated by the scoring function, the greater the possibility of entity alignment. Ultimately, the entities determined to be aligned are mapped to the same instance in the ontology.

[0041] Relationship / attribute conflict resolution and confidence assessment: for the extracted relationship and attribute triplesT , design a conflict resolution mechanism based on rule reasoning and statistical verification. First, perform logical consistency verification through preset ontology constraint rules (such as domain, range constraints, function dependencies, etc.) ∈{0,1}, eliminating knowledge that obviously violates the ontology definition.

[0042] Then, the reliability score of the data source of the knowledge is comprehensively considered (For example, data sources from authoritative reports are more reliable than ordinary sensor data), the original confidence level of the output during extraction , calculate the final confidence of the fused knowledge .

[0043] An exemplary calculation formula is:

[0044] Where f(⋅) is a fusion function, such as a weighted linear combination or a more complex belief propagation algorithm. This process is used to resolve conflicts when there are multiple different extraction results for the same entity relationship or attribute (for example, selecting the one with the highest confidence or merging them according to a specific strategy). Knowledge confirmation: Ultimately, only those who have passed the ontology consistency check ( =1) and the final confidence The knowledge triples above the preset threshold are output as fused knowledge triples consisting of high-quality structured knowledge.

[0045] Step S14, graph construction phase: format the fused knowledge triples according to the preset data model, and then store them in the preset graph database to complete the construction of the multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring.

[0046] This phase integrates knowledge triples (or other graph representations, such as property graphs) and formats them according to standard RDF (Resource Description Framework) or property graph data models. Efficient batch import tools (such as Bulk Loader) are then used to load and persist this knowledge data into a performance-optimized distributed graph database system (for example, JanusGraph, Nebula Graph, or Neo4j, which are indexed and optimized for railway spatiotemporal data queries). This completes the construction of a multi-source, heterogeneous data knowledge graph for railway disaster prevention and monitoring. This graph stores a wealth of knowledge on railway facilities, operating environment, disaster risks, monitoring status, and early warning responses in a structured and semantically organized manner.

[0047] In summary, the method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring in the above embodiment of the present invention generates a password by combining the unique code of the medical device with a customized preset character, thereby ensuring a strong correlation between the password and the medical device, and the password of each medical device is unique. It prevents medical devices from being accessed by unauthorized personnel, protects the patient's privacy data and key parameters of the device, and greatly improves the security of the medical device. The user's initial password and the maintenance-level password are generated separately, meeting the security requirements of different management levels of medical equipment. Medical staff use the user's initial password for daily operations, while equipment maintenance personnel use the maintenance password for advanced management operations such as equipment maintenance and upgrades, avoiding confusion in password permissions. It solves the problem of low security and easy confusion of permissions in the existing method of generating medical device passwords.

[0048] Example 2 See also Figure 2 , shown is a device for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring proposed in the second embodiment of the present invention, the device comprising: Preparation module 100 is used to: gather multi-source heterogeneous data related to railway disaster prevention and monitoring, and perform standardized preprocessing on the gathered multi-source heterogeneous data during the data preparation phase. The standardized preprocessing includes data cleaning, format conversion, spatiotemporal coordinate reference system 1, and missing value filling. The mining module 200 is used to analyze the domain documents containing railway domain knowledge using a preset topic model in the ontology construction stage, mine the core concepts in the railway domain and the potential semantic relationships between the core concepts, and screen, modify, supplement and structure the potential semantic relationships between the core concepts to form an initial version of the railway disaster prevention domain ontology; The extraction module 300 is used to: start a hybrid knowledge extraction engine in the knowledge extraction stage. The hybrid knowledge extraction engine extracts entities, attributes, and relationships between entities from text data, time series data, and image / video data using corresponding extraction algorithms from multi-source heterogeneous data after standardization preprocessing to form knowledge triples; Fusion module 400 is used, during the knowledge fusion stage, to: utilize the railway disaster prevention domain ontology as a semantic reference, sequentially employ an entity alignment algorithm based on vector space embedding to align entities extracted from multi-source heterogeneous data, perform logical consistency verification and confidence assessment on knowledge triples using preset ontology constraint rules, and determine knowledge triples that pass the logical consistency verification and have a confidence level above a threshold as fused knowledge triples; Construction module 500 is used to format the fused knowledge triples according to the preset data model during the graph construction phase, and then store them in the preset graph database to complete the construction of the multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring.

[0049] The functions or operation steps implemented when the above modules are executed are substantially the same as those in the above method embodiments and will not be repeated here.

[0050] Example 3 Another aspect of the present invention further provides a readable storage medium having a computer program stored thereon, which implements the steps of the method described in the above embodiment 1 when the program is executed by a processor.

[0051] Example 4 Another aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the program, the steps of the method described in the above embodiment 1 are implemented.

[0052] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0053] Those skilled in the art will appreciate that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or for use in conjunction with such instruction execution system, apparatus, or device. For purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device.

[0054] More specific examples (a non-exhaustive list) of computer-readable storage media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable storage medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0055] It should be understood that various components of the present invention may be implemented using hardware, software, firmware, or a combination thereof. In the aforementioned embodiments, multiple steps or methods may be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following technologies known in the art may be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0056] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0057] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A method for constructing a knowledge graph of multi-source heterogeneous data for railway disaster prevention monitoring, characterized in that: The method comprises: Data preparation phase: Aggregate multi-source heterogeneous data related to railway disaster prevention and monitoring, and perform standardized preprocessing on the aggregated multi-source heterogeneous data. Standardization preprocessing includes data cleaning, format conversion, spatiotemporal coordinate reference system 1, and missing value filling. Ontology construction phase: Analyze domain documents containing railway domain knowledge using a preset topic model, discover the core concepts and potential semantic relationships between them within the railway domain, and screen, revise, supplement, and structure the potential semantic relationships between them to form the initial version of the railway disaster prevention domain ontology. Knowledge extraction phase: A hybrid knowledge extraction engine is launched. The hybrid knowledge extraction engine extracts entities, attributes, and relationships between entities from text data, time series data, and image / video data using corresponding extraction algorithms from multi-source heterogeneous data after standardization and preprocessing, forming knowledge triples. Knowledge fusion stage: Using the railway disaster prevention domain ontology as a semantic reference, the entity alignment algorithm based on vector space embedding is used to align the entities extracted from multi-source heterogeneous data. The knowledge triples are then logically verified for consistency and confidence assessed using preset ontology constraint rules. Knowledge triples that pass the logical consistency check and have a confidence level above the threshold are identified as fused knowledge triples. Graph construction stage: The fused knowledge triples are formatted according to the preset data model and then stored in the preset graph database to complete the construction of the multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring.

2. The method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring according to claim 1 is characterized in that: Multi-source heterogeneous data includes at least one or more of sensor monitoring data, text data, meteorological data, and geographic spatial information; Sensor monitoring data at least includes real-time or historical monitoring data streams generated by various sensors deployed on lines, bridges, and tunnels; Text data includes at least manually completed inspection records, equipment maintenance logs, historical disaster event reports, and accident analysis documents; Meteorological data shall at least include weather forecasts and real-time meteorological data issued by meteorological departments; Geographical spatial information includes at least the geographical location of the railway, the line topology, and the surrounding geological and hydrological environment.

3. The method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring according to claim 1 is characterized in that: Domain documents include at least railway safety regulations, accident reports, technical standards, and research papers; The steps of screening, correcting, supplementing and structuring the core concepts and the potential semantic relationships between the core concepts include: The core concepts and the potential semantic relationships between them are compared and verified with a pre-built rule base containing the experience and knowledge of railway experts. The core concepts and the potential semantic relationships between them are screened, corrected, supplemented and structured through human-computer collaboration.

4. The method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring according to claim 1 is characterized in that: The hybrid knowledge extraction engine includes: At least one natural language processing model optimized for text data and trained on domain corpus to extract entities, attributes, and relationships from text; At least one spatiotemporal correlation mining algorithm for time series data that integrates graph neural networks and attention mechanisms to identify abnormal patterns of sensor collaboration and infer their association with related entities or events; and At least one recognition model that integrates a computer vision model for image / video data to identify specific objects in the image / video and treat them as entities, while extracting their location and time and associating them with the corresponding geographic location or device entity.

5. The method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring according to claim 4 is characterized in that: The steps of extracting entities, attributes, and relationships between entities from text data, time series data, and image / video data using corresponding extraction algorithms to form knowledge triples include: A natural language processing model based on the Transformer architecture is developed through continuous pre-training and instruction fine-tuning based on railway domain corpus. The natural language processing model is used to extract entities, attributes, and relationships between entities from text data to form knowledge triples. The sensor network in the monitoring area is modeled as a graph structure using a spatiotemporal association mining algorithm. The time series data and graph structure information are fused through a specific GNN model to dynamically learn the state representation of each sensor node. Analyze the changes in the state representation of sensor nodes or the differences with other nodes, identify abnormal readings of single sensors and collaborative abnormal patterns among multiple sensor data, and infer their association with related entities or events to form knowledge triples; Using the recognition model, specific objects in the image are identified and treated as entities. At the same time, their location and time are extracted and associated with the corresponding geographic location or device entity to form a knowledge triple.

6. The method for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring according to claim 1 is characterized in that: The steps of using the railway disaster prevention domain ontology as a semantic reference and sequentially adopting a vector space embedding-based entity alignment algorithm to align entities extracted from multi-source heterogeneous data include: The entity vector space embedding representation based on the constraints of the domain ontology is adopted, and the similarity or distance score between the embedded vectors is calculated to determine whether the entities are aligned. The entities determined to be aligned are mapped to the same instance in the domain ontology.

7. The method according to claim 6, characterized in that The step of performing logical consistency verification and confidence evaluation on the knowledge triples using preset ontology constraint rules, and determining the knowledge triples that pass the logical consistency verification and have a confidence level higher than a threshold as fused knowledge triples includes: For the extracted knowledge triples, the knowledge that obviously violates the domain ontology definition is eliminated through the preset ontology constraint rules; Based on the reliability score of the knowledge source in the knowledge triple, the confidence output during extraction, and the logical consistency verification results with the domain ontology, a final confidence score is calculated and assigned to the fused knowledge, and the knowledge triple with a final confidence score higher than the threshold is determined as the fused knowledge triple.

8. A device for constructing a multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring, characterized in that: The device comprises: The preparation module is used to aggregate multi-source heterogeneous data related to railway disaster prevention and monitoring during the data preparation phase, and perform standardized preprocessing on the aggregated multi-source heterogeneous data. Standardized preprocessing includes data cleaning, format conversion, spatiotemporal coordinate reference system 1, and missing value filling. The mining module is used in the ontology construction phase to analyze domain documents containing railway domain knowledge using a preset topic model, to mine the core concepts and potential semantic relationships between them in the railway domain, and to screen, amend, supplement, and structure the potential semantic relationships between them to form the initial version of the railway disaster prevention domain ontology. The extraction module is used to start a hybrid knowledge extraction engine in the knowledge extraction stage. The hybrid knowledge extraction engine extracts entities, attributes, and relationships between entities from text data, time series data, and image / video data using corresponding extraction algorithms from multi-source heterogeneous data after standardization preprocessing to form knowledge triples. The fusion module is used in the knowledge fusion stage to: use the railway disaster prevention domain ontology as a semantic reference, sequentially adopt the entity alignment algorithm based on vector space embedding to align the entities extracted from multi-source heterogeneous data, perform logical consistency verification and confidence assessment on the knowledge triples based on preset ontology constraint rules, and determine the knowledge triples that pass the logical consistency verification and have a confidence level above the threshold as the fused knowledge triples; The construction module is used in the graph construction stage to format the fused knowledge triples according to the preset data model and then store them in the preset graph database to complete the construction of the multi-source heterogeneous data knowledge graph for railway disaster prevention monitoring.

9. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the steps of the method according to any one of claims 1 to 7 are implemented when the processor executes the program.

Citation Information

Patent Citations

  • Knowledge graph construction method based on disaster scene

    CN109992672A

  • Multi-source heterogeneous network security knowledge graph construction method and device

    CN112131882A

  • Multi-source heterogeneous data fusion method based on geographic entity

    CN113065000A

  • Space-time knowledge graph construction system and method oriented to dynamic analysis

    CN114860884A

  • Analysis processing and knowledge graph construction method based on multi-source heterogeneous big data

    CN115640406A

Cited By

  • Multi-source heterogeneous data fusion analysis method and system

    CN120654077A

  • Railway construction simulation model generation method based on full-line design parameters

    CN120745261A

  • Self-adaptive risk assessment method and system for disaster prevention and reduction emergency drill

    CN120782267A

  • Railway emergency aid decision-making method and system

    CN120822711A

  • Cross-modal knowledge graph construction method

    CN120851177A