Railway accident factor identification and relationship extraction method, system and device and medium

By using text mining technology to clean and extract features from historical railway accident documents, a knowledge base with a multi-layered causative factor system is constructed. This solves the problem of insufficient utilization of unstructured data in existing technologies, enables effective identification and relationship extraction of causative factors of railway accidents, and improves the early warning capability of railway safety operation.

CN116821360BActive Publication Date: 2026-05-01RD CENT CHINA ACADEMY OF RAILWAY SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RD CENT CHINA ACADEMY OF RAILWAY SCI
Filing Date
2023-06-01
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing railway accident analysis methods neglect the multidimensionality and correlation of accident-causing factors, make it difficult to effectively utilize massive amounts of unstructured data, and rely on expert experience, which has limitations and cannot comprehensively analyze the safety and reliability of railway systems.

Method used

Using a text mining-based approach, we cleaned, extracted features, and transformed the data from historical railway accident documents to construct a knowledge base of multi-layered causative factors. We then identified and extracted accident causative factors and their relationships, and used expert domain knowledge for feature annotation.

Benefits of technology

It enables the effective utilization of massive amounts of unstructured data, constructs a unified standard domain knowledge base, identifies accident causative factors and their relationships, provides data support, and provides basic theoretical support for railway safety operation early warning and rectification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821360B_ABST
    Figure CN116821360B_ABST
Patent Text Reader

Abstract

The application discloses a railway accident factor identification and relationship extraction method, comprising the following steps: defining a regular expression design for accident investigation and analysis reports, accident identification books and other multi-source railway historical accident documents, extracting text paragraphs describing events in the documents, performing data cleaning, and obtaining an accident text dataset; performing sentence segmentation, word segmentation, part-of-speech tagging, named entity recognition and dependency syntax structure analysis on the accident text, performing feature extraction and structured storage; constructing a multi-layer cause factor system containing man-machine loop management, performing feature labeling by related field experts to form a knowledge base, and then proposing a cause factor identification and relationship extraction method based on the text features and containing a three-layer structure. The application also discloses a railway accident factor identification and relationship extraction system. The method is reasonable and effective in utilizing railway historical accident documents, forming a knowledge base by using expert field knowledge, and then identifying accident cause factors and constructing relationships between the factors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of railway accident analysis technology. More specifically, it relates to a method for identifying causal factors and extracting relationships in railway accidents based on text mining. Background Technology

[0002] Currently, the railway system is one of my country's most crucial infrastructures, holding an irreplaceable and critical position in the comprehensive transportation system. As a large-scale ground transportation system vital to the safety of passengers' lives and property, the safety and reliability of the railway system are of paramount importance. However, as a complex system, the railway system has a high degree of coupling between its various elements and complex interface relationships between its subsystems. Even a small change in a single factor can lead to a rapid deterioration in the behavior of the entire system, thus creating hidden dangers for safe railway operations and potentially triggering major railway accidents. Therefore, how to avoid accidents and improve the railway's safe transportation capacity is a critical issue that urgently needs to be addressed for safe railway operations. Analyzing historical railway accident documents, identifying accident-causing factors, and recognizing the nonlinear relationships between these factors are of significant practical importance for effectively predicting accident risk points, improving risk warning technology, refining operational management strategies, and achieving accident prevention and control in the railway system.

[0003] Current railway accident analyses primarily rely on expert experience to conduct single-factor analyses of historical accident data, neglecting the multidimensionality and interrelationships of accident causative factors, or employing comprehensive evaluation methods that artificially weight different factors for accident assessment. While existing research has developed certain theoretical methods, these methods have limitations. Firstly, they are constrained by the experiential knowledge of domain experts. Railway safety operations involve different stages and specialties, and knowledge barriers exist between experts in different fields, making comprehensive analysis from a systemic perspective difficult. Secondly, they primarily rely on structured data, neglecting the effective utilization of the ever-accumulating massive amounts of unstructured data. With the continuous development of railway system operation and management, the railway industry has established a sensor network covering fixed railway facilities, mobile equipment, and the environment along the railway lines nationwide, accumulating massive amounts of business information related to railway traffic safety. Among these, the largest, longest-retained, and most valuable text files in the field of railway traffic safety are railway accident factual documents. These unstructured text data, as carriers of key accident information, contain rich value and urgently need to be explored through text mining to uncover the hidden patterns of accident occurrence within the text, thereby providing decision support for shifting railway traffic safety from passive to proactive safety.

[0004] Text mining is the entire process of extracting unknown, understandable, and useful knowledge from unstructured text data, involving sub-tasks such as data acquisition, storage, retrieval, feature extraction, and mining analysis. Text mining methods have been widely applied in various fields and have achieved high practical value.

[0005] Therefore, compared with current traditional accident analysis methods, there is an urgent need to propose a text mining-based method for identifying causative factors and extracting relationships in railway accidents. Taking historical railway accident documents as the research object, this method transforms unstructured text data into structured data through document conversion, data cleaning, and feature extraction techniques. Key features are extracted using text mining methods, and feature annotation is performed using expert domain knowledge to form a knowledge base. A three-layered causative factor identification method is then constructed to identify accident causative factors and extract their relationships. This method, based on the full utilization of massive amounts of historical railway accident data, transforms expert domain knowledge into a knowledge base, avoiding the limitations and subjectivity of domain experts, and constructing a standardized domain knowledge base. This effectively learns causative factors and their relationships from historical accidents. Summary of the Invention

[0006] This application provides a method for identifying causal factors and extracting relationships in railway accidents based on text mining, in order to solve the problem of effectively utilizing massive amounts of unstructured data.

[0007] In a first aspect, embodiments of this application provide a method for identifying railway accident factors and extracting relationships, including:

[0008] Steps for obtaining historical accident text dataset: Based on the paragraph and chapter layout features of historical railway accident documents from various sources, define regular expressions to extract text paragraphs describing events from the historical accident documents, perform data cleaning, and obtain a valid historical accident text dataset.

[0009] The structured feature extraction steps are as follows: After segmenting the effective historical accident text dataset into words based on a pre-constructed railway domain vocabulary, part-of-speech tagging and named entity recognition are performed based on the word segmentation. After generating dependency syntax structure from the part-of-speech tagging results, the structured features of the historical accident text dataset are extracted and stored in a structured manner.

[0010] The steps for identifying and classifying causative factors are as follows: 1. A knowledge base is constructed by labeling the structured features of historical accident text data. 2. Based on the knowledge base, an accident causative factor identification method containing multiple layers of causative factors is constructed to identify and classify the accident causative factors into a multi-layered accident causative factor set.

[0011] The steps for extracting causal factor relationships are as follows: Based on the multi-layer accident causal factor set, sort and combine them to construct an accident causal factor chain, thereby realizing the extraction of accident causal factor relationships.

[0012] Preferably, the above-mentioned historical accident text dataset acquisition step further includes:

[0013] Text format conversion steps: For historical railway accident documents containing multiple formats and sources, a unified file encoding method is used to convert the file types to obtain files with recognizable formats;

[0014] The steps for obtaining valid text are as follows: Analyze the paragraph and chapter layout features of files with recognizable formats, design regular expressions, filter and clean irrelevant railway historical accident texts, and obtain a railway historical accident text dataset composed of valid railway historical accident texts.

[0015] Preferably, the above-mentioned structured feature extraction further includes:

[0016] Word segmentation steps: For historical railway accident texts, sentences are divided according to punctuation marks to obtain a sentence set. The railway domain vocabulary includes: a railway domain stop word list and a railway domain personalized word segmentation list. A pre-trained word segmentation model is used, combined with the railway domain stop word list and the railway system personalized word segmentation list, to perform word segmentation on the sentence set to obtain the word segmentation results.

[0017] Part-of-speech tagging steps: Based on the word segmentation results of the railway historical accident text, a pre-trained part-of-speech tagging model is used for part-of-speech tagging;

[0018] Named entity recognition steps: Based on the word segmentation results of the railway historical accident text, a pre-trained named entity recognition model is used to perform named entity recognition;

[0019] Part-of-speech tagging step: Based on the part-of-speech tagging results obtained from the part-of-speech tagging step, the railway historical accident text is segmented and categorized by word, retaining the preset valid word classes, and the categorized results are spliced ​​together to form a new text and new corpus corresponding to the railway historical accident text;

[0020] The steps for calculating the term frequency-inverse document frequency (TF-IDF) value are as follows: For the new corpus, the TF-IDF values ​​of the selected effective word classes in each railway historical accident text are calculated, and the representative scores of the words in different documents are calculated.

[0021] Supplementary optimization steps: By selecting representative words with high scores from each document, the word segmentation step is repeated up to the word frequency-reverse file frequency value calculation step to supplement and optimize the stop word list and personalized word segmentation list for the railway field;

[0022] Dependency syntactic structure recognition steps: Based on the word segmentation results and part-of-speech tagging results of the railway historical accident text, a pre-trained dependency syntactic analysis model is used to identify dependency syntactic structures. Multiple dependency syntactic structures can be obtained from the segmentation, forming tuple features and storing them in a structured manner.

[0023] Preferably, the above-mentioned causative factor identification and classification steps further include:

[0024] Steps for constructing the causative factor system: Construct a multi-layered causative factor system based on human-machine-environment-management, and form a mapping relationship between causative factor classification labels and descriptions;

[0025] Knowledge base construction steps: Annotate the structured features of the text data to build a multi-layered knowledge base, which includes: mapping relationships, keyword thesaurus, and dependency structure table;

[0026] Steps for obtaining the causative factor set: After cleaning and calculating text features for the historical railway accident text dataset, the causative factors of the accident are classified and identified by constructing a causative factor identification system with a multi-layer structure based on the knowledge base, generating a candidate set of causative factors. By fusing and deduplicating the causative factors, the causative factor set corresponding to the test historical railway accident text dataset is obtained.

[0027] Preferably, the above knowledge base construction steps further include:

[0028] Foreign language words with preset tags are obtained from the text of historical railway accidents. The foreign language words correspond to the railway accident level in the text. Based on the foreign language words, they are annotated in a multi-layer causal factor system to obtain specific tags, and a mapping relationship is constructed in the knowledge base.

[0029] The TF-IDF values ​​of words in the railway historical accident text dataset are sorted from high to low. Verbs, nouns, and entity words in the preset effective word classes are selected as candidates. Experts annotate the keywords, and the keywords are accumulated in the knowledge base into the keyword list corresponding to the label category.

[0030] For the subject-predicate relation, verb-object relation, and adverbial-head structure in the dependency syntax structure of railway historical accident texts, the representative scores of the dependency syntax structure are sorted from high to low. Experts annotate the key dependency structures and accumulate the key dependency structures in the dependency structure table corresponding to the tag category in the knowledge base.

[0031] Preferably, the above-mentioned step of obtaining the causative factor set further includes:

[0032] Foreign language words with preset tag types in the test railway historical accident text are mapped according to the mapping relationship to obtain the first layer of causal factors;

[0033] The verb and noun results of the test railway historical accident text were retrieved using a keyword thesaurus to obtain the hit keyword sequence. The hit keyword sequence was then filtered for adjacent occurrences of the same causative factor, and the filtered hit keyword sequence was obtained as the second layer of causative factors.

[0034] The dependency structure table is used to retrieve the dependency syntax structure of the test railway historical accident text to obtain the sequence of hit dependency syntax structure. Adjacent hit structures of the same cause factor are filtered to obtain the filtered sequence of hit dependency structure as the third-level cause factor. The candidate set of cause factors includes: first-level cause factor, second-level cause factor and third-level cause factor.

[0035] Preferably, the above-mentioned causal factor relationship extraction step further includes:

[0036] The hit sequences of keywords and dependency syntax structures are deduplicated and fused. The fusion process combines the writing logic of railway historical accident texts and the inherent relationship of the multi-layered causal factor system of human-machine-environment-management. Factors belonging to the preset category are sorted in advance to obtain composite sequences.

[0037] Based on the complex sequence, the factors are classified according to the factor categories of the first-level causative factors as follows:

[0038] If the causative factors corresponding to the first-level causative factors are contained in the composite sequence, then the causative factor relationship of the accident is a composite sequence, thus obtaining the causative factor relationship chain;

[0039] If the causative factors corresponding to the first-level causative factors are human or management factors and are not included in the composite sequence, then the causative factors are combined into the composite sequence and ordered as the first position after the management causative factors to obtain the causative factor relationship chain.

[0040] If the causative factors corresponding to the first-level causative factors do not belong to the human or management categories and are not included in the composite sequence, then they are sorted into the last item of the composite sequence to obtain the causative factor relationship chain.

[0041] Secondly, embodiments of this application provide a railway accident factor identification and relation extraction system, employing the above-mentioned railway accident factor identification and relation extraction method. The railway accident factor identification and relation extraction system includes:

[0042] Historical accident text dataset acquisition module: Based on the paragraph and chapter layout features of historical railway accident documents from various sources, regular expressions are defined to extract text paragraphs describing events from historical accident documents, perform data cleaning, and obtain a valid historical accident text dataset.

[0043] The structured feature extraction module performs word segmentation on the effective historical accident text dataset based on a pre-constructed railway domain vocabulary, performs part-of-speech tagging and named entity recognition based on the word segmentation, generates dependency syntax structure from the part-of-speech tagging results, and then performs structured feature extraction and structured storage of the historical accident text dataset.

[0044] Causative Factor Identification and Classification Module: The structured features of historical accident text data are labeled to construct a knowledge base. Based on the knowledge base, a causative factor identification method containing multiple causative factors is constructed to identify accident causative factors and classify them to obtain a multi-level accident causative factor set.

[0045] Causative Factor Relationship Extraction Module: Based on the multi-layer accident causative factor set, sort and combine them to construct an accident causative factor chain, thereby realizing the extraction of accident causative factor relationships.

[0046] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the railway accident factor identification and relationship extraction method described above.

[0047] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the railway accident factor identification and relationship extraction method described above.

[0048] Compared to existing technologies, it has the following outstanding advantages:

[0049] 1) The method of this invention is based on the theoretical method of text mining. It makes full use of the historical documents of railway accidents and transforms unstructured data into structured data through document conversion, data cleaning and feature extraction. Furthermore, it uses text mining methods to extract key features, identify the causative factors of railway accidents and their relationships, which can effectively guide on-site personnel to prevent key accidents and provide data support for early warning and rectification of actual accident failure risks, thereby ensuring the safe operation of railways.

[0050] 2) When mining railway accident texts, the method of this invention needs to construct a knowledge base related to railway text data. Taking historical railway accident documents as the research object, the unstructured text data is transformed into a structured form through document conversion, data cleaning, feature extraction and other techniques. Key features are extracted using text mining methods, and feature annotation is performed using expert domain knowledge to form a knowledge base.

[0051] 3) The method of this invention constructs a causative factor identification method with a three-layer structure, identifies the causative factors of accidents, and extracts the relationships between the causative factors. Based on the full utilization of massive historical railway accident data, this method transforms expert domain knowledge into a knowledge base, avoiding the limitations and subjectivity of domain experts, and constructing a unified standard domain knowledge base, effectively learning causative factors and their relationships from historical accidents. Attached Figure Description

[0052] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0053] Figure 1 This is a flowchart of the railway accident factor identification and relationship extraction method of the present invention;

[0054] Figure 2 This is a schematic diagram of the railway accident factor identification and relationship extraction process according to a specific embodiment of the present invention;

[0055] Figure 3 This is a schematic diagram of the railway accident factor identification and relationship extraction system of the present invention;

[0056] Figure 4 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of this application.

[0057] In the above image:

[0058] 10. Historical accident text dataset acquisition module; 20. Structured feature extraction module.

[0059] 30 Causative Factor Identification and Classification Module; 40 Causative Factor Relationship Extraction Module Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the application is described and illustrated below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0061] Obviously, the accompanying drawings described below are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without any inventive effort. Furthermore, it is understood that although the efforts made in this development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, any changes to design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as insufficient disclosure of the content of this application.

[0062] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent.

[0063] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that is mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0064] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms “a,” “an,” “an,” “the,” and similar words used in this application do not indicate quantity limitation and may indicate singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units not listed, or may include other steps or units inherent to these processes, methods, products, or devices. The terms “connected,” “linked,” “coupled,” and similar words used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following objects are in an "or" relationship. The terms "first," "second," and "third" used in this application are merely to distinguish similar objects and do not represent a specific ordering of the objects.

[0065] This invention aims to provide a method for identifying causal factors and extracting relationships in railway accidents based on text mining. The method includes the following steps: S1: For multi-source historical railway accident documents such as accident investigation and analysis reports and accident determination reports, define regular expressions to extract text paragraphs describing events from the documents, perform data cleaning, and obtain an accident text dataset; S2: Perform sentence segmentation, word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing analysis on the accident text, extract features, and store them in a structured manner; S3: Construct a multi-layered causal factor system containing "human-machine-environment-management" (hereinafter also referred to as human-machine-environment-management), have relevant domain experts perform feature annotation to form a knowledge base, and then propose a causal factor identification and relationship extraction method based on text features and containing a three-layer structure. This method is based on text mining theory, overcomes the limitations of expert domain experience and manual analysis, fully utilizes massive amounts of unstructured text information, and employs intelligent analysis methods to identify and extract causal factors and relationships in railway accidents. Experimental results show that this method can make reasonable and effective use of historical railway accident documents, form a knowledge base from expert domain knowledge, identify accident causative factors, construct relationships between factors, and thus uncover the accident occurrence patterns hidden in the accident texts. This provides data support for actual accident and fault risk early warning and remediation, and provides basic theoretical support for accident prevention and control in the railway system.

[0066] like Figure 1 As shown in the embodiments of this application, a method for identifying railway accident factors and extracting relationships is provided, including:

[0067] Step S10 for obtaining historical accident text dataset: Based on the paragraph and chapter layout features of historical railway accident documents from various sources, define regular expressions, extract text paragraphs describing events from the historical accident documents, perform data cleaning, and obtain a valid historical accident text dataset.

[0068] S20: After segmenting the effective historical accident text dataset into words based on a pre-constructed railway domain vocabulary, part-of-speech tagging and named entity recognition are performed based on the word segmentation. After generating dependency syntax structure from the part-of-speech tagging results, the structured features of the historical accident text dataset are extracted and stored in a structured manner.

[0069] Step S30: The structured features of historical accident text data are labeled to construct a knowledge base. Based on the knowledge base, the accident causative factors are identified by constructing a causative factor identification method containing multiple causative factors, and a multi-level accident causative factor set is obtained by classification.

[0070] Step S40: Based on the multi-layer accident causative factor set, sort and combine them to construct an accident causative factor chain, thereby realizing the extraction of accident causative factor relationships.

[0071] Preferably, the above-mentioned historical accident text dataset acquisition step S10 further includes:

[0072] Text format conversion steps: For historical railway accident documents containing multiple formats and sources, a unified file encoding method is used to convert the file types to obtain files with recognizable formats;

[0073] The steps for obtaining valid text are as follows: Analyze the paragraph and chapter layout features of files with recognizable formats, design regular expressions, filter and clean irrelevant railway historical accident texts, and obtain a railway historical accident text dataset composed of valid railway historical accident texts.

[0074] Preferably, the above-mentioned structured feature extraction S20 further includes:

[0075] Word segmentation steps: For historical railway accident texts, sentences are divided according to punctuation marks to obtain a sentence set. The railway domain vocabulary includes: a railway domain stop word list and a railway domain personalized word segmentation list. A pre-trained word segmentation model is used, combined with the railway domain stop word list and the railway system personalized word segmentation list, to perform word segmentation on the sentence set to obtain the word segmentation results.

[0076] Part-of-speech tagging steps: Based on the word segmentation results of the railway historical accident text, a pre-trained part-of-speech tagging model is used for part-of-speech tagging;

[0077] Named entity recognition steps: Based on the word segmentation results of the railway historical accident text, a pre-trained named entity recognition model is used to perform named entity recognition;

[0078] Part-of-speech tagging step: Based on the part-of-speech tagging results obtained from the part-of-speech tagging step, the railway historical accident text is segmented and categorized by part of speech, and the preset valid word classes are retained. The categorized results are then concatenated to form a new text and new corpus corresponding to the railway historical accident text. In a specific embodiment of the present invention, the preset valid word classes are nouns and verbs, but the present invention is not limited to these, and other valid word classes can also be set.

[0079] The steps for calculating the term frequency-inverse document frequency (TF-IDF) value are as follows: For the new corpus, the TF-IDF values ​​of the selected effective word classes in each railway historical accident text are calculated, and the representative scores of the words in different documents are calculated.

[0080] Supplementary optimization steps: By selecting representative words with high scores from each document, the word segmentation step is repeated up to the word frequency-reverse file frequency value calculation step to supplement and optimize the stop word list and personalized word segmentation list for the railway field;

[0081] Dependency syntactic structure recognition steps: Based on the word segmentation results and part-of-speech tagging results of the railway historical accident text, a pre-trained dependency syntactic analysis model is used to identify dependency syntactic structures. Multiple dependency syntactic structures can be obtained from the segmentation, forming tuple features and storing them in a structured manner.

[0082] Preferably, the above-mentioned causative factor identification and classification step S30 further includes:

[0083] Steps for constructing the causative factor system: Construct a multi-layered causative factor system based on human-machine-environment-management, and form a mapping relationship between causative factor classification labels and descriptions;

[0084] Knowledge base construction steps: Annotate the structured features of the text data to build a multi-layered knowledge base, which includes: mapping relationships, keyword thesaurus, and dependency structure table;

[0085] Steps for obtaining the causative factor set: After cleaning and calculating text features for the historical railway accident text dataset, the causative factors of the accident are classified and identified by constructing a causative factor identification system with a multi-layer structure based on the knowledge base, generating a candidate set of causative factors. By fusing and deduplicating the causative factors, the causative factor set corresponding to the test historical railway accident text dataset is obtained.

[0086] Preferably, the above knowledge base construction steps further include:

[0087] Foreign language words with preset tags are obtained from historical railway accident texts. The foreign language words correspond to the railway accident level in the text. Based on the foreign language words, they are annotated in a multi-layer causal factor system to obtain specific tags, and a mapping relationship is constructed in a knowledge base. In a specific embodiment of the present invention, the preset foreign language word is ws, but the present invention is not limited to this and other preset tags can also be set.

[0088] The TF-IDF values ​​of words in the railway historical accident text dataset are sorted from high to low. Words with parts of speech such as verbs, nouns, and entities are selected as candidates. Experts annotate the keywords, and the keywords are accumulated in the knowledge base into the keyword list corresponding to the label category.

[0089] For the subject-predicate relation, verb-object relation, and adverbial-head structure in the dependency syntax structure of railway historical accident texts, the representative scores of the dependency syntax structure are sorted from high to low. Experts annotate the key dependency structures and accumulate the key dependency structures in the dependency structure table corresponding to the tag category in the knowledge base.

[0090] Preferably, the above-mentioned step of obtaining the causative factor set further includes:

[0091] Foreign language words with preset tag types in the test railway historical accident text are mapped according to the mapping relationship to obtain the first layer of causal factors;

[0092] The verb and noun results of the test railway historical accident text were retrieved using a keyword thesaurus to obtain the hit keyword sequence. The hit keyword sequence was then filtered for adjacent occurrences of the same causative factor, and the filtered hit keyword sequence was obtained as the second layer of causative factors.

[0093] The dependency structure table is used to retrieve the dependency syntax structure of the test railway historical accident text to obtain the sequence of hit dependency syntax structure. Adjacent hit structures of the same cause factor are filtered to obtain the filtered sequence of hit dependency structure as the third-level cause factor. The candidate set of cause factors includes: first-level cause factor, second-level cause factor and third-level cause factor.

[0094] Preferably, the above-mentioned causal factor relationship extraction step S40 further includes:

[0095] The keyword and dependency syntax structure hit sequences are deduplicated and fused. The fusion process combines the writing logic of railway historical accident texts and the inherent relationship of the multi-layered causal factor system of human-machine-environment-management. Factors belonging to the preset category are sorted in advance to obtain a composite sequence. In a specific embodiment of the present invention, the preset category factor is the management factor, but the present invention is not limited to this and other category factors can also be set.

[0096] Based on the complex sequence, the factors are classified according to the factor categories of the first-level causative factors as follows:

[0097] If the causative factors corresponding to the first-level causative factors are contained in the composite sequence, then the causative factor relationship of the accident is a composite sequence, thus obtaining the causative factor relationship chain;

[0098] If the causative factors corresponding to the first-level causative factors are human or management factors and are not included in the composite sequence, then the causative factors are combined into the composite sequence and ordered as the first position after the management causative factors to obtain the causative factor relationship chain.

[0099] If the causative factors corresponding to the first-level causative factors do not belong to the human or management categories and are not included in the composite sequence, then they are sorted into the last item of the composite sequence to obtain the causative factor relationship chain.

[0100] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings:

[0101] Figure 2 The diagram shown is a flowchart of a railway accident factor identification and relationship extraction method according to a specific embodiment of the present invention. Figure 2As shown, the purpose of this invention is to provide a method for identifying causal factors and extracting relationships in railway accidents based on text mining. The method includes the following steps:

[0102] S1: For multi-source historical railway accident documents (training sample data), such as accident investigation and analysis reports and accident determination letters, define regular expressions to extract text paragraphs describing events from the documents, perform data cleaning, and obtain an effective historical railway accident text dataset.

[0103] S1.1: Convert the file types of historical railway accident documents containing multiple sources such as PDF and Word using a document parser, and unify the file encoding to UTF-8 to obtain a TXT file that can be used for subsequent identification and analysis;

[0104] S1.2: Analyze the paragraph and chapter layout features of the txt file, design a regular expression Re to filter and clean irrelevant railway historical accident text, such as chapter titles, thereby obtaining an effective railway historical accident text dataset D = {d1, d2, ..., d...} N}, and d i =Re(d i ′), d i ' represents the railway historical accident text corresponding to the initial txt file.

[0105] S2: For the effective railway historical accident text dataset obtained in S1, feature extraction is performed to achieve structured storage of railway historical accident text features.

[0106] S2.1: For a specific railway historical accident text d i Sentences are divided according to punctuation marks to obtain a sentence set d. i ={s1,s2,…,s n Based on the commonly used Chinese stop word list, personalized stop words for the railway field are added to construct a stop word list T-hit_stopwords for railway historical accident texts. Simultaneously, a personalized word segmentation list T_lexicon is constructed for railway system professional terminology. The pre-trained word segmentation model cws.model from the Python open-source library pyltp is used, combined with the stop word list T-hit_stopwords and the personalized word segmentation list T_lexicon, to target railway historical accident texts. i Clauses s j (j=1,2,…n) is segmented into words to obtain the segmentation result s. j ={w j1 ,w j2 ,…,w jm}

[0107] S2.2: Regarding historical railway accident texts di Clause s j ={w j1 ,w j2 ,...,w jm The word segmentation results were processed using the pre-trained part-of-speech tagging model pos.model from the pyltp tool, which includes 29 parts of speech such as adjectives (labeled as a), adverbs (labeled as d), nouns (labeled as n), and verbs (labeled as v). That is, for any word in a clause, it has a unique part-of-speech tag pos(s) within its respective clause. j ,w jk ).

[0108] S2.3: Regarding historical railway accident texts d i Clause s j ={w j1 ,w j2 ,…,w jm The word segmentation results were processed using the pre-trained named entity recognition model ner.model from the pyltp tool for named entity recognition, specifically including four types: person names (labeled as Nh), organization names (labeled as Ni), place names (labeled as Ns), and non-entity names (labeled as O). That is, for any word in a sentence, it has a unique named entity label ner(s) within its respective sentence. j ,w jk ).

[0109] S2.4: Based on the part-of-speech tagging results obtained in S2.2, the part-of-speech tagging of each accident text d is performed. i Word segmentation and part-of-speech filtering were performed, retaining only nouns and verbs. The filtered words were then concatenated using spaces to form a text corresponding to historical railway accidents. i The new text f(d) i Thus, D = {d1, d2, ..., d} N} Form a new corpus corpus={f(d1),f(d2),…,f(d N )}. Calculate the historical accident text f(d) for each railway in corpus. i The TF-IDF (Term Frequency-Inverse Document Frequency) values ​​of nouns and verbs after filtering represent the representativeness of each word to the document. Specifically,

[0110]

[0111] Where, n ij This indicates that word i is in document f(d) j The number of times it appears in ) ∑ k n kj This indicates the total number of words contained in the document. |corpus| represents the total number of documents in the corpus, |j:ti ∈f(d j )| represents the number of documents containing word i. From this, the representative score of word i in different documents can be calculated. By selecting words with high representative scores from each document, the personalized word segmentation table and stop word table are supplemented and optimized, thus repeating steps S2.1 to S2.4, and then proceeding to S2.5 after optimization.

[0112] S2.5: Include historical railway accident texts. i Clause s j ={w j1 ,w j2 ,...,w jm The word segmentation results and the part-of-speech tags pos(s) obtained in S2.2 j ,w jk Using the input as features, the pre-trained dependency parsing model `parser.model` from the pyltp tool is employed to identify dependency syntax structures. The dependency syntax structure labels include 14 structures such as subject-verb relations (labeled SBV), verb-object relations (labeled VOB), and adverbial-head structures (labeled ADV). That is, clauses s j ={w j1 ,w j2 ,...,w jm Multiple dependency syntax structures can be obtained from}, forming triplet features arc j,k = (relation, head, rely_word), where j and k represent the order of the clause and the order of the matched dependency parent node word, respectively; relation represents the dependency syntactic structure tag; head represents the dependency parent node word; and rely_word represents the matched dependency parent node word. The average of head and rely_word is calculated using the TF-IDF values ​​of each segment obtained in S2.4, denoted as avg_score(arc j,k This is used as a representative score for the dependency syntax structure.

[0113] S3: Construct a multi-layered causal factor system including "human-machine-environment-management", form a knowledge base by feature annotation by experts in relevant fields, and propose a causal factor identification and relation extraction method with a three-layer structure based on text features based on the knowledge base.

[0114] S3.1: Considering that the main causes of railway accidents are unsafe human behavior, unsafe equipment conditions, unsafe environmental conditions, and management deficiencies and technical shortcomings, this design incorporates a multi-layered causative factor system L={l1,l2,...,l m This forms a mapping relationship between the causative factor classification labels and their descriptions.

[0115] S3.2: A knowledge base is formed by experts using their domain knowledge and experience to annotate the structured features of the text data. Specifically, this begins by acquiring historical railway accident text d. i Foreign words tagged with "ws" in the text correspond to railway accident levels (e.g., general A / B / C / D category accidents). Based on this, they are labeled in a multi-level causal factor system L to obtain specific labels l, and their mapping relationship is recorded as ws_dict. l Secondly, based on the TF-IDF values ​​of the words, they are sorted from high to low, and words with parts of speech such as verbs, nouns, and entities are selected as candidates. Experts then annotate the keywords, and these keywords are accumulated in the keyword list kw_dict corresponding to the tag category. l In the middle; finally, regarding the text of historical railway accidents d i Subject-verb relations (labeled SBV), verb-object relations (labeled VOB), and adverbial-head structures (labeled ADV) in dependency syntax are analyzed using avg_score(arc j,k The scores are sorted from highest to lowest, and key dependency structures are annotated by experts. These key dependency structures are then accumulated in the dependency structure table `arc_dict` corresponding to the label category. l This forms a three-layer knowledge base, specifically containing ws_dict = {ws_dict} l ,l∈L},kw_dict={kw_dict l ,l∈L} and arc_dict={arc_dict l ,l∈L}.

[0116] S3.3: A text feature knowledge base built upon S3.2, targeting the test railway historical accident text dataset TD={td1,td2,…,td N After completing the data cleaning and text feature calculation for S1 and S2, a three-layer causative factor identification method was constructed to identify accident causative factors. Specifically, firstly, foreign words labeled as ws in the test railway historical accident text td were mapped according to ws_dict to obtain the first-layer causative factor class1 (class1∈L); secondly, kw_dict was used to retrieve the verb and noun results after word segmentation of the test railway historical accident text td to obtain the hit keyword sequence. Considering that the description of the accident in the railway historical accident text contains multiple causative factors and the expression feature of repeated occurrence of the same causative factor, the hit keywords of adjacent occurrence of the same causative factor in the hit keyword sequence were filtered. The filtered hit keyword sequence is denoted as [(ind1,w1,TF-IDF1,l1),...,(ind i ,w iTF-IDF i ,l i ),...], where ind i Indicates keyword w i TF-IDF indicates the location of the historical railway accident text (td) during testing. i Indicates keyword w i TF-IDF value, l i (l i ∈L) is the keyword w i The corresponding causative factor categories in the knowledge base kw_dict are used to obtain the filtered keyword sequence, which is the second-level causative factor class2. Finally, the dependency syntax structure of the test railway historical accident text td is retrieved using arc_dict to obtain the sequence of hit dependency syntax structures. Similarly, the hit structures of adjacent occurrences of the same causative factor are filtered to obtain the filtered hit dependency structure sequence, which is the third-level causative factor class3, denoted as [(ind1,arc1,avg_score1,l1),...,(ind i ,arc i ,avg_score i ,l i ),...], where ind i representing dependency syntax structure arc i The location of avg_score in the test railway historical accident text td. i For the dependency syntax structure importance score obtained from S2.5, l i (l i ∈L) is a dependency syntax structure arc i The corresponding causative factor categories are found in the knowledge base arc_dict. Therefore, class1, class2, and class3 together constitute the candidate set of causative factors. By fusing and deduplicating the causative factors, the causative factor set corresponding to the test railway accident text dataset td is obtained.

[0117] S3.4 sorts and combines the candidate causative factors class1, class2, and class3 obtained in S3.3 to construct a causative factor chain. During the construction process, the hit sequences of keywords and dependency syntax structures are arranged according to ind... i The process involves deduplication and fusion. Taking into account the writing logic of historical railway accident texts and the inherent relationship between the "human-machine-environment-management" factors, management-related factors are prioritized, resulting in a composite sequence l_list, specifically represented as l1→...→l i →..., where l i ≠l i+1 And l i∈L. Based on this, three cases are divided according to the factor category of class1. Case 1: The causative factors corresponding to class1 are included in l_list, then the causative factor relationship of the accident is l1→...→l i →...; Case 2: If the causative factor l corresponding to class1 belongs to the human or management category and is not included in l_list, then it is combined into l_list and sorted into the first position after the management category causative factors to obtain the causative factor relationship chain; Case 3: If the causative factor corresponding to class1 does not belong to the human or management category and is not included in l_list, then it is sorted into the last item of l_list to obtain the causative factor relationship chain.

[0118] In a specific embodiment of the present invention, a preferred embodiment uses 123 accident and fault investigation reports from a railway bureau's train operation system from 2011 to 2020 (a total of ten years) as training sample data. First, the training sample data undergoes file format conversion and regular expression extraction to obtain an accident text dataset (corresponding to S1). Next, feature extraction is performed on the accident text, including word segmentation, part-of-speech tagging, named entity tagging, TF-IDF calculation, and dependency syntax structure annotation, thereby achieving structured storage of the accident text features (corresponding to S2). Furthermore, a stop word list T_hit_stopwords (see Table 1) with capacities of 817 and 48, and a personalized word segmentation word list T_lexicon (see Table 2), both oriented towards railway system terminology, are constructed.

[0119] Table 1. Stop Word List (Partial)

[0120] Column 1 Column 2 Column 3 Column 4 —— Not only take and 》), Not only rush Moreover )÷(1- not only remove and ”, in spite of besides Instead )、 not only unless And outside =( but Apart from In other words : not only this That's all → Unrestricted here After that ℃ whether take advantage of in turn & Not afraid Taking advantage of Conversely … … … …

[0121] Table 2. Personalized Word Segmentation Table (Partial)

[0122] Column 1 Column 2 Column 3 Iron shoes Anti-slip iron shoes Connector brake Sliding car group Duty Officer Remove hat Aolibugao Electrical workers Wearing a hat Train number frame Brake operator Vehicle Service Unloading ballast switchman The bell rang. Public Works Liaison signal mutation Dispatcher Exhaust brake operator Reverse Lock Shunting foreman Car number clerk Arrival Line Protective personnel Train inspection personnel … … …

[0123] Using the accident text from a railway bureau's train operation system on March 24, 2014, as test data, the results of its structured features, including text segmentation, part-of-speech tagging, named entity tagging, and TF-IDF calculation, are shown in Tables 3 and 4. In Table 4, sub_sentence represents the sentence identifier of the entire accident text, and the sentence number starts from 0.

[0124] Table 3. Text Structure Results of the March 24, 2014 Event (Partial)

[0125]

[0126]

[0127] Table 4. Results of text dependency segmentation for the event on March 24, 2014

[0128]

[0129]

[0130] Based on feature extraction from the training sample data, a multi-layered causal factor system was designed by domain experts (corresponding to S3.1). This multi-layered causal factor system involves 17 labels from different perspectives of "human (R)-machine (J)-environment (H)-management (G)", and feature annotation was performed to obtain a three-layered knowledge base (corresponding to S3.2), namely ws_dict, kw_dict, and arc_dict (see Tables 5, 6, and 7).

[0131] Table 5. A single-layer knowledge base based on text-based incident levels

[0132] Causative factor label Accident levels in the text R-1 C8 R-1 C9 R-1 D15 R-2 D6 … … J-6 C19 J-8 C23 J-9 D3 J-10 D5

[0133] Table 6. Keyword-based two-tier knowledge base (partial)

[0134]

[0135]

[0136] Table 7. Partial Three-Layer Knowledge Base Based on Dependency Structure

[0137]

[0138] Based on the knowledge base built in S3.2, the causative factor identification and relation extraction methods (corresponding to S3.3 and S3.4) are used to obtain the accident causative factor chain. Using the accident text of a railway operation system accident on March 24, 2014, of a certain railway bureau as test data, its causative factors and relation table are obtained (see Table 8). The causative factor chain is G-2→J-5→J-9→J-1, that is, unreasonable regulations → insufficient braking force of the track shoe → turnout derailment → vehicle derailment.

[0139] Table 8. Causative Factors and Relationships of the Event on March 24, 2014

[0140] Causative factors Representative information Word segmentation index TF-IDF G-2 Regulations --> Formulation 106 0.324778862 J-5 Anti-slip iron shoes 23 0.132846472 J-9 turnout --> squeeze 77 0.193440405 J-1 All wheels --> derailed 84 0.344119656

[0141] Secondly, embodiments of this application provide a railway accident factor identification and relation extraction system, employing the above-mentioned railway accident factor identification and relation extraction method, such as... Figure 3 As shown, the railway accident factor identification and relationship extraction system includes:

[0142] Historical accident text dataset acquisition module 10: Based on the paragraph and chapter layout features of historical railway accident documents from various sources, regular expressions are defined to extract text paragraphs describing events from historical accident documents for data cleaning and to obtain a valid historical accident text dataset.

[0143] Structured Feature Extraction Module 20: After segmenting the effective historical accident text dataset into words based on a pre-constructed railway domain vocabulary, part-of-speech tagging and named entity recognition are performed based on the word segmentation. After generating dependency syntax structure from the part-of-speech tagging results, structured feature extraction and structured storage of the historical accident text dataset are performed.

[0144] Causative Factor Identification and Classification Module 30: The structured features of historical accident text data are labeled to construct a knowledge base. Based on the knowledge base, a causative factor identification method containing multiple causative factors is constructed to identify accident causative factors and classify them to obtain a multi-level accident causative factor set.

[0145] Causative Factor Relationship Extraction Module 40: Based on the multi-layer accident causative factor set, sort and combine them to construct an accident causative factor chain, thereby realizing the extraction of accident causative factor relationships.

[0146] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the railway accident factor identification and relationship extraction method described above.

[0147] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, wherein the program, when executed by a processor, implements the railway accident factor identification and relation extraction method described above.

[0148] In addition, combined Figure 1 The railway accident factor identification and relationship extraction method described in this application embodiment can be implemented by computer equipment. Figure 4 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of this application.

[0149] The computer device may include a processor 81 and a memory 82 storing computer program instructions.

[0150] Specifically, the processor 81 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0151] The memory 82 may include a mass storage device for data or instructions. For example, and not limitingly, the memory 82 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, the memory 82 may include removable or non-removable (or fixed) media. Where appropriate, the memory 82 may be internal or external to a data processing device. In a particular embodiment, the memory 82 is non-volatile memory. In a particular embodiment, the memory 82 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.

[0152] The memory 82 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 81.

[0153] The processor 81 reads and executes computer program instructions stored in the memory 82 to implement any of the railway accident factor identification and relationship extraction methods in the above embodiments.

[0154] In some embodiments, the computer device may further include a communication interface 83 and a bus 80. For example, Figure 4 As shown, the processor 81, memory 82, and communication interface 83 are connected through bus 80 and complete communication with each other.

[0155] The communication interface 83 is used to enable communication between the various modules, devices, units, and / or equipment in the embodiments of this application. The communication port 83 can also enable data communication with other components such as external devices, image / data acquisition devices, databases, external storage, and image / data processing workstations.

[0156] Bus 80 includes hardware, software, or both, that couples components of a computer device together. Bus 80 includes, but is not limited to, at least one of the following: data bus, address bus, control bus, expansion bus, and local bus. For example, and not as a limitation, bus 80 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 80 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0157] Compared with existing technologies, the beneficial effects of the present invention are as follows: Based on the theoretical method of text mining, the present invention makes full use of historical railway accident documents, and transforms unstructured data into structured data through document conversion, data cleaning, feature extraction and other technologies. Furthermore, it uses text mining methods to extract key features, identify the causative factors of railway accidents and their relationships, which can effectively guide on-site personnel to prevent key accidents, provide data support for early warning and remediation of actual accident and fault risks, and thus ensure the safe operation of railways.

[0158] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for identifying railway accident factors and extracting relationships, characterized in that, The method for identifying and extracting railway accident factors includes: Steps for obtaining historical accident text dataset: Based on the paragraph and chapter layout features of railway historical accident documents from various sources, define regular expressions, extract text paragraphs describing events from the historical accident documents, perform data cleaning, and obtain a valid historical accident text dataset. The structured feature extraction steps are as follows: After the effective historical accident text dataset is segmented into words based on a pre-constructed railway domain vocabulary, part-of-speech tagging and named entity recognition are performed based on the word segmentation. After the part-of-speech tagging results are generated using dependency syntax, the structured features of the historical accident text dataset are extracted and stored in a structured manner. The steps for identifying and classifying causative factors are as follows: the structured features of the historical accident text data are labeled to construct a knowledge base; based on the knowledge base, an accident causative factor identification method containing multiple layers of causative factors is constructed to identify and classify the accident causative factors to obtain a multi-layer accident causative factor set. The steps for extracting causal factor relationships are as follows: Based on the multi-layered accident causal factor set, the accident causal factor chain is constructed by sorting and combining the factors to extract the accident causal factor relationships.

2. The railway accident factor identification and relationship extraction method according to claim 1, characterized in that, The step of obtaining the historical accident text dataset further includes: Text format conversion steps: For historical railway accident documents containing multiple formats and sources, a unified file encoding method is used to convert the file types to obtain files with recognizable formats; The effective text acquisition steps are as follows: Analyze the paragraph and chapter layout features of the recognizable format files, design regular expressions, filter and clean irrelevant railway historical accident texts, and obtain a railway historical accident text dataset composed of effective railway historical accident texts.

3. The railway accident factor identification and relationship extraction method according to claim 1, characterized in that, The structured feature extraction further includes: Word segmentation steps: For historical railway accident texts, sentences are divided according to punctuation marks to obtain a sentence set. The railway domain vocabulary includes: a railway domain stop word list and a railway domain personalized word segmentation list. A pre-trained word segmentation model is used, combined with the railway domain stop word list and the railway system personalized word segmentation list, to perform word segmentation on the sentence set to obtain the word segmentation results. Part-of-speech tagging steps: Based on the word segmentation results of the railway historical accident text sentences, a pre-trained part-of-speech tagging model is used to perform part-of-speech tagging; Named entity recognition steps: Based on the word segmentation results of the railway historical accident text sentences, a pre-trained named entity recognition model is used to perform named entity recognition; Part-of-speech tagging step: Based on the part-of-speech tagging results obtained in the part-of-speech tagging step, the railway historical accident text is segmented and categorized by word, and the preset valid word classes are retained. The categorized results are then concatenated to form a new text and new corpus corresponding to the railway historical accident text. The steps for calculating the term frequency-inverse document frequency value are as follows: For the new corpus, the term frequency-inverse document frequency (TF-IDF) values ​​of the preset effective word classes after screening in each of the railway historical accident texts are calculated, and the representative score of the words in different documents is calculated. Supplementary optimization steps: By selecting representative words with high scores from each document, repeat the word segmentation step up to the word frequency-reverse file frequency value calculation step to supplement and optimize the railway field stop word list and the railway field personalized word segmentation list; Dependency syntactic structure recognition steps: Based on the word segmentation results and part-of-speech tagging results of the railway historical accident text sentences as feature inputs, a pre-trained dependency syntactic analysis model is used to recognize dependency syntactic structures. Multiple dependency syntactic structures can be obtained from the sentences, forming tuple features and storing them in a structured manner.

4. The railway accident factor identification and relationship extraction method according to claim 1, characterized in that, The causative factor identification and classification step further includes: Steps for constructing the causative factor system: Construct a multi-layered causative factor system based on human-machine-environment-management, and form a mapping relationship between causative factor classification labels and descriptions; Knowledge base construction steps: Annotate the structured features of the text data to construct a multi-layered knowledge base, which includes: mapping relationships, keyword vocabulary, and dependency structure table; Steps for obtaining the causative factor set: After completing data cleaning and text feature calculation for the railway historical accident text dataset, the accident causative factors are classified and identified based on the knowledge base by constructing a causative factor identification system with a multi-layer structure, generating a candidate set of causative factors, and obtaining the causative factor set corresponding to the test railway historical accident text dataset by fusing and deduplicating the causative factors.

5. The railway accident factor identification and relationship extraction method according to claim 4, characterized in that, The knowledge base construction steps further include: Foreign language words with preset tags are obtained from the text of historical railway accidents. The foreign language words correspond to the railway accident level in the text. Based on the foreign language words, specific tags are obtained in the multi-layer causal factor system. The mapping relationship is constructed in the knowledge base. The TF-IDF values ​​of the words in the railway historical accident text dataset are sorted from high to low. Words with parts of speech of verbs, nouns and entity words in the preset effective word classes are selected as candidates. Experts annotate the keywords, and the keywords are accumulated in the keyword list corresponding to the tag categories in the knowledge base. For the subject-predicate relation, verb-object relation, and adverbial-head structure in the dependency syntax structure of the railway historical accident text, the representative score of the dependency syntax structure is used to sort them from high to low. Experts mark the key dependency structures and accumulate the key dependency structures in the dependency structure table corresponding to the tag category in the knowledge base.

6. The railway accident factor identification and relationship extraction method according to claim 4, characterized in that, The step of obtaining the causative factor set further includes: Foreign language words with preset tag types in the test railway historical accident text are mapped according to the mapping relationship to obtain the first layer of causative factors; The keyword list is used to retrieve the verb and noun results after word segmentation of the test railway historical accident text to obtain the hit keyword sequence. The hit keyword sequence is then filtered for adjacent occurrences of the same causative factor, and the filtered hit keyword sequence is the second layer of causative factors. The dependency structure table is used to retrieve the dependency syntax structure of the test railway historical accident text to obtain the sequence of hit dependency syntax structures. Adjacent hit structures of the same cause factor are filtered to obtain the filtered sequence of hit dependency structures as the third-level cause factors. The candidate set of cause factors includes: first-level cause factors, second-level cause factors and third-level cause factors.

7. The railway accident factor identification and relationship extraction method according to claim 6, characterized in that, The step of extracting causative factor relationships further includes: The keyword and dependency syntax structure hit sequences are deduplicated and fused. The fusion process combines the writing logic of railway historical accident texts and the inherent relationship of the multi-layered causal factor system of the human-machine-environment-management system. Factors belonging to the preset category are sorted in advance to obtain a composite sequence. Based on the composite sequence, the factors are classified according to the factor categories of the first-level causative factors as follows: If the causative factors corresponding to the first-level causative factors are included in the composite sequence, then the causative factor relationship of the accident is the composite sequence, thus obtaining the causative factor relationship chain; If the causative factors corresponding to the first-level causative factors are human or management-related factors and are not included in the composite sequence, then the causative factors are combined in the composite sequence and ordered as the first position after the management-related causative factors to obtain the causative factor relationship chain. If the causative factor corresponding to the first-level causative factor does not belong to the human or management category and is not included in the composite sequence, then it is sorted as the last item in the composite sequence to obtain the causative factor relationship chain.

8. A railway accident factor identification and relation extraction system, employing the railway accident factor identification and relation extraction method as described in any one of claims 1-7, characterized in that, The railway accident factor identification and relationship extraction system includes: Historical accident text dataset acquisition module: Based on the paragraph and chapter layout features of historical railway accident documents from various sources, regular expressions are defined to extract text paragraphs describing events from the historical accident documents, perform data cleaning, and obtain a valid historical accident text dataset. Structured feature extraction module: After segmenting the effective historical accident text dataset into sentences based on a pre-constructed railway domain vocabulary, part-of-speech tagging and named entity recognition are performed based on the segmented words. After generating dependency syntax structure from the part-of-speech tagging results, the structured features of the historical accident text dataset are extracted and structured storage is performed. Causative Factor Identification and Classification Module: The structured features of the historical accident text data are labeled to construct a knowledge base. Based on the knowledge base, a causative factor identification method containing multiple causative factors is constructed to identify accident causative factors and classify them to obtain a multi-layer accident causative factor set. Causative Factor Relationship Extraction Module: Based on the multi-layered accident causative factor set, sort and combine them to construct an accident causative factor chain, thereby realizing the extraction of accident causative factor relationships.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the railway accident factor identification and relation extraction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the railway accident factor identification and relation extraction method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Railway accident cause analysis method based on word extension LDA

    CN110472225A

  • Railway accident root cause identification system and identification method based on unstructured data

    CN113609302A