A method for constructing aviation safety incident logic graph based on causality
Through the multi-feature fusion of aviation safety accident causality extraction model CEASM and deep learning model to calculate text similarity, an efficient and accurate aviation safety story rational map was constructed, which solved the shortcomings of extracting knowledge and constructing a rational map in the existing technology, and realized the causality analysis and prediction and early warning of aviation safety accidents.
Patent Information
- Application Number
- CN202311323330.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-12
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2043-10-12
AI Technical Summary
The prior art is difficult to extract knowledge efficiently and accurately from aviation safety reports and build a rational map, and cannot fully capture semantic information in text.
The causal relationship extraction model of aviation safety accidents with multi-feature fusion is adopted, and text features are extracted and captured through the Transformer architecture. The text similarity is calculated using deep learning models to construct a rational map of aviation safety events.
It improves the accuracy and efficiency of causal relationship extraction of aviation safety accidents, reveals the causal logic evolution laws of accidents, provides scientific prediction, prevention and early warning basis, and improves the safety of aviation operations.
Smart Images

Figure CN117291264B_ABST
Abstract
Description
Background Art
[0002] With the continuous development and progress of the civil aviation industry, air traffic has become more and more popular. But at the same time, aviation safety incidents occur from time to time. Aviation accidents have cast a serious shadow on the development of the civil aviation industry. Once a safety accident occurs, it will inevitably bring heavy casualties and economic losses. Therefore, it is very important to prevent and avoid aviation accidents, and analyzing the causes of accidents can identify safety hazards in advance and effectively ensure flight safety.
[0003] Aviation safety accident reports collect a large number of unsafe incidents involving aircraft operations from aviation practitioners. For example, the National Transportation Safety Board (NTSB) of the United States collects a large amount of data related to aviation safety accident reports, including detailed descriptions of the accident process, personnel information, aircraft status, weather conditions, and possible causes of the accident. These safety reports are the best source of information for identifying aviation safety hazards and explaining the causes of aviation accidents, providing the industry with a data foundation for mining past accident information and learning lessons. However, due to the diversity and complexity of the causes of aviation safety accidents, the analysis of aviation safety reports faces new challenges.
[0004] At present, the research on the causes of aviation safety accidents mainly uses data statistics to analyze the distribution of accident causes, or uses Bayesian networks and complex networks to analyze the probability distribution of the impact of known causes on accident results. These methods cannot effectively explore the complete process of the evolution of aviation safety accidents. In recent years, artificial intelligence technology with event logic graph as the core has emerged, which has provided help to solve the above problems. Event logic graph (ELG) has a strong semantic representation ability and can well represent the complex relationship between events in aviation safety accidents. It can automatically discover valuable information from massive unstructured data and analyze the ins and outs, causes and consequences of events. It is of great significance for event risk warning, decision-making assistance and other activities in many fields such as politics, economy and military. Using event logic graph technology to conduct qualitative and quantitative analysis of aviation safety accidents can explore the causal relationship of aviation safety accidents, clarify the development context of accidents, reveal the laws of accident development, assist in the investigation of the causes of accidents, and provide decision support for future accident prediction, prevention and warning, so as to avoid secondary injuries as much as possible and reduce casualties and losses caused by accidents. However, existing circumstance graph construction methods generally use word2vec technology to generate text word vectors, which is difficult to fully capture the semantic information in the text; in addition, the current research on the construction method of aviation safety incident circumstance graph still lacks an efficient and accurate method to extract knowledge and information from aviation safety reports and construct circumstance graphs. Summary of the Invention
[0005] To solve the above problems, that is, the prior art is difficult to fully capture the semantic information in the text, resulting in poor accuracy and efficiency in extracting knowledge and information from aviation safety reports and constructing an event logic graph, the present invention provides a method for constructing an aviation safety accident event logic graph based on causal relationships, and the method includes:
[0006] Step S100, obtaining historical data of aviation safety accident reports and preprocessing; the preprocessing includes data conversion, text extraction, cleaning, word segmentation, sentence segmentation, and annotation;
[0007] Step S200, extracting causal relationships of various safety accident events in the preprocessed historical data of aviation safety accident reports through a trained aviation safety accident causal relationship extraction model CEASM with multi-feature fusion, and constructing a text dataset of aviation safety accident causal relationship event pairs;
[0008] Step S300, calculating the comprehensive similarity feature vectors between the causal relationship event pair texts in the text dataset of aviation safety accident causal relationship event pairs and classifying them, and then constructing a set of similar accident causal texts;
[0009] Step S400, obtaining the central vector of each set of similar accident causal texts, calculating the similarity scores between each causal relationship event pair text in the same set of similar accident causal texts and the central vector, and taking the causal relationship event pair text with the highest similarity score as the standard event description; normalizing the text dataset of aviation safety accident causal relationship event pairs according to the standard event description;
[0010] Step S500, creating a node for each event in the normalized text dataset of aviation safety accident causal relationship event pairs and connecting them in series. After connecting them in series, calculate the conditional transition probability between the nodes and assign the corresponding directed edges, thereby obtaining an aviation safety accident event logic graph.
[0011] In some preferred embodiments, step S100 further includes:
[0012] Step S110, obtaining historical data of aviation safety accident reports in Microsoft Access format, and converting the historical data of aviation safety accident reports in Microsoft Access format into CSV format;
[0013] Step S120, extracting the accident narrative text from the historical data of aviation safety accident reports in CSV format; deleting the basic data in the accident narrative text where the accident cause narrative is repeated or empty, and obtaining the first text; the accident narrative text includes accident number, full accident narrative, long accident narrative sentence, and accident cause narrative;
[0014] Step S130: Clean, segment, and clause-segment the accident cause description of the first text to obtain a second accident description text;
[0015] Step S140: Based on the second accident description text, use the BIO annotation method for annotation, and use the obtained BIO-annotated accident description text data as the second text.
[0016] In some preferred embodiments, the step S200 further includes:
[0017] Step S210: For the second text, use the CCNN module and the BERT module to generate character-level feature vectors and word-level feature vectors;
[0018] Step S220: Perform feature fusion on the character-level feature vectors and the word-level feature vectors. After fusion, use the Transformer encoding layer and the Transformer decoding layer to perform encoding and decoding processing on the fused feature vectors to further extract and capture the feature relationships;
[0019] Input the feature vectors after encoding and decoding processing into the CRF layer for causal relationship extraction to obtain an aviation safety accident causal relationship event pair text, the content of which includes a cause event description text, a result event description text, and other description texts;
[0020] Step S230: Divide the second text into a training set, a validation set, and a test set according to a set ratio, and complete the training of the aviation safety accident causal relationship extraction model CEASM with multi-feature fusion;
[0021] Step S240: Perform causal relationship extraction on the first text through the trained aviation safety accident causal relationship extraction model CEASM with multi-feature fusion to construct an aviation safety accident causal relationship event pair text dataset.
[0022] In some preferred embodiments, the step S300 further includes:
[0023] Step S310: Vectorize the causal relationship event pair text in the aviation safety accident causal relationship event pair text dataset through the BERT model;
[0024] Step S320: Obtain the labeled event text similarity labels, and train the BERT pre-trained model to obtain a trained BERT text similarity calculation model; the BERT pre-trained model includes BERT-Large Uncased, BERT-Large Cased, BERT-Base Uncased, and BERT-Base Cased;
[0025] Step S330, for each trained BERT text similarity calculation model, obtaining the BERT text similarity calculation model with the best effect according to preset model evaluation indicators; the model evaluation indicators include: precision, accuracy, recall and F1 value;
[0026] Step S340, using the BERT text similarity calculation model with the best effect and each similarity calculation method to respectively calculate and merge the text similarity feature vectors of the vectorized causal event pair texts to obtain a comprehensive similarity feature vector; the similarity calculation methods include cosine similarity, Jaccard similarity, Manhattan distance method and Euclidean distance method;
[0027] Step S350, classifying the comprehensive similarity feature vectors by a multi-layer perceptron, and after classification, classifying aviation safety accident causal relationship events with a similarity of 1 into the same similar accident causal text set.
[0028] In some preferred embodiments, the step S400 further comprises:
[0029] Step S410, obtaining the central vector of the text in each similar accident causal text set; and calculating the similarity calculation result between each causal event pair text in the same similar accident causal text set and the corresponding central vector, and then obtaining the accident causal event pair text with the highest similarity as the standard event description;
[0030] Step S420: Replace and update the causal relationship event pair texts in the aviation safety accident causal relationship event pair text data set one by one using the standard event description.
[0031] In some preferred embodiments, the step S500 further comprises:
[0032] Step S510, for each event in the normalized aviation safety accident causal relationship event pair text data set, create a corresponding node; the node includes an accident cause node and an accident result node;
[0033] Connecting directed edges from the accident cause node to the corresponding accident result node, and merging identical nodes, ultimately connecting all aviation safety accident causal relationship event pairs in series;
[0034] Step S520, calculate the conditional transition probabilities between nodes and assign them to the corresponding directed edges to obtain an aviation safety incident logic graph.
[0035] Beneficial effects of the present invention:
[0036] The present invention adopts natural language processing methods such as causal relationship extraction, event fusion normalization, and event relationship quantification to construct a causal graph that describes the causal relationship and occurrence probability between various abnormal events in aviation safety accidents, revealing the causal logic evolution law and pattern of aviation safety accidents, which helps airlines to judge the various evolution directions and corresponding results of aviation safety accidents, improves the accuracy and efficiency of the constructed causal graph, and can provide a more reliable scientific basis for the prediction, prevention and early warning of aviation safety accidents; specifically:
[0037] (1) The present invention proposes a multi-feature fusion aviation safety accident causal relationship extraction model, which uses the CCNN module and the BERT module to extract text character-level and word-level features respectively and fuses them, and uses the Transformer architecture to further extract and capture feature relationships, which greatly reduces the loss of semantic information features and can effectively improve the accuracy of aviation safety incident causal relationship extraction;
[0038] (2) The present invention proposes a comprehensive text similarity calculation model, which uses the deep learning text feature extraction model BERT to extract text feature vectors, and integrates the traditional text similarity calculation method with the BERT similarity calculation model, thereby overcoming the defect that the traditional text similarity calculation method does not take into account the features such as semantics, grammar, and word order, and can greatly improve the calculation accuracy of event text similarity;
[0039] (3) The present invention uses a deep learning model to quickly and accurately extract causal event pairs from aviation safety accident texts, and then uses a text similarity calculation model to calculate the similarity of various event texts in the causal event pairs. The text is normalized based on the similarity results combined with the professional context in the aviation field, and finally the causal relationships of aviation safety events are connected in series, and the probability of causal relationships between events is calculated to form an aviation safety incident logic map. This method effectively reduces the difficulty of aviation accident report analysis, improves the efficiency of accident report analysis, and makes up for the shortcomings of existing aviation safety accident evolution analysis. In addition, the construction of an aviation safety incident logic map provides a more reliable scientific guidance for accident prediction, prevention, and early warning, which helps to improve the safety of aviation operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0041] Figure 1 It is a flow chart of a method for constructing an aviation safety incident logic map based on causal relationships according to an embodiment of the present invention;
[0042] Figure 2is a schematic diagram of the data format of a CSV data set extracted according to an embodiment of the present invention;
[0043] Figure 3 It is a schematic diagram of a BIO annotation method of aviation safety accident text data according to an embodiment of the present invention;
[0044] Figure 4 is a schematic diagram of a CEASM model architecture of an embodiment of the present invention;
[0045] Figure 5 is a schematic diagram of a causal relationship extraction result of an embodiment of the present invention;
[0046] Figure 6 It is a structural diagram of a text similarity comprehensive calculation model of an embodiment of the present invention;
[0047] Figure 7 It is a schematic diagram of a normalization process of a textual representation of an aviation safety incident according to an embodiment of the present invention;
[0048] Figure 8 is a flowchart of aviation safety incident text merging according to an embodiment of the present invention;
[0049] Figure 9 It is a schematic diagram of a runway incursion event event spectrum result according to an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0051] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0052] In order to more clearly illustrate the method for constructing an aviation safety incident logic map based on causal relationships of the present invention, each step in an embodiment of the present invention is described in detail below with reference to the figures.
[0053] In the first embodiment of the present invention, according to the characteristics of aviation safety accident texts, a causality extraction model for aviation safety accidents based on multi-feature fusion, CEASM (Causality Extraction Model for Aviation Safety Accident Based on Multi-Feature Fusion), is proposed. The CCNN module and the BERT module are used to extract text character-level and word-level features respectively and fuse them. The Transformer architecture is used to further extract and capture feature relationships, minimizing the loss of semantic features to a great extent. Finally, the CRF layer is used for causality extraction. Secondly, the present invention proposes a comprehensive text similarity calculation model. The advanced deep learning text feature extraction model BERT is used to calculate text similarity feature vectors, which are fused with the similarity feature vectors obtained by four traditional text similarity calculation methods. Then, the MLP classifier is used to judge whether the causality events are similar or not. Finally, according to the similarity calculation results, the aviation accident causality event pairs are normalized, the accident causality chain is concatenated, and the transition probability between the causality event pairs is calculated to form an event logic graph of aviation safety accidents, providing a reference for aviation safety accident early warning. As Figure 1 shown, the specific steps are as follows:
[0054] Step S100, obtain and preprocess the historical data of aviation safety accident reports; the preprocessing includes data conversion, text extraction, cleaning, word segmentation, sentence splitting and annotation;
[0055] In this embodiment, the specific process of preprocessing the historical data of aviation safety accident reports is as follows:
[0056] Step S110, obtain the historical data of aviation safety accident reports in the Microsoft Access format, and convert the historical data of aviation safety accident reports in the Microsoft Access format into the CSV format;
[0057] In this embodiment, it is preferably to extract the historical data of aviation safety accident reports in the Microsoft Access format from the data of aviation accident reports included in the NTSB data.
[0058] Step S120, extract the accident narrative text from the historical data of aviation safety accident reports in the CSV format; delete the basic data in the accident narrative text where the accident cause narrative is repeated or empty to obtain the first text; the accident narrative text includes the accident number, the full text of the accident narrative, the long accident narrative sentence and the accident cause narrative;
[0059] In this embodiment, the accident narrative text (narratives) in the historical data of aviation safety accident reports in CSV format is extracted. The accident narrative text includes fields such as accident number (ev_id), full text of accident narrative (narr_accp), long accident narrative sentence (narr_accf), and accident cause narrative (narr_cause), as Figure 2 shown; the duplicate or empty basic data in the accident cause narrative in the accident narrative text is deleted, and finally 9,954 pieces of data are obtained as the research data for subsequent causal relationship extraction;
[0060] Step S130: Clean, segment words, and split sentences for the accident cause narrative of the first text to obtain a second accident narrative text;
[0061] In this embodiment, cleaning includes converting abbreviations to full forms, removing punctuation marks and line breaks in words, etc. Since English words are separated by spaces, different words can be distinguished by spaces to complete the word segmentation process. Based on punctuation marks and other text features (such as abbreviations, capitalization, etc.) to determine the boundaries of sentences, preferably use the sent_tokenize function of nltk to split the text into individual sentences;
[0062] Step S140: Based on the second accident narrative text, use the BIO annotation method for annotation, and use the obtained BIO-annotated accident narrative text data as the second text;
[0063] In this embodiment, preferably use the BIO annotation method for manual annotation, which can convert the causal relationship extraction problem into a label classification task. The BIO three-digit annotation (B-begin, I-inside, O-outside): B-X represents the beginning of entity X, I-X represents the end of entity X, and O represents not belonging to any type; among them, "B-X" means that the segment where this element is located belongs to type X and this element is at the beginning of this segment, "I-X" means that the segment where this element is located belongs to type X and this element is in the middle position of this segment, and "O" means not belonging to any type; for the accident cause narrative text of aviation safety accidents, the BIO annotation method is as Figure 3 shown.
[0064] Step S200: Use the trained aviation safety accident causal relationship extraction model CEASM with multi-feature fusion to extract causal relationships for various safety accident events in the preprocessed historical data of aviation safety accident reports, and construct an aviation safety accident causal relationship event pair text data set;
[0065] In this embodiment, a multi-feature fusion aviation safety accident causality extraction model CEASM is used to extract the causality of various safety accident events in the aviation accident report, and a dataset of aviation safety accident causality event pairs is constructed. The structure of the multi-feature fusion aviation safety accident causality extraction model CEASM is as shown in Figure 4 follows. The input layer is the aviation safety accident description text. The CCNN module (character CNN encoder, that is, the CCNN model) and the BERT module (that is, the BERT model) respectively extract character-level and word-level features of the text. After fusing the word-level and character-level feature vectors, the Transformer encoder layer and decoder layer are used to further extract and capture the feature relationships. Finally, the CRF layer is used to perform the causality extraction work, and five vocabulary categories of B-C, I-C, O, B-E, and I-E are output. Among them, B-C and I-C are combined into the cause event description text, B-E and I-E are combined into the result event description text, and O is other description text. The specific process is as follows:
[0066] Step S210, for the second text, use the CCNN module and the BERT module to generate character-level feature vectors and word-level feature vectors;
[0067] Step S220, fuse the character-level feature vector and the word-level feature vector. After fusion, use the Transformer encoding layer and the Transformer decoding layer to perform encoding and decoding processing on the fused feature vector to further extract and capture the feature relationship;
[0068] Input the feature vector after encoding and decoding processing into the CRF layer for causality extraction to obtain the aviation safety accident causality event pair text, the content of which includes the cause event description text, the result event description text, and other description text;
[0069] Step S230, divide the second text into the training set, the validation set, and the test set according to a set ratio to complete the training of the multi-feature fusion aviation safety accident causality extraction model CEASM;
[0070] In this embodiment, based on the second text, it is preferably divided into a training set, a validation set, and a test set according to a ratio of 7:2:1. The training set data is used to train the feature fusion-based aviation safety accident causality extraction model CEASM. After each training cycle, the model parameters are saved and the recognition effect of the model is verified on the validation set. A total of 200 training cycles are performed to iterate the model parameters. The model parameters with the highest F1-score in the validation set are selected as the optimal parameters of the feature fusion-based aviation safety accident causality extraction model CEASM, and the performance of the model is tested with the test set to complete the training of the feature fusion-based aviation safety accident causality extraction model CEASM. The specific training process is as follows:
[0071] The second text is divided into a training set, a validation set, and a test set according to a set ratio to obtain training samples;
[0072] The training samples are input into the pre-constructed multi-feature fusion-based aviation safety accident causality extraction model CEASM to obtain aviation safety accident causality event pair texts;
[0073] Based on the aviation safety accident causality event pair texts and their corresponding labels, the loss value is calculated, and then the model parameters are updated;
[0074] The multi-feature fusion-based aviation safety accident causality extraction model CEASM is trained in a loop until the trained multi-feature fusion-based aviation safety accident causality extraction model CEASM is obtained.
[0075] The feature fusion-based aviation safety accident causality extraction model CEASM is trained using the negative log-likelihood loss function, and its calculation formula is as follows:
[0076] Loss(C)=-∑ (x,y) logP(y|x;W) (1)
[0077] In the formula, C is the annotated corpus; x=(x1,…,x m ) and y are the strings and labels in the corpus C respectively; W is the weight matrix of the probabilities of each character output by the model;
[0078] To evaluate the performance of the aviation safety accident causality extraction model CEASM, the evaluation metrics used include: accuracy (Accuracy, Acc), precision (Precision, P), recall (Recall, R), and F1-score (F);
[0079] Accuracy refers to the percentage of correctly predicted results in the total samples. Precision means the probability that the actual positive samples are among all the samples predicted as positive. Recall means the probability that the samples predicted as positive are among the actual positive samples. F1-score is the weighted average of accuracy and recall. The calculation formulas for the above indicators are as follows:
[0080]
[0081]
[0082]
[0083]
[0084] Wherein, TP: the number of samples with the true class being positive and the predicted class being positive; FP: the number of samples with the true class being negative but the predicted class being positive; FN: the number of samples with the true class being positive but the predicted class being negative; TN: the number of samples with the true class being negative and the predicted class being negative.
[0085] Step S240, perform causal relationship extraction on the first text through the trained multi-feature fusion aviation safety accident causal relationship extraction model CEASM to construct an aviation safety accident causal relationship event pair text data set, as Figure 5 shown.
[0086] Step S300, calculate the comprehensive similarity feature vectors between the causal relationship event pair texts in the aviation safety accident causal relationship event pair text data set and classify them, thereby constructing a set of similar accident causal texts;
[0087] In this embodiment, the specific process of constructing a set of similar accident causal texts is as follows:
[0088] Step S310, vectorize the causal relationship event pair texts in the aviation safety accident causal relationship event pair text data set through the BERT model;
[0089] Step S320, obtain the labeled event text similarity labels, train the BERT pre-trained model to obtain the trained BERT text similarity calculation model; the BERT pre-trained model includes BERT-Large Uncased, BERT-Large Cased, BERT-Base Uncased, and BERT-Base Cased;
[0090] In this embodiment, event text similarity tags are obtained as experimental data, and four scales of BERT pre-trained models are trained respectively in combination with the event text similarity tags. In each group of model training, a training set, a validation set, and a test set are obtained according to the ratio of 7:2:1. The training set is used to train the BERT model, and the trained BERT model is used to predict the validation set and the test set respectively to obtain the trained BERT text similarity calculation model. In the present invention, four types are selected, namely BERT-Large Uncased, BERT-Large Cased, BERT-Base Uncased, and BERT-Base Cased;
[0091] Step S330, for each trained BERT text similarity calculation model, obtain the BERT text similarity calculation model with the best effect according to the preset model evaluation indexes; the model evaluation indexes include: precision, accuracy, recall rate, and F1 value;
[0092] Step S340, calculate and fuse the text similarity feature vectors of the vectorized causal relationship event pair texts through the BERT text similarity calculation model with the best effect and each similarity calculation method respectively to obtain a comprehensive similarity feature vector; the similarity calculation methods include cosine similarity, Jaccard similarity, Manhattan distance method, and Euclidean distance method;
[0093] In this embodiment, the architecture of the text similarity comprehensive calculation model is as Figure 6 shown;
[0094] Set S i and S j to be the text vector representations of events i and j respectively, and set the traditional similarity calculation methods (also preferably four types in the present invention) as follows:
[0095] Cosine similarity: Use the cosine value of the angle between two vectors in the vector space to measure the difference between two events. The closer the cosine value is to 1, the more similar the two events are. Cosine similarity is a commonly used similarity measurement method, and the calculation result is accurate, suitable for processing short texts. The formula is as follows:
[0096]
[0097] Jaccard (Jaccard) similarity: The Jaccard coefficient is equal to the ratio of the intersection to the union of two sets, that is, the intersection divided by the union, the words that two events have in common divided by all the words of the two events, and can be used to measure the correlation between two sets. The advantage of the similarity function based on the Jaccard coefficient is that the set intersection processing is independent of the order of the strings in the set. Therefore, the order of the strings has basically no influence on the similarity measurement result. The calculation formula is as follows:
[0098]
[0099] Manhattan distance: It refers to the actual "block" distance between two points in a plane space, and the calculation formula is as follows:
[0100] Manhat(S i ,S j )=|S i -S j | (8)
[0101] Euclidean distance: It is usually used to measure the absolute distance between two points in a multi-dimensional space, and the calculation formula is as follows:
[0102]
[0103] Step S350, classify the comprehensive similarity feature vector through a multi-layer perceptron. After classification, the aviation safety accident causal relationship events with a similarity of 1 are classified into the same similar accident causal text set;
[0104] In this embodiment, the comprehensive similarity feature vector is classified into 0-1 through a multi-layer perceptron (MLP). 0 represents that the event texts are not similar, and 1 represents that the event texts are similar. After classification, a similar accident causal text set is constructed based on the comprehensive similarity feature vectors with a similarity of 1;
[0105] The classifier training of the multi-layer perceptron uses the BCE loss function, and its calculation formula is as follows:
[0106] Loss=-(1-y)log(1-x)-ylog(x) (10)
[0107] In the formula, y is the data label, and x is the model prediction result.
[0108] Step S400, obtain the central vector of each similar accident causal text set, calculate the similarity score between each causal relationship event pair text in the same similar accident causal text set and the central vector, and use the causal relationship event pair text with the highest similarity score as the standard event description; normalize the aviation safety accident causal relationship event pair text data set according to the standard event description;
[0109] In this embodiment, referring to the idea of "cluster center", the standard event description of the causal events in the same set is extracted, and the aviation safety accident causal relationship event pair text data set is normalized. The specific process is as Figure 7 shown, and specifically includes:
[0110] Step S410, obtaining the central vector of the text in each similar accident causal text set; and calculating the similarity calculation result between each causal event pair text in the same similar accident causal text set and the corresponding central vector, and then obtaining the accident causal event pair text with the highest similarity as the standard event description;
[0111] In this embodiment, the cosine similarity calculation result is preferably used as the similarity calculation result between each causal relationship event pair text and the central vector in the similar accident causal text set;
[0112] Step S420: Replace and update the causal relationship event pair texts in the aviation safety accident causal relationship event pair text data set one by one using the standard event description.
[0113] Step S500, for each event in the normalized aviation safety accident causal relationship event text data set, a node is created and connected in series, after which the conditional transition probabilities between the nodes are calculated and the corresponding directed edges are assigned, thereby obtaining an aviation safety accident causal relationship graph;
[0114] In this embodiment, the specific process of constructing the aviation safety incident logic map is as follows:
[0115] Step S510, for each event in the normalized aviation safety accident causal relationship event pair text data set, create a corresponding node; the node includes an accident cause node and an accident result node; connect a directed edge from the accident cause node to the corresponding accident result node, merge the same nodes, and finally connect all aviation safety accident causal relationship event pairs in series;
[0116] In this embodiment, all aviation safety accident causal events are finally connected in series through the above method. There are three main ways to merge the same nodes, such as Figure 8 As shown:
[0117] If the cause event A in the causal relationship event pair text I and J is expressed the same, then the cause event is merged; if the result event C in the causal relationship event pair text I and J is expressed the same, then the result event is merged; if the result event B in the causal relationship event pair text I and the cause event B in the causal relationship event pair text J are expressed the same, then the result event in I and the cause event in J are merged;
[0118] Step S520, calculating the conditional transition probabilities between nodes and assigning them to the corresponding directed edges, and obtaining an aviation safety incident logic graph;
[0119] In this embodiment, it is preferred to use the Markov Chain algorithm to obtain the conditional transition probability. Markov chain (MC) is a random process that undergoes transitions from one state to another in the state space. At each step of the Markov chain, the system can change from one state to another state or maintain the current state according to the probability distribution. The change of state is called transition, and the probability associated with different state changes is called transition probability.
[0120] The present invention uses the Markov Chain algorithm to calculate the event transition probability to represent the evolution probability of the event causal relationship. The calculation formula is as follows:
[0121]
[0122] count(E i ,E j ) is E i Occurs and E j The number of events that also occurred, ∑ k count(E i ,E k ) is E i The total number of all possible events that could have occurred in the event of an occurrence;
[0123] Finally, Neo4j is used to visualize the event graph, such as Figure 9 The figure shows the event graph of runway incursion. The circular nodes in the figure are cause and result events. The direction of the arrow indicates the direction from the cause event to the result event. The value on the arrow indicates the transition probability of the cause event evolving into the result event.
[0124] The aviation safety incident logic map reveals the causal logic evolution laws and patterns of aviation safety accidents, which helps airlines judge the various evolution directions and corresponding results of aviation safety accidents, provides a reference for aviation safety accident warning, and provides a more reliable scientific basis for aviation safety accident prediction, prevention and warning.
[0125] Although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art can understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of the present invention.
[0126] The aviation safety incident logic diagram construction system based on causality according to the second embodiment of the present invention includes:
[0127] A data preprocessing module configured to obtain and preprocess historical data of aviation safety accident reports; the preprocessing includes data conversion, text extraction, cleaning, word segmentation, sentence segmentation and annotation;
[0128] The relationship extraction module is configured to extract the causal relationships of various safety accident events in the pre-processed aviation safety accident report historical data through the trained multi-feature fusion aviation safety accident causal relationship extraction model CEASM, and construct an aviation safety accident causal relationship event pair text dataset;
[0129] A set construction module is configured to calculate and classify the comprehensive similarity feature vectors between the causal relationship event pair texts in the aviation safety accident causal relationship event pair text data set, and then construct a similar accident causal relationship text set;
[0130] A normalization processing module is configured to obtain a central vector of each similar accident causal text set, calculate a similarity score between each causal event pair text in the same similar accident causal text set and the central vector, and use the causal event pair text with the highest similarity score as a standard event description; and normalize the aviation safety accident causal event pair text data set according to the standard event description;
[0131] The graph construction module is configured to create a node for each event in the normalized aviation safety accident causal relationship event text data set and connect them in series. After the connection, the conditional transition probability between the nodes is calculated and the corresponding directed edges are assigned, thereby obtaining an aviation safety accident causal graph.
[0132] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process and related instructions of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0133] It should be noted that the aviation safety incident logic diagram construction system based on causal relationships provided in the above embodiment is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the modules or steps, and are not regarded as improper limitations on the present invention.
[0134] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A method for constructing an aviation safety incident logic map based on causality, characterized in that: The method comprises: Step S100, obtaining and preprocessing historical data of aviation safety accident reports; the preprocessing includes data conversion, text extraction, cleaning, word segmentation, sentence segmentation and annotation; Step S200, extracting causal relationships of various types of safety accident events in the pre-processed aviation safety accident report historical data through the trained multi-feature fusion aviation safety accident causal relationship extraction model CEASM, and constructing an aviation safety accident causal relationship event pair text dataset; Step S300, calculating and classifying the comprehensive similarity feature vectors between the causal relationship event pair texts in the aviation safety accident causal relationship event pair text data set, and then constructing a similar accident causal relationship text set; Step S400, obtaining the central vector of each similar accident causal text set, calculating the similarity score between each causal event pair text in the same similar accident causal text set and the central vector, taking the causal event pair text with the highest similarity score as the standard event description; and normalizing the aviation safety accident causal event pair text data set according to the standard event description; Step S500, for each event in the normalized aviation safety accident causal relationship event text data set, a node is created and connected in series, after which the conditional transition probabilities between the nodes are calculated and the corresponding directed edges are assigned, thereby obtaining an aviation safety accident causal graph.
2. The method for constructing an aviation safety incident logic map based on causality according to claim 1 is characterized in that: Step S100 further includes: Step S110, obtaining historical data of aviation safety accident reports in Microsoft Access format, and converting the historical data of aviation safety accident reports in Microsoft Access format into CSV format; Step S120, extracting the accident narrative text from the aviation safety accident report history data in CSV format; deleting the basic data of the accident cause description in the accident narrative text that is repeated or empty, to obtain a first text; the accident narrative text includes the accident number, the full text of the accident narrative, the long accident narrative sentence and the accident cause description; Step S130, cleaning, word segmentation and sentence segmentation of the accident cause description in the first text to obtain a second accident description text; Step S140: annotating the second accident narrative text using a BIO annotation method, and using the obtained BIO annotated accident narrative text data as the second text.
3. The method for constructing an aviation safety incident logic map based on causal relationships according to claim 2 is characterized in that: The step S200 further comprises: Step S210, for the second text, using the CCNN module and the BERT module to generate a character-level feature vector and a word-level feature vector; Step S220, performing feature fusion on the character-level feature vector and the word-level feature vector, and after fusion, using the Transformer encoding layer and the Transformer decoding layer to perform encoding and decoding processing on the fused feature vector to further extract and capture the feature relationship; The encoded and decoded feature vector is input into the CRF layer to extract the causal relationship, and the aviation safety accident causal relationship event pair text is obtained, which includes the cause event description text, the result event description text and other description texts; Step S230, dividing the second text into a training set, a validation set, and a test set according to a set ratio, and completing the training of the aviation safety accident causal relationship extraction model CEASM with multi-feature fusion; Step S240, extracting causal relationships from the first text using the trained multi-feature fusion aviation safety accident causal relationship extraction model CEASM, and constructing an aviation safety accident causal relationship event pair text dataset.
4. The method for constructing an aviation safety incident logic map based on causality according to claim 1 is characterized in that: The step S300 further comprises: Step S310, vectorizing the causal relationship event pair text in the aviation safety accident causal relationship event pair text dataset by using a BERT model; Step S320, obtaining the annotated event text similarity label, training the BERT pre-trained model, and obtaining a trained BERT text similarity calculation model; the BERT pre-trained model includes BERT-Large Uncased, BERT-LargeCased, BERT-Base Uncased, and BERT-Base Cased; Step S330, for each trained BERT text similarity calculation model, obtaining the BERT text similarity calculation model with the best effect according to preset model evaluation indicators; the model evaluation indicators include: precision, accuracy, recall and F1 value; Step S340, using the BERT text similarity calculation model with the best effect and each similarity calculation method to respectively calculate and merge the text similarity feature vectors of the vectorized causal event pair texts to obtain a comprehensive similarity feature vector; the similarity calculation methods include cosine similarity, Jaccard similarity, Manhattan distance method and Euclidean distance method; Step S350, classifying the comprehensive similarity feature vectors by a multi-layer perceptron, and after classification, classifying aviation safety accident causal relationship events with a similarity of 1 into the same similar accident causal text set.
5. The method for constructing an aviation safety incident logic map based on causality according to claim 1, characterized in that: The step S400 further comprises: Step S410, obtaining the central vector of the text in each similar accident causal text set; and calculating the similarity calculation result between each causal event pair text in the same similar accident causal text set and the corresponding central vector, and then obtaining the accident causal event pair text with the highest similarity as the standard event description; Step S420: Replace and update the causal relationship event pair texts in the aviation safety accident causal relationship event pair text data set one by one using the standard event description.
6. The method for constructing an aviation safety incident logic map based on causality according to claim 1, characterized in that: The step S500 further comprises: Step S510, for each event in the normalized aviation safety accident causal relationship event pair text data set, create a corresponding node; the node includes an accident cause node and an accident result node; Connecting directed edges from the accident cause node to the corresponding accident result node, and merging identical nodes, ultimately connecting all aviation safety accident causal relationship event pairs in series; Step S520, calculate the conditional transition probabilities between nodes and assign them to the corresponding directed edges to obtain an aviation safety incident logic graph.
Citation Information
Patent Citations
Fault positioning method fusing Bayesian network and performance-fault relation graph
CN116360387A
Aviation field fault analysis affair graph construction method
CN116662571A