A method for extracting accident-causing keywords based on text mining
By constructing an accident information text library and word library, combining word network analysis and verb network analysis, the word weight is calculated, the problem of ignoring the influence of sentences in the existing technology is solved, and the accurate extraction of keywords caused by accidents is achieved, and the scientificity and comprehensiveness of the analysis is improved.
Patent Information
- Application Number
- CN202211483146.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-24
AI Technical Summary
In the prior art, the keyword extraction method based on text mining is based only on the frequency of words appearing, ignoring the importance of keywords in sentences, resulting in insufficient scientificity and accuracy of the cause analysis due to accidents.
By constructing an accident information text library and word library, combining word network analysis and sentence network analysis, the weight of words is calculated, and the impact of words in sentences is comprehensively considered, so as to achieve the extraction of keywords caused by accidents.
It improves the scientificity and accuracy of keyword text mining caused by accidents, ensures that key information is not lost, and improves the comprehensiveness of accident analysis.
Smart Images

Figure CN115759064B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text mining, and particularly relates to a method for extracting accident cause keywords based on text mining. Background Technique
[0002] To scientifically prevent and effectively contain the occurrence of accidents, it is necessary to clarify the causes of accidents. Risk identification is an effective safety management method and has currently been widely applied in fields such as energy extraction, construction, transportation, and chemical manufacturing. Currently, commonly used accident cause analysis methods include the fault tree method, neural network method, Bayesian network method, text mining method, analytic hierarchy process, etc. Among them, the text mining method focuses on the risks that have occurred, takes text data as the basis, and mines and interprets risk factors. An accident report is an objective description of the accident occurrence process and causes, and is a comprehensive review of the accident. Therefore, text mining based on accident reports can significantly improve the scientific nature of risk identification.
[0003] For accident cause analysis based on text mining, an important step is to extract accident cause keywords. Keywords can reflect the core idea of the text content, are important and meaningful information in the text content, and are the basis for accident cause analysis. Currently, keyword extraction based on text mining mostly takes the frequency of word appearance as the statistical basis. The greater the keyword frequency, the more important this keyword is considered. However, according to the accident cause theory, the occurrence of an accident is caused by the interaction of a series of reasons. The complete accident cause is to further carry out clustering analysis or network analysis on the basis of keyword analysis, and its essence is the combined action of word networks and sentence networks. Therefore, only using the frequency of a single keyword appearance as the importance judgment basis and ignoring the influence of the sentence where the keyword is located on the keyword importance will lead to the loss of key information for accident causes and affect the scientific nature and accuracy of the conclusions of accident cause text mining. Summary of the Invention
[0004] To solve the above problems, the purpose of the present invention is to provide a method for extracting accident cause keywords based on text mining, which realizes the extraction of accident cause keywords by integrating the analysis results of word networks and sentence networks.
[0005] To achieve the above purpose, the method for extracting accident cause keywords based on text mining provided by the present invention includes the following steps carried out in sequence:
[0006] Step 1) Collect accident investigation reports related to the subject to be analyzed and construct an accident information text library;
[0007] Step 2) Combine the characteristics of the subject field to construct a professional word library, a filtering word library, and a merged word library for the subject field, and form an accident cause word segmentation rule;
[0008] Step 3) using the accident cause word segmentation rules formed in step 2), the accident information text library constructed in step 1) is segmented, thereby converting the accident information in the accident information text library into structured data that is easy for computers to recognize, and extracting all the accident cause words contained in the accident information text library in combination with the subject area to form an accident cause word library;
[0009] Step 4) Based on the accident information text library and the accident cause word library, the weights of the accident cause words in the accident cause word library are calculated to obtain the importance ranking of the accident cause words.
[0010] The specific method is as follows:
[0011] Step 4.1) Based on the content of the accident information text library, construct the sentence vector Sentences of the accident information text library n ×1 , assuming that the accident information text library contains n sentences, the expression of the sentence vector is:
[0012] Sentences n×1 =[S1S2…S n ] T
[0013] Among them, S i (i=1,2,…n) is the i-th sentence in the accident information text library;
[0014] Step 4.2) Construct the word vector of the accident cause word library based on the content of the accident cause word library Assume that the accident cause word library contains m accident cause words, then the word vector of the accident cause word library is The expression is:
[0015]
[0016] Among them, W j (j=1,2,…m) is the jth accident cause word in the accident cause word database;
[0017] Step 4.3) Analyze the accident information text library sentence vectors Sentences obtained in step 4.1) n×1 The i-th statement S in i , get the i-th statement S i The word vector of the accident cause word library obtained in step 4.2) is included The jth accident cause word W j The number of (j=1,2,…m) is recorded as z ij (j=1,2,…,m), then the number z ij (j=1,2,…,m) constitutes the i-th statement Si Accident-causing Word Frequency Matrix The expression is:
[0018]
[0019] Step 4.4) Repeat Step 4.3), calculate the number of times each of the m accident-causing words in the accident-causing word library appears in the n sentences of the accident information text library, and obtain the m accident-causing word frequency matrices Then sum all the accident-causing word frequency matrices to obtain the total frequency of the jth accident-causing word W j appearing in the accident information text library, that is, obtain the comprehensive accident-causing word frequency matrix L m×m , and the expression is:
[0020]
[0021] Among them, L j (j = 1, 2,..., m) is the jth word W of the accident-causing word library after summation calculation j (j = 1, 2,..., m) the frequency of appearance in the accident information text library;
[0022] Step 4.5) Based on the results obtained in Step 4.4, normalize the comprehensive accident-causing word frequency matrix L m×m to obtain the normalized matrix The expression is:
[0023]
[0024] Among them, tr(L m×m ) represents the trace or trace number of the comprehensive accident-causing word frequency matrix L m×m , which is expressed as the sum of the elements on the main diagonal of the comprehensive accident-causing word frequency matrix L m×m ;
[0025] Step 4.6) Based on the normalized matrix obtained in Step 4.5 The accident-causing word frequency matrix obtained in Step 4.3 and the accident-causing word library word vector obtained in Step 4.2 Then the sentence vector Sentences of the accident information text library n×1 The ith sentence S i The corrected word vector of the accident-causing word library contained in The expression is:
[0026]
[0027] Define the accident cause word vector correction coefficient matrix as:
[0028]
[0029] Then the corrected word vector of the accident cause word library The expression is simplified to:
[0030]
[0031] Then the trace of the accident cause word vector correction coefficient matrix constitutes the accident information text library sentence weight vector of the accident information text library Accident information text library sentence weight vector The expression is:
[0032]
[0033] Step 4.7) Based on the results obtained in Step 4.6), normalize the accident information text library sentence weight vector to obtain the normalized weight vector The expression is:
[0034]
[0035] Then the sentence vector Sentences of the accident information text library n×1 The correction coefficient K i of the i-th sentence S i The expression is:
[0036] K i = 1 + Y i g
[0037] where Y i g represents the normalized weight value of the i-th sentence in the accident information text library;
[0038] Furthermore, the correction frequency L' of the j-th accident cause word W j (j = 1, 2, … m) in the accident cause word vector j The expression is:
[0039]
[0040] Furthermore, based on the above correction frequency L' j , the accident cause word frequency correction matrix The expression is:
[0041]
[0042] L' m That is, it represents the optimized accident cause word weight. The larger the weight value, the more important the word is as an accident cause. Then, this word is used as the keyword of the accident cause.
[0043] The accident cause keyword extraction method based on text mining provided by the present invention has the following beneficial effects: The present invention makes full use of the key information of accident texts, introduces a sentence weight vector based on word frequency, obtains a word frequency correction matrix, and realizes the comprehensive calculation of the dual influence of the word network and the sentence network on the word weight, improving the scientificity and accuracy of accident cause keyword text mining. Description of the Drawings
[0044] Figure 1 It is a method flow chart of the accident cause keyword extraction method based on text mining provided by the present invention. Detailed Embodiments
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0046] The accident cause keyword extraction method based on text mining provided by the present invention is applicable to the analysis of any accident investigation report. The following takes general aviation accidents as an example to illustrate the method of the present invention. The accident cause keyword extraction method based on text mining provided in this embodiment includes the following steps carried out in sequence:
[0047] Step 1: Collect general aviation accident investigation reports and construct a general aviation accident information text library.
[0048] Step 2: Combine the general aviation operation management characteristics and practices to construct a general aviation professional word library, a filtering word library, and a general aviation combined word library, and form a general aviation accident cause word segmentation rule
[0049] Step 3: Use the general aviation accident cause word segmentation rule formed in Step 2 to perform word segmentation on the general aviation accident information text library constructed in Step 1), convert the accident information in the general aviation accident information text library into structured data that is easy for a computer to recognize, and extract all the general aviation accident cause words included in the general aviation accident information text library to form a general aviation accident cause word library.
[0050] Step 4 comprehensively considers the influence of the sentence network information in the general aviation accident information text library and the word information in the general aviation accident cause word library, calculates the weights of the general aviation accident cause words in the general aviation accident cause word library, and obtains the importance ranking of the general aviation accident cause words.
[0051] The specific method is as follows:
[0052] (1) Based on the content of the general aviation accident information text library, extract the sentence vector Sentences of the general aviation accident information text library n×1 = [S1 S2 S3] T , where n = 3, that is, the general aviation accident information text library contains 3 sentences in total. S1 = "The flight crew violated relevant regulations and still carried out an approach and landing when the visibility was lower than the company's minimum descent altitude, resulting in the aircraft crashing into the ground"; S2 = "The flight crew violated the relevant regulations of the Civil Aviation Administration. When the aircraft entered the radiation fog and did not see the airport runway and did not establish the visual reference necessary for landing, they still crossed the minimum descent altitude to carry out the landing"; S3 = "Before the aircraft crashed into the ground, there was a radio altitude voice prompt, and when the airport runway was not seen, the flight crew still did not take the go-around measure and continued to blindly carry out the landing, resulting in the aircraft crashing into the ground."
[0053] (2) Construct the word vector of the general aviation accident cause word library Among them, m = 14, that is, the general aviation accident information text library contains 14 general aviation accident cause words in total, which are: W1 = "flight crew", W2 = "violate", W3 = "approach", W4 = "minimum descent altitude", W5 = "regulation", W6 = "radiation fog", W7 = "landing", W8 = "runway", W9 = "visual reference", W 10 = "go-around", W 11 = "did not see", W 12 = "did not take", W 13 = "blindly", W 14 = "crash into the ground".
[0054] (3) Analyze the i-th sentence S 3×1 in the sentence vector Sentences of the general aviation accident information text library i , and obtain the number of the j-th accident cause word W i in the word vector of the general aviation accident cause word library contained in the i-th sentence S j (j = 1, 2,... 14). Denote the number of the j-th accident cause word W i contained in the i-th sentence S j as z ij (i = 1, 2, 3; j = 1, 2,..., 14), then the quantity zij (i = 1, 2, 3; j = 1, 2, …, 14) constitutes the i-th statement S i of the accident cause word frequency matrix It can be calculated that:
[0055]
[0056]
[0057]
[0058] Among them, diag(A) represents the diagonal square matrix formed by vector A, and the values of the diagonal square matrix are the frequencies of the j-th accident cause word W j (j = 1, 2, … 14) in the i-th statement S i (i = 1, 2, 3).
[0059] (4) Summing all the general aviation accident cause word frequency matrices can obtain the total frequency of the j-th accident cause word W j appearing in the general aviation accident information text library, that is, the general aviation accident cause word frequency comprehensive matrix L m×m , it can be calculated that:
[0060] L 14×14 = diag([3, 2, 1, 2, 2, 1, 3, 2, 1, 1, 2, 1, 1, 2])
[0061] Among them, the values of the diagonal square matrix are the sum of the frequencies of the corresponding j-th accident cause word W j (j = 1, 2, … 14) in the i-th statement S i (i = 1, 2, 3).
[0062] (5) Based on the calculation results of the previous step, normalize the general aviation accident cause word frequency comprehensive matrix L m×m to calculate the normalized matrix:
[0063]
[0064] Then the modified word vector of the general aviation accident cause word library composed of the j-th accident cause word W i (i = 1, 2, 3) contained in the i-th statement S j (j = 1, 2, … 14) can be expressed as:
[0065]
[0066]
[0067]
[0068] Among them, the word vector correction coefficient matrices of the general aviation accident cause word library are respectively:
[0069]
[0070]
[0071]
[0072] (6) Then the statement weight vector of the accident information text library of the general aviation accident information text library The calculation result is
[0073]
[0074] That is, the weights of statements S1, S2, and S3 in the general aviation accident information text library are 0.625, 0.75, and 0.625 respectively.
[0075] (7) Normalize the statement weight vector of the general aviation accident information text library Perform normalization calculation, and the normalized weight vector can be obtained:
[0076]
[0077] That is, the weights of statements S1, S2, and S3 in the general aviation accident information text library after normalization are 0.4651, 0.5581, and 0.4651 respectively.
[0078] (8) According to the following formula, the word vector of the general aviation accident cause word library The j-th word W j (j = 1, 2,... m) of the correction frequency L' j ,
[0079]
[0080] Among them, K i = 1 + Y i g (i = 1, 2, 3)
[0081] Calculated, the general aviation accident cause word frequency correction matrix is as follows:
[0082]
[0083] The values of the above diagonal square matrix respectively represent the words W1 to W in the general aviation accident cause word library 14The weight values. That is, the weight of W1 "flight crew" is 1.4884, the weight of W2 "violation" is 1.0233, the weight of W3 "approach" is 0.4651, the weight of W4 "minimum descent altitude" is 1.0233, the weight of W5 "regulation" is 1.0233, the weight of W6 "radiation fog" is 0.5581, the weight of W7 "landing" is 1.4884, the weight of W8 "runway" is 1.0233, the weight of W9 "visual reference" is 0.5581, W 10 The weight of "go-around" is 0.4651, W 11 The weight of "not seen" is 1.0233, W 12 The weight of "not taken" is 0.4651, W 13 The weight of "blind" is 0.4651, W 14 The weight of "ground collision" is 0.9302. It shows that based on the text analysis of this general aviation accident information text library, the most important accident-causing words for general aviation accidents are: "flight crew", "landing", and the second most important accident-causing words are: "violation", "minimum descent altitude", "runway", "not seen", and so on.
Claims
1. A method for extracting accident-causing keywords based on text mining, characterized in that: The accident cause keyword extraction method based on text mining includes the following steps carried out in sequence: Step 1) Collect accident investigation reports related to the subject to be analyzed, and construct an accident information text library; Step 2) Combine the characteristics of the subject field to construct a professional word library, a filtering word library, and a combined word library for the subject field, and form accident cause word segmentation rules; Step 3) Use the accident cause word segmentation rules formed in Step 2) to perform word segmentation on the accident information text library constructed in Step 1), thereby converting the accident information in the accident information text library into structured data that is easy for a computer to recognize, and combining the subject field, extract all accident cause words contained in the accident information text library to form an accident cause word library; Step 4) Based on the accident information text library and the accident cause word library, calculate the weights of the accident cause words in the accident cause word library to obtain the importance ranking of the accident cause words; Sentences of the accident information text library n×1 The i-th sentence S i The corrected word vector of the accident cause word library contained in The expression of is: Among them, represents the accident cause word frequency matrix, represents the normalization matrix, represents the word vector of the accident cause word library; Define the accident cause word vector correction coefficient matrix as follows: Then the expression of the accident causation word library correction word vector is simplified to: The accident-causing word vector correction coefficient matrix of the accident-causing word library The trace of which constitutes the accident information text library statement weight vector of the accident information text library Accident information text library statement weight vector The expression of which is: Based on the obtained results, the statement weight vector of the accident information text library is normalized to obtain a normalized weight vector The expression is as follows: Then the sentence vector of the accident information text library, Sentences n×1 The i-th sentence S i The correction coefficient K i The expression is: K i = 1 + Y i g Among them, Y i g represents the normalized weight value of the i-th sentence in the accident information text library; Furthermore, the word vector of the accident cause word library The j-th accident cause word W j (j = 1, 2, … m) of the corrected frequency L′ j The expression is as follows: Among them, z ij (j = 1, 2, …, m) represents the j-th accident-causing word W i in the accident-causing word vector library contained in the i-th statement S ; and j (j = 1, 2, …, m) is the quantity of Furthermore, based on the above-mentioned corrected frequency L′ j , the expression of the accident cause word frequency correction matrix is as follows: L′ m That is, it represents the optimized accident causation word weight. The larger the weight value, the more important the word is as an accident causation. Then, this word is used as the keyword for accident causation.
2. The accident cause keyword extraction method based on text mining according to claim 1, characterized in that: In Step 4), the specific method for calculating the weights of the accident cause words in the accident cause word library based on the accident information text library and the accident cause word library to obtain the importance ranking of the accident cause words is as follows: Step 4.1) Based on the content of the accident information text library, construct the sentence vector Sentences of the accident information text library n×1 , assuming that the accident information text library contains n sentences, the expression of the sentence vector is as follows: Sentences n×1 = [S1 S2 … S n T Among them, S i (i = 1, 2, … n) is the i-th statement in the accident information text library; Step 4.2) Based on the content of the accident causation word library, construct the word vectors of the accident causation word library Suppose the accident causation word library contains m accident causation words, then the word vectors of the accident causation word library The expression is as follows: Among them, W j (j = 1, 2, …, m) is the j-th accident cause word in the accident cause word library; Step 4.3) Analyze the sentence vectors Sentences in the accident information text library obtained in Step 4.1) n×1 the i-th sentence S in i , and obtain the i-th sentence S i containing the accident cause word vectors in the accident cause word library obtained in Step 4.2) the j-th accident cause word W j (j = 1, 2, … m) and denote the quantity as z ij (j = 1, 2, …, m), then from this quantity z ij (j = 1, 2, …, m) constitutes the accident cause word frequency matrix of the i-th sentence S i The expression is: Expression: Step 4.4) Repeat Step 4.3), calculate the number of times each of the m accident-causing terms in the accident-causing term library appears in the n sentences of the accident information text library, and obtain an m×n accident-causing term frequency matrix Then, for all the accident-causing term frequency matrices Sum them up to obtain the total frequency of occurrence of the jth accident-causing term W j in the accident information text library, that is, obtain the comprehensive accident-causing term frequency matrix L m×m , and the expression is: Among them, L j (j = 1, 2, …, m) is the j-th word W in the accident cause word library after summation calculation j (j = 1, 2, …, m) the frequency of occurrence in the accident information text library; Step 4.5) Based on the results obtained in Step 4.4, normalize the accident-causing word frequency comprehensive matrix L m×m to obtain a normalized matrix The expression is as follows: Among them, tr(L m×m ) represents the trace or the number of traces of the comprehensive matrix L of accident-causing word frequencies m×m , which is expressed as the sum of the elements on the main diagonal of the comprehensive matrix L of accident-causing word frequencies m×m .
Citation Information
Patent Citations
Safety production accident analysis method and device based on text mining, electronic equipment and storage medium
CN112364627A
Text rhetorical sentence generation method, apparatus and device, and readable storage medium
WO2021139229A1