Expressway traffic accident cause network construction method based on text mining
By applying text mining technology in highway traffic accident data analysis, the keywords of the cause cause characteristics of accidents are extracted and the cause network is built, which solves the problem of insufficient utilization of unstructured text data in the existing technology, and achieves more accurate accident cause cause analysis and network quality improvement.
Patent Information
- Application Number
- CN202510083898.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology lacks the full utilization of unstructured text data in the analysis of highway traffic accident data, resulting in high-frequency co-occurrence between terms but weak semantic relationships, making it difficult to effectively explore the mechanism of accident cause.
Using text mining based methods, unstructured text data of highway traffic accidents are collected and processed, keywords of accident causes are extracted, and the accident cause causes are built using point mutual information values to build an accident cause network to explore the correlation between different key causes.
This method can further explore the high-interaction mode in the highway traffic accident-induced cause network, clearly display the meaning logic and correlation between causes, and improve the quality and interpretability of the accident-induced cause network.
Smart Images

Figure CN120144776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic safety cause analysis, and in particular to a method for building a highway traffic accident cause network based on text mining. Background Art
[0002] The frequent occurrence of road traffic accidents has become a focus of attention. Highways have the characteristics of various types of vehicles, high vehicle speeds, and closed driving environments, which leads to higher levels of serious traffic accidents on them, with serious injuries and mortality rates far exceeding those on ordinary roads, which have a significant impact on government finances, citizens' lives and health, medical treatment, and social stability. The theory of traffic accident causes points out that the occurrence of accidents is caused by the imbalance of causal molecules such as people-vehicle-road-environment. Therefore, based on accident data mining technology, the mechanism of highway traffic accidents is explored, the characteristic values of accident causes are extracted and identified, and the key information and potential laws hidden in them are revealed, which is of great value for effectively preventing the occurrence of highway traffic accidents or reducing the risk level of accidents.
[0003] At present, research in the field of highway traffic safety has achieved certain results, mainly based on traffic accident big data to explore the causes of various accidents and their mechanisms. However, after sorting out, it was found that there are still some problems in the following research stages, including: 1. In the stage of traffic data preparation, most studies are committed to qualitative or quantitative analysis of the potential content in structured data, and the exploration of abstract unstructured text data is not sufficient; 2. In the stage of traffic data understanding, for highway traffic accident big data, more methods such as data statistical analysis, mining and visualization are used to explore the causes and action paths of accidents. Among them, the visualization part is mostly based on the co-occurrence matrix to build the corresponding semantic network, and the association between terms is accidental.
[0004] In order to make full and effective use of the unstructured text data of highway traffic accidents and identify the causal mechanism characteristics that lead to the severity of each accident, when using text mining technology, more in-depth semantic mining should be carried out on the basis of key feature word extraction, such as mining the highly interactive patterns of accident causes and optimizing the meaning logic between nodes in the semantic network. Summary of the invention
[0005] The present invention provides a method for constructing a highway traffic accident causal network based on text mining. The method makes full use of unstructured text data of highway traffic accidents, constructs highly interactive vocabulary patterns in significant accident causal data, and can better restore the situation when the accident occurred compared to the application of structured data; furthermore, the accident causal network constructed based on point mutual information values can effectively solve the problem of high-frequency co-occurrence between terms but weak semantic relationships, clearly display the meaning logic and association between causes, and improve the quality and interpretability of the highway traffic accident causal network.
[0006] The technical solution of the present invention is as follows: A method for constructing a highway traffic accident cause network based on text mining, comprising the following steps: S1: Collect highway traffic accident investigation reports, perform data selection, conversion and data cleaning, and build an accident text corpus; S2: Target Chinese word segmentation, load merged word list, custom professional dictionary and stop word list, and complete corpus dataization; S3: Taking into account the weight and frequency of text feature words, further dimensionality reduction screening is performed to extract the characteristic keywords of accident causes; S4: Establish an index relationship between text and keyword terms, and then calculate the point mutual information between high-weight causes to mine the highly interactive patterns of key causes of highway traffic accidents; S5: Construct an accident risk cause network based on point mutual information and explore the correlation between different key causes.
[0007] Furthermore, in step S1, the database information exported from the traffic management system of the Ministry of Public Security generally includes two levels of content: on-site situation records and post-investigation. Information irrelevant to the text mining object, such as the alarm receiving point information and responsibility demarcation, is eliminated, and only the time of the accident, road name, accident type, number of injured people, total number of deaths, and initial cause of the accident are retained in the data, so as to reduce the redundancy of text mining and improve the processing speed of the text mining process; furthermore, the abnormal items and non-standard descriptions in the accident records are corrected and converted into a unified file format.
[0008] Further, in step S2, the data is segmented based on the Pyphon and Jieba Chinese word segmentation packages, and the traffic accident text is divided into several valid feature items according to the dictionary. The specific process is as follows; Load the merged vocabulary and normalize the terms with the same meaning but different expressions; Refer to national road traffic accident standard documents and collect and organize vocabulary in the fields of traffic safety, traffic accidents and road transportation, and build a custom professional dictionary for text analysis of the causes of highway traffic accidents; Dynamically update the stop word list according to the word segmentation results to add or delete new words until the final result meets the requirements of structured data.
[0009] Furthermore, after word segmentation of the text on the causes of highway traffic accidents, the word frequency statistics results are obtained and a word cloud diagram of the accident cause characteristics is drawn accordingly. The correlation degree between the font size and the word frequency statistics value is set to visually display the word segmentation results of the cause text.
[0010] Furthermore, in step S3, the specific process of extracting the accident cause characteristic keywords is as follows: S301: Eliminate the word items that are irrelevant or have ambiguous relationships with the accident cause theme shown in the word cloud diagram; S302: Secondarily merge the characteristic word items with the same meaning but not recognized and normalized in the word segmentation results; S303: Apply the improved term frequency-inverse document frequency (TF-IDF) algorithm to normalize the weight formula of the characteristic words, accurately quantify the importance of the word items in the text, and the formula is as follows:
[0011] In the formula: represents the characteristic word in the text the weight value; represents the characteristic word in the text the frequency of occurrence; represents the total number of texts in the corpus; represents the text the characteristic word the number of; represents the number of texts including the characteristic word in the corpus.
[0012] Furthermore, extract the top 38 highway traffic accident cause characteristic keyword items with a weight value ranking (the minimum weight value ≥ 0.01), including 11 traffic elements such as drivers, motor vehicles, highways, leading vehicles, low visibility, meteorological conditions, pedestrians, driving parts, warning signs, emergency lanes, driver's licenses, etc., and 27 situation elements such as excessive fatigue, speeding, etc.
[0013] Furthermore, in step S4, after obtaining the accident key characteristic word items, the frequent co-occurrence vocabulary mining algorithm PMI is used to extract the development mode of highway traffic accident causes. For the word items and the word item in the corpus, their point mutual information can be expressed as and the formula is as follows:
[0014] In the formula: , respectively represent the probabilities of the term , the term appearing in the corpus; represents the joint probability of the term and the term , that is, the probability of co-occurrence; is the conditional probability value, that is, the probability that the term also appears on the premise that the term appears; then represents the probability that the term appears on the premise that the term appears.
[0015] When the term and the term are independent of each other, , when the correlation between the two is stronger, the value is larger, and the calculated PMI value is also larger.
[0016] Furthermore, based on the theoretical basis of the PMI algorithm, the steps to mine highly interactive vocabulary patterns from high-weight accident-causing data are as follows: Based on the extraction results of accident feature terms by the improved TF-IDF algorithm, establish an index relationship between the text and the keyword terms to obtain a two-dimensional table composed of text information and corresponding terms; Run the programming code to calculate the co-occurrence frequency between keyword terms, construct a one-mode co-occurrence matrix, mine the point mutual information between high-weight causes, and calculate the PMI value.
[0017] Set a threshold according to business requirements and data characteristics, and screen out highly interactive key cause terms whose PMI values and co-occurrence frequencies simultaneously meet the threshold conditions.
[0018] Furthermore, in step S5, import the PMI value between accident cause keywords as the weight of the edge into Gephi (network visualization) software, set the resolution value for modular processing, and complete the drawing of the highway traffic accident cause feature term network; furthermore, calibrate relevant parameters and calculate the accident cause network density and standard deviation.
[0019] Furthermore, in step S5, the specific process of exploring the correlation between different key causes is as follows: S501: Use three measurement indicators, namely degree centrality, betweenness centrality, and closeness centrality, to represent the centralization degree of each cause node in the accident cause network, as follows: By calculating the node and each node Obtain the degree centrality by the number of connected edges; Count the nodes by other nodes Take the ratio of the number of shortest paths passed by a node to the total number of shortest paths in the network as the betweenness centrality. The calculation formula is:
[0020] In the formula: and are nodes, represents node to node the number of shortest paths, represents from node to node among the shortest paths from the number of paths passing through node By calculating the shortest distance between nodes to reflect the closeness of node and other nodes to obtain the closeness centrality of the node. The calculation formula is:
[0021] In the formula: represents the shortest distance between node and node .
[0022] S502: Based on UCINET and the Core / Periphery module, load the core-periphery structure of the causes of highway traffic accidents and calculate the average density of the core and periphery regions; S503: Use the iterative correlation convergence method to repeatedly calculate the correlation coefficients between rows and columns in the co-occurrence matrix of accident cause keywords to obtain a relationship matrix composed of two types of coefficients, 1 and -1, and complete the construction of the cohesive subgroup.
[0023] The beneficial effects of the present invention are: The present invention provides a method for constructing a highway traffic accident causation network based on text mining. This method is based on the unstructured text data of highway traffic accidents, uses text mining technology to achieve the formatting processing of accident investigation reports, takes into account the weights and frequencies of text feature words, extracts accident cause keywords, establishes an index relationship between the text and the keyword items, calculates the point mutual information between the feature values, and further mines the highly interactive lexical patterns in the highway traffic accident cause text. Moreover, based on the point mutual information values, an accident cause network analysis schema is constructed to explore the correlation between different key causes. Compared with the existing analysis methods, it can better restore the situation at the time of the accident and effectively solve the problem of weak semantic relationships despite high-frequency co-occurrence between word items. Description of the Drawings
[0024] Figure 1 This is a flow chart of the method for building a highway traffic accident cause network based on text mining in the present invention.
[0025] Figure 2 This is a network diagram of characteristic items of causes of highway traffic accidents in an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The drawings are only for illustrative purposes and cannot be construed as limiting the present invention. To better illustrate the present embodiment, some parts of the drawings may be omitted, enlarged, or reduced, and do not represent the size of the actual product. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the drawings. The positional relationships described in the drawings are only for illustrative purposes and cannot be construed as limiting the present invention.
[0027] Embodiment 1: like Figure 1 As shown, a method for constructing a highway traffic accident cause network based on text mining includes the following steps: S1: Collect highway traffic accident investigation reports, perform data selection, conversion and data cleaning, and build an accident text corpus; S2: Target Chinese word segmentation, load merged word list, custom professional dictionary and stop word list, and complete corpus dataization; S3: Taking into account the weight and frequency of text feature words, further dimensionality reduction screening is performed to extract the characteristic keywords of accident causes; S4: Establish an index relationship between text and keyword terms, and then calculate the point mutual information between high-weight causes to mine the highly interactive patterns of key causes of highway traffic accidents; S5: Construct an accident risk cause network based on point mutual information and explore the correlation between different key causes.
[0028] In this embodiment, the database information exported from the traffic management system of the Ministry of Public Security in step S1 generally includes two levels of content: on-site situation records and post-investigation. Information irrelevant to the text mining object, such as the alarm receiving point information and responsibility demarcation, is eliminated, and only the time of the accident, road name, accident type, number of injured people, total number of deaths, and initial cause of the accident are retained in the data, so as to reduce the redundancy of text mining and improve the processing speed of the text mining process; furthermore, abnormal items and non-standard descriptions in the accident records are corrected in a semi-manual manner, and data joint query, regular expressions, text annotation and other technologies are comprehensively used to convert the unified file format, and the field content required for the research is extracted.
[0029] In this embodiment, step S2 performs word segmentation processing on the data based on Pyphon and Jieba Chinese word segmentation packages, and divides the traffic accident text into several valid feature items according to the dictionary. The specific process is as follows; Load the merged word list and normalize the terms with the same meaning but different expressions, such as driving against the flow and driving against the flow, violating regulations and not following regulations, etc. Refer to national road traffic accident standard documents and collect and organize vocabulary in the fields of traffic safety, traffic accidents and road transportation, and build a custom professional dictionary for text analysis of the causes of highway traffic accidents; The stop word list is dynamically updated according to the word segmentation results to add and delete new words until the final result meets the requirements of structured data.
[0030] Furthermore, after word segmentation of the text of the cause of highway traffic accidents, the word frequency statistics results are obtained and a word cloud diagram of the accident cause characteristics is drawn based on this. The correlation between the font size and the word frequency statistics value is set to 0.5 to intuitively display the word segmentation results of the cause text.
[0031] In this embodiment, the specific process of extracting the accident cause characteristic keywords in step S3 is as follows: S301: Eliminate the terms such as lane, scene, occupation, driving, etc. shown in the word cloud that are irrelevant or vaguely related to the accident cause topic; S302: merging the feature terms with the same meaning but not recognized and normalized in the secondary word segmentation results, such as passenger, passenger, and passenger; S303: Apply the improved term frequency-inverse document frequency (TF-IDF) algorithm to normalize the weight formula of the feature words and accurately quantify the importance of the terms in the text. The formula is as follows:
[0032] Where: Characteristic words In text The weight value in ; Characteristic words In text The frequency of occurrence in Represents the text in the corpus Total number of; Represents text Feature words The number of Indicates that the corpus includes feature words The number of texts.
[0033] Extract the top 38 highway traffic accident cause characteristic keyword items with the highest weight value (minimum weight value ≥ 0.01), including driver T1 , motor vehicle T 3 , highway T 5 , leading vehicle T 8 , low visibility T 14 , meteorological condition T 15 , pedestrian T 16 , driving components T 25 , warning sign T 33 , emergency lane T 35 , driver's license T 37 and 11 other traffic elements, excessive fatigue T 7 , speeding T 10 and 27 other situation elements.
[0034] In this embodiment, after obtaining the accident key feature terms in step S4, the frequent co-occurrence vocabulary mining algorithm PMI is used to extract the development pattern of the causes of highway traffic accidents. For the terms and the term , the point mutual information between them can be expressed as , and the formula is as follows:
[0035] In the formula: , respectively represent the probabilities of the term , the term appearing in the corpus; represents the joint probability of the term and the term , that is, the probability of co-occurrence; is the conditional probability value, that is, the probability that the term appears on the premise that the term appears; then represents the probability that the term appears on the premise that the term appears.
[0036] When the term is independent of the term , , when the correlation between the two is stronger, the value is larger, and the calculated PMI value is also larger.
[0037] Based on the theoretical basis of the PMI algorithm, the steps for mining highly interactive vocabulary patterns from high-weight accident cause data are as follows: Based on the extraction results of accident feature terms using the improved TF-IDF algorithm, establish an index relationship between the text and the keyword terms to obtain a two-dimensional table composed of text information and corresponding terms; Run Python and the Pandas library to calculate the co-occurrence frequency between keyword terms, construct a one-mode co-occurrence matrix, mine the point mutual information between high-weight causes, and calculate the PMI value.
[0038] Set the threshold to co-occurrence frequency > 15 and PMI value > 2.5 according to business requirements and data characteristics, and then screen out the highly interactive key cause terms whose PMI values and co-occurrence frequencies both meet the threshold conditions.
[0039] In this embodiment, in step S5, the PMI value between accident cause keywords is imported into Gephi (network visualization) software as the weight of the edge, and the resolution is set to 1.5. Thus, the modularity value is 0.586, and the result is higher than the threshold of 0.4, indicating that a stable cluster is obtained, and the drawing of the highway traffic accident cause feature item network is completed; furthermore, the calculated graph density result is 0.154, and the standard deviation is 0.351, indicating that the associations between the network nodes composed of the key cause items of highway traffic accidents are close.
[0040] Compare the results of the accident cause feature item network diagram with the content of the "initial accident investigation reasons" part extracted in step S1, and find that the two results are highly consistent, indicating that the proposed method can effectively mine and display the semantic logic and associations between causes, which is in line with the current highway traffic accident cause analysis.
[0041] After constructing the accident cause network, the specific process of exploring the associations between different key causes is as follows: S501: Use three measurement indicators, namely degree centrality, betweenness centrality, and closeness centrality, to represent the centralization degree of each cause node in the accident cause network, as follows: By calculating the number of edges connecting node and each node to obtain the degree centrality; Count the ratio of the number of times node is passed through by other nodes along the shortest path to the total number of shortest paths in the network as the betweenness centrality, and the calculation formula is:
[0042] In the formula: and are nodes, represents the number of shortest paths from node to node , represents the number of paths passing through node from node to node ; By calculating the shortest distance between nodes to reflect the nodes in the network and other nodes closeness, and thus obtain the closeness centrality of the nodes. The calculation formula is:
[0043] In the formula: represents the node and the node the shortest distance between them.
[0044] S502: Based on UCINET and the Core / Periphery module, load the core-periphery structure of the causes of highway traffic accidents. The calculated average density of the core area is 382.33, and the average density of the edge area is 8.387. The core causal factors in the former play a leading role in the whole network; S503: Use the iterative correlation convergence method to repeatedly calculate the correlation coefficients between rows and columns in the co-occurrence matrix of accident cause keywords, and obtain a relationship matrix composed of two types of coefficients, 1 and -1, to complete the construction of the cohesive subgroup.
[0045] The method of the present invention is based on the unstructured text data of highway traffic accidents, uses text mining technology to realize the formatting processing of accident investigation reports, takes into account the weights and frequencies of text feature words, extracts accident cause keywords, and establishes an index relationship between the text and the keyword items, calculates the point mutual information between the eigenvalues, and further mines the highly interactive patterns of the key causes of highway traffic accidents. Moreover, based on the point mutual information value, an accident cause network analysis schema is built to explore the correlation between different key causes. Compared with the existing analysis methods, it can better restore the situation at the time of the accident and effectively solve the problem of weak semantic relationships despite high-frequency co-occurrence between terms.
[0046] The method of the present invention can improve the quality and interpretability of the highway traffic accident cause network, clearly display the meaning logic and correlation between the causes, effectively restore the situation at the time of the accident, and is of great significance for preventing traffic accidents or reducing losses in the highway scenario.
[0047] Example 2: The data source used in this example is the highway traffic accident investigation reports in Hubei Province from 2006 to 2012, with a total of 2,886 records. Each record contains two levels of content: on-site situation description and post-accident investigation.
[0048] S1: Perform data selection, transformation, and data cleaning. Extract the content of the "initial accident investigation reasons" part from the accident investigation reports and number them. After formatting the data, use it as the corpus for accident text mining; S2: Load the Jieba component in Python, import the custom professional dictionary, stop word list, and combined word list, complete the word segmentation process of the text on the causes of highway traffic accidents, obtain the word frequency statistics results, and draw a word cloud diagram of the accident cause characteristics based on this. It can be found that the causes of highway traffic accidents mostly originate from aspects such as fatigue driving (word frequency is 558), lane violation (word frequency is 553), insufficient vehicle distance (word frequency is 445), speeding (word frequency is 314), weather conditions (word frequency is 141), etc.; S3: Apply the improved TF-IDF algorithm to extract the top 38 keyword items of the characteristics of the causes of highway traffic accidents with the weight value ranking (the minimum weight value ≥ 0.01), including the driver T 1 , motor vehicle T 3 , highway T 5 , leading vehicle T 8 , low visibility T 14 , meteorological conditions T 15 , pedestrian T 16 , driving components T 25 , warning sign T 33 , emergency lane T 35 , driver's license T 37 and other 11 traffic elements, excessive fatigue T 7 , speeding T 10 and other 27 situation elements. The weight values decrease in sequence according to T 1 to T 38 's order, as shown in Table 1.
[0049] Table 1 Keywords of the causes of highway traffic accidents Among them, speeding T 10 means speeding less than 50% on the highway, and exceeding the speed limit T 21 means that the motor vehicle travels more than 50% of the speed limit; overtaking T 31 means four situations such as overtaking when the leading vehicle overtakes, overtaking on the highway ramp, overtaking when the leading vehicle turns left, and overtaking on the highway not in accordance with regulations. Since the term of overtaking on the right appears more frequently and belongs to the common occurrence form of overtaking behavior, it is listed separately as T 32 .
[0050] After extracting the accident key feature term list, use the frequent co-occurrence vocabulary mining algorithm to calculate the PMI values and co-occurrence frequencies among the 38 keyword items, construct a one-mode co-word matrix of the key causes of highway traffic accidents, extract the key cause term pairs with co-occurrence frequency > 15 and PMI value > 2.5. They are more likely to become important elements in forming the accident development pattern. Furthermore, complete the design of the highly interactive mode of the key causes of highway traffic accidents.
[0051] S5: Import the PMI values between the accident cause keywords as the weights of the edges into Gephi (network visualization) software, set the resolution value to 1.5 for modular processing, and adjust the sizes of the nodes and their corresponding labels based on the weighted degree. Further, use the Fruchterman Reingold graph layout algorithm to complete the drawing of the highway traffic accident cause feature item network, as Figure 2 shown; The cause social network connects each accident cause feature item to form a system. The size of the node reflects its importance in the network, and the thickness, color depth, and number of connection lines between nodes reflect the degree of connection tightness between accident cause feature items. By observing the accident cause feature item network diagram, it can be found that terms such as {failure to comply with regulations, driving, leading vehicle, safety distance}, {excessive fatigue, continuing to drive}, {highway, speeding}, {low visibility, meteorological conditions}, {obstructing, safe driving} interact highly with each other and are likely to form a development model of accident causes.
[0052] In contrast, by statistically classifying the highway traffic accident investigation reports exported from the Ministry of Public Security's traffic management system, it can be obtained that the most common causes of cases are "failure to maintain a necessary safe distance from the leading vehicle in the same lane in accordance with regulations", "motor vehicle driving in violation of the specified lane", "driver continuing to drive while overly fatigued", "driving on the highway in violation of regulations under low visibility meteorological conditions", "speeding less than 50% on the highway", etc., which are attributed to factors such as fatigue driving, lane violations, insufficient vehicle distance, speeding, and weather conditions. Among them, the driver's driving state factor plays a dominant role, which is highly consistent with the results of the accident cause feature item network diagram, indicating that the proposed method can effectively mine and display the semantic logic and associations between causes, and improve the quality and interpretability of the highway traffic accident cause network.
[0053] Example 3: This example is similar to Example 2. The difference is that this example can perform centrality analysis on the highway traffic accident cause feature network diagram, and use three measurement indicators, namely degree centrality, betweenness centrality, and closeness centrality, to represent the centralization degree of each cause node in the accident cause network. The results are shown in Table 2; Table 2 Centrality analysis of highway traffic accident cause network The centrality analysis in this embodiment, as a main aspect of social network analysis, is a way to further quantitatively analyze the importance of nodes. According to the degree centrality index, the hidden danger behaviors of drivers (fatigue driving, violations, insufficient safety distance) and unstable factors (low visibility) are the most common risk causes leading to highway traffic accidents. The betweenness centrality of illegal parking ranks higher compared to the degree centrality and closeness centrality rankings, indicating that this feature term has a strong mediating effect. Many other accident cause items will be associated through this node, thereby triggering different forms of traffic accidents. Therefore, taking corresponding measures to inhibit the role of this causal node will be beneficial to interrupting the potential connections between multiple risk factors and effectively preventing the occurrence of traffic accidents in the highway scenario.
[0054] Obviously, the above embodiments of the present invention are only examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for constructing a highway traffic accident cause network based on text mining, characterized in that: The following steps are involved: S1: Collect highway traffic accident investigation reports, perform data selection, conversion and data cleaning, and build an accident text corpus; S2: Target Chinese word segmentation, load merged word list, custom professional dictionary and stop word list, and complete corpus dataization; S3: Taking into account the weight and frequency of text feature words, further dimensionality reduction screening is performed to extract the characteristic keywords of accident causes; S4: Establish an index relationship between text and keyword terms, and then calculate the point mutual information between high-weight causes to mine the highly interactive patterns of key causes of highway traffic accidents; S5: Construct an accident risk cause network based on point mutual information and explore the correlation between different key causes.
2. The method for constructing a highway traffic accident cause network according to claim 1, characterized in that: In step S1, the database information exported from the traffic management system of the Ministry of Public Security generally includes two levels of content: on-site situation records and post-incident investigations. Information irrelevant to the text mining object, such as the alarm receiving station information and responsibility demarcation, is removed, and only the accident time, road name, accident type, number of injured, total number of deaths, and initial cause of the accident are retained in the data, so as to reduce the redundancy of text mining and improve the processing speed of the text mining process; Furthermore, correct abnormal items and non-standard descriptions in accident records and convert them into a unified file format.
3. The method for constructing a highway traffic accident cause network according to claim 2, characterized in that: In step S2, the data is segmented based on Pyphon and Jieba Chinese word segmentation packages, and the traffic accident text is divided into several valid feature items according to the dictionary. The specific process is as follows; Load the merged vocabulary and normalize the terms with the same meaning but different expressions; Refer to national road traffic accident standard documents and collect and organize vocabulary in the fields of traffic safety, traffic accidents and road transportation, and build a custom professional dictionary for text analysis of the causes of highway traffic accidents; The stop word list is dynamically updated according to the word segmentation results to add and delete new words until the final result meets the requirements of structured data.
4. The method for constructing a highway traffic accident cause network according to claim 3, characterized in that: After word segmentation of the text of the causes of highway traffic accidents, the word frequency statistics results are obtained and a word cloud diagram of the accident cause characteristics is drawn based on this. The correlation between the font size and the word frequency statistics value is set to intuitively display the word segmentation results of the cause text.
5. The method for constructing a highway traffic accident cause network according to claim 4, characterized in that: In step S3, the specific process of extracting the accident cause characteristic keywords is as follows: S301: Eliminate the terms in the word cloud that are irrelevant or vaguely related to the topic of accident causes; S302: secondary merging of feature terms in the word segmentation results that have the same meaning but have not been identified and normalized; S303: Apply the improved term frequency-inverse document frequency (TF-IDF) algorithm to normalize the weight formula of the feature words and accurately quantify the importance of the terms in the text. The formula is as follows: Where: Characteristic words In text The weight value in ; Characteristic words In text The frequency of occurrence in Represents the text in the corpus Total number of; Represents text Feature words The number of Indicates that the corpus includes feature words The number of texts.
6. The method for constructing a highway traffic accident cause network according to claim 5, characterized in that: The top 38 highway traffic accident cause characteristic keyword items with the highest weight values were extracted (minimum weight value ≥ 0.01), including 11 traffic elements such as driver, motor vehicle, highway, preceding vehicle, low visibility, weather conditions, pedestrians, driving parts, warning signs, emergency lane, and driver's license, and 27 event elements such as excessive fatigue and speeding.
7. The method for constructing a highway traffic accident cause network according to claim 6, characterized in that: In step S4, after obtaining the key feature terms of the accident, the frequent co-occurrence vocabulary mining algorithm PMI is used to extract the cause development pattern of highway traffic accidents. and terms , the point mutual information between them can be expressed as , the formula is as follows: Where: , Respectively represent terms , terms The probability of occurrence in the corpus; Representing terms and terms The joint probability of , that is, the probability of simultaneous occurrence; is the conditional probability value, that is, in the term Precondition The probability of also appearing; then Indicated in the term Precondition The probability of occurrence, When the term With terms When independent of each other, , the stronger the correlation between the two, The larger the value of , the larger the calculated PMI value.
8. The method for constructing a highway traffic accident cause network according to claim 7, characterized in that: Based on the theoretical foundation of the PMI algorithm, the steps for mining highly interactive vocabulary patterns from high-weight accident cause data are as follows: Based on the results of the accident feature terms extracted by the improved TF-IDF algorithm, an index relationship is established between the text and the keyword terms to obtain a two-dimensional table consisting of text information and corresponding terms; Run the programming code to calculate the co-occurrence frequency between keywords, build a co-occurrence matrix, mine the point mutual information between high-weight causes, and calculate the PMI value. Thresholds are set according to business requirements and data characteristics to filter out highly interactive key causal terms whose PMI values and co-occurrence frequencies meet the threshold conditions.
9. The method for constructing a highway traffic accident cause network according to claim 8, characterized in that: In step S5, the PMI values between the accident cause keywords are imported into the Gephi network visualization software as the edge weights, and the resolution value is set for modular processing to complete the drawing of the characteristic item network of highway traffic accident causes; furthermore, the relevant parameters are calibrated to calculate the accident cause network density and standard deviation.
10. The method for constructing a highway traffic accident cause network according to claim 9, characterized in that: In step S5, the specific process of exploring the correlation between different key causes is as follows: S501: The three measurement indicators of point degree centrality, betweenness centrality and closeness centrality are used to represent the centralization degree of each causal node in the accident cause network, as follows: Through the computing node With each node The number of connected edges to obtain the point degree centrality; Statistics Node By other nodes The ratio of the number of shortest paths to the total number of shortest paths in the network is taken as the betweenness centrality, and the calculation formula is: Where: and For nodes, Representation Node To Node The number of shortest paths, Represents a slave node To Node The shortest path passes through the node The number of paths; Reflect the nodes in the network by calculating the shortest distance between nodes With other nodes The closeness of the node is obtained by calculating the closeness of the node. The calculation formula is: Where: Representation Node With Node The shortest distance between S502: Load the core-edge structure of the causes of highway traffic accidents based on UCINET and Core / Periphery modules, and calculate the average density of the core and edge areas; S503: The correlation coefficients between rows and columns in the accident cause keyword co-occurrence matrix are repeatedly calculated using the iterative correlation convergence method to obtain a relationship matrix consisting of two types of coefficients, 1 and -1, and complete the construction of the cohesive subgroup.