A method and system for investigating, reporting, processing and analyzing water traffic accidents

By using technologies such as large language models and conditional random airports, in-depth analysis of the water traffic accident investigation report is performed, the causes and risk points of the accident are extracted, and the coupled relationship matrix is ​​constructed, which solves the problems of low analysis efficiency, insufficient quantitative analysis and incomplete identification in the existing technology, and achieves efficient, accurate and reliable analysis results.

CN119904107BActive Publication Date: 2025-06-20TRANSPORT PLANNING & RES INST MINIST OF TRANSPORT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510389090.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-20
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The prior art has problems such as inefficiency, lack of accurate quantitative analysis, incomplete identification of causes and risk points of accidents, and insufficient reliability and stability of results in the analysis of water traffic accident investigation reports.

Method used

The text data of water traffic accidents is pre-trained by a large language model, combined with conditional random airports and similarity algorithms, the cause and risk points of accidents are extracted, and the coupling relationship matrix is ​​constructed, and the importance analysis and scenario classification are performed through word frequency statistics and thematic model.

Benefits of technology

It significantly improves the analysis efficiency, enhances the accuracy of identification of accident causes and risk points, realizes the accuracy of quantitative analysis, improves the reliability and stability of results, and provides credible decision-making support for water traffic safety management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904107B_ABST
    Figure CN119904107B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of water transportation, and discloses a method and system for processing and analyzing investigation reports of water traffic accidents, including: pre-training a language model using text data of water traffic accidents to obtain a keyword extraction model; annotating the output of the keyword extraction model using conditional random fields; extracting accident causes and risk points from the water traffic accident report to be analyzed based on the keyword extraction model; calculating the degree of association between each accident cause and between each risk point, and constructing a coupling relationship matrix; screening out strong coupling relationship pairs of accident causes and strong coupling relationship pairs of risk points; performing word frequency analysis on the accident causes and risk points to obtain the importance of the accident causes and risk points; classifying the scenarios of the water traffic accident report to be analyzed, and constructing an accident scenario set. The analysis efficiency of the present invention is greatly improved, the accuracy is significantly enhanced, the quantitative analysis is more precise, and the result reliability is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of water transportation, and particularly to a method and system for processing, analyzing and reporting water traffic accident investigation reports. Background Art

[0002] In the field of water transportation, the analysis of accident investigation reports has mainly relied on manual methods in the past, with expert groups conducting interpretation and evaluation. This traditional method has many drawbacks, such as consuming a large amount of human and time costs, and the analysis results are easily affected by subjective factors of experts, lacking objectivity and consistency. With the development of information technology, some studies have begun to attempt to adopt some simple data statistics and analysis methods, but for unstructured accident report text data, these methods are difficult to deeply mine the key information and potential associations therein.

[0003] In recent years, artificial intelligence technology has been gradually applied in various fields, but in the processing of water traffic accident investigation reports, it is still in the exploratory stage. Although some related studies have attempted to use machine learning algorithms, there are deficiencies in the accuracy and adaptability of the models, and they fail to fully meet the actual needs. For example, some early text analysis models often cannot accurately extract accident causes and risk points when dealing with accident reports with numerous professional terms and complex text structures, and it is also difficult to effectively classify and analyze accident scenarios.

[0004] The main defects of the prior art are as follows: the manual analysis is inefficient and cannot quickly process a large number of accident reports; most of the analysis is qualitative, lacking accurate quantitative analysis and it is difficult to accurately measure the risk level; the ability to mine complex semantics and implicit information in accident reports is limited, resulting in incomplete identification of accident causes and risk points; the results obtained by different analysis methods and personnel vary greatly, lacking reliability and stability.

[0005] Therefore, how to provide a method and system for processing, analyzing and reporting water traffic accident investigation reports is an urgent problem to be solved at present. Summary of the Invention

[0006] Embodiments of the present invention provide a method and system for processing, analyzing and reporting water traffic accident investigation reports to solve the problems of low efficiency of manual analysis, lack of accurate quantitative analysis, incomplete identification of accident causes and risk points, and lack of reliability and stability in the prior art.

[0007] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0008] According to the first aspect of the embodiments of the present invention, a method for processing and analyzing an investigation report of a water traffic accident is provided.

[0009] In one embodiment, a method for processing and analyzing an investigation report of a water traffic accident includes:

[0010] Pre-training a language model using water traffic accident text data so that the language model acquires semantic features and language patterns in the water traffic accident text data and obtains a keyword extraction model; annotating the output of the keyword extraction model using a conditional random field;

[0011] Based on the keyword extraction model, extracting accident causes and risk points from the water traffic accident report to be analyzed; calculating the degree of association between each accident cause and between each risk point using a similarity algorithm, and constructing a coupling relationship matrix; screening out strongly coupled pairs of accident causes and strongly coupled pairs of risk points using the coupling relationship matrix and a coupling degree threshold;

[0012] Performing word frequency analysis on the accident causes and risk points according to word frequency statistics analysis technology to obtain the importance of the accident causes and risk points;

[0013] Using a topic model and combining with the accident causes, classifying the scenes of the water traffic accident report to be analyzed and constructing an accident scene set.

[0014] In one embodiment, before pre-training the language model using water traffic accident text data, it further includes:

[0015] Collecting investigation reports of water traffic accidents and converting the investigation reports of water traffic accidents into text data in a unified format;

[0016] Performing a cleaning operation on the text data in the investigation report of the water traffic accident and removing noise data, irrelevant symbols, and duplicate information;

[0017] Performing lexical and syntactic analysis on the text data in the investigation report of the water traffic accident and standardizing professional terms to obtain water traffic accident text data.

[0018] In one embodiment, annotating the output of the keyword extraction model using a conditional random field includes:

[0019] Constructing a feature function and a transition function of the conditional random field; using the conditional random field and combining with expert experience to annotate the output of the keyword extraction model.

[0020] In one embodiment, calculating the degree of association between each accident cause and between each risk point using a similarity algorithm and constructing a coupling relationship matrix includes:

[0021] Calculate the coupling degrees between each accident cause and each risk point through the cosine similarity calculation formula, and save the coupling degrees in the form of a matrix.

[0022] In one embodiment, using the coupling relationship matrix and the coupling degree threshold, the strongly coupled relationship pairs of accident causes and the strongly coupled relationship pairs of risk points are screened out, including:

[0023] Preset the coupling degree threshold, establish a coupling connection between two keywords that exceed the coupling degree threshold, display the coupling connection situation using a chord diagram, and count the number of coupling connections for each keyword, where the keywords include accident causes and risk points;

[0024] According to the coupling connections established between the keywords, obtain the strongly coupled relationship pairs of accident causes and the strongly coupled relationship pairs of risk points.

[0025] In one embodiment, according to the word frequency statistical analysis technology, perform word frequency analysis on accident causes and risk points, and obtain the importance degrees of accident causes and risk points, including:

[0026] Through the accident cause list and the risk point list, perform word frequency statistics on accident causes and risk points to obtain the word frequency parameters of each accident cause and risk point;

[0027] Arrange in descending order according to the word frequency parameters and display them using a bar chart; determine the coordinates of the corresponding keyword in the scatter plot according to the number of coupling connections of each keyword statistically, and based on the coordinates of each keyword in the scatter plot, determine the importance degrees of the corresponding accident causes and risk points.

[0028] In one embodiment, using the topic model and combining with accident causes, perform scenario classification on the water traffic accident reports to be analyzed, and construct an accident scenario set, including:

[0029] Determine the parameters for training the topic model, including the word frequency amplification factor, the number of training topics, and the number of topic words to be displayed; amplify the word frequency of accident causes and remove the stop words;

[0030] Through the perplexity experiment, observe and select the required number of scenarios and the amplification factor; use the corpus to train the topic model to obtain the trained topic model;

[0031] Combine the number of scenarios, the amplification factor, and the trained topic model to obtain the accident scenario set.

[0032] In one embodiment, when amplifying the word frequency of accident causes, the word frequency amplification formula is: ;

[0033] In the formula, represents the updated word 's word frequency, Indicating its original word frequency;

[0034] Indicating the word frequency increase coefficient of a certain word, Indicating the accident cause set, where k represents the words in the accident report.

[0035] In one embodiment, after constructing the accident scenario set, it includes:

[0036] Classify and name each scenario in the accident scenario set; through the analysis of the relevance between scenarios, accident causes, and risk points, calculate the coupling degree between accident causes and risk points in a certain scenario, establish a coupling connection through a preset threshold, and display it through a Sankey diagram.

[0037] According to the second aspect of the embodiments of the present invention, a water traffic accident investigation report processing and analysis system is provided.

[0038] In one embodiment, the water traffic accident investigation report processing and analysis system includes:

[0039] A keyword extraction model training module, which is used to pre-train a language model using water traffic accident text data, so that the language model can obtain semantic features and language patterns in the water traffic accident text data, and obtain a keyword extraction model; use a conditional random field to label the output of the keyword extraction model;

[0040] A coupling analysis module, which is used to extract accident causes and risk points from the water traffic accident report to be analyzed based on the keyword extraction model; use a similarity algorithm to calculate the degree of association between each accident cause and between each risk point, and construct a coupling relationship matrix; use the coupling relationship matrix and a coupling degree threshold to screen out strong coupling relationship pairs of accident causes and strong coupling relationship pairs of risk points;

[0041] An importance analysis module, which is used to perform word frequency analysis on accident causes and risk points according to word frequency statistics analysis technology to obtain the importance of accident causes and risk points;

[0042] A scenario classification module, which is used to classify the water traffic accident report to be analyzed using a topic model and in combination with accident causes, and construct an accident scenario set.

[0043] According to the third aspect of the embodiments of the present invention, a computer device is provided.

[0044] In some embodiments, the computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0045] According to the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided.

[0046] In one embodiment, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0047] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0048] (1) Greatly improved analysis efficiency: The device of the present invention can quickly process a large number of accident reports through automated data collection and intelligent analysis algorithms, greatly shortening the analysis cycle compared with manual analysis, and improving the timeliness and response speed of water traffic safety management.

[0049] (2) Significantly enhanced accuracy: With the help of advanced large language models and optimized algorithms, the identification of accident causes and risk points is more accurate and comprehensive, reducing omissions and misjudgments, and being able to more precisely mine key information in accident reports, providing a reliable basis for risk assessment. For example, in the analysis of complex accident scenarios, the coupling effects of multiple factors can be accurately identified, while existing technologies may only be able to discover some obvious factors.

[0050] (3) More precise quantitative analysis: Through quantitative analysis methods such as cosine similarity, not only can the correlation between accident causes and risk points be determined, but also the coupling degree and risk importance can be accurately measured, providing an effective means for the quantitative assessment of water traffic risks and making up for the deficiencies of existing technologies in quantitative analysis.

[0051] (4) High reliability of results: Based on a large amount of accident data and scientific algorithm models, the analysis results of the present invention have good stability and consistency, are not affected by human subjective factors, and can provide reliable decision-making support for water traffic safety management departments and shipping enterprises. Compared with traditional methods and some simple data analysis technologies, the reliability has made a qualitative leap.

[0052] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0054] Figure 1 is a flowchart of a method for processing and analyzing a water traffic accident investigation report shown according to an exemplary embodiment;

[0055] Figure 2 is a schematic block diagram of a system for processing and analyzing a water traffic accident investigation report shown according to an exemplary embodiment;

[0056] Figure 3 is a schematic structural diagram of a computer device shown according to an exemplary embodiment;

[0057] Figure 4 is a diagram showing the causes of an accident shown according to an exemplary embodiment;

[0058] Figure 5 is a diagram showing risk factors shown according to an exemplary embodiment;

[0059] Figure 6 is a diagram showing importance shown according to an exemplary embodiment;

[0060] Figure 7 is a diagram showing the calculation results of perplexity shown according to an exemplary embodiment. Detailed implementation manners

[0061] The following description and the drawings fully illustrate the specific embodiments herein, enabling those skilled in the art to practice them. Parts and features of some embodiments may be included in or substituted for parts and features of other embodiments. The scope of the embodiments herein includes the entire scope of the claims and all available equivalents of the claims. Herein, the terms "first", "second", etc. are only used to distinguish one element from another element, and do not require or imply any actual relationship or order between these elements. In fact, the first element can also be called the second element, and vice versa. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a structure, device or equipment including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such structure, device or equipment. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the structure, device or equipment including the said element. The embodiments herein are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same and similar parts among the embodiments can be referred to each other.

[0062] As used herein, the terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. These are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In the description of the present application, unless otherwise specified and defined, the terms "mounted", "connected", and "coupled" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the communication inside two elements. It can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0063] As used herein, unless otherwise specified, the term "plurality" means two or more.

[0064] As used herein, the character " / " indicates that the objects before and after are in an "or" relationship. For example, U / V means: U or V.

[0065] As used herein, the term "and / or" is an associative relationship describing an object, indicating that three relationships can exist. For example, U and / or V means: U or V, or, the three relationships of U and V.

[0066] It should be understood that although the steps in the flowchart are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially according to the order indicated by the arrows. Unless there is a clear indication in the present application, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0067] Each module in the device or system of the present application can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0068] Without conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0069] Figure 1An embodiment of a method for processing, analyzing and reporting on water traffic accident investigations according to the present invention is shown.

[0070] In this alternative embodiment, the method for processing, analyzing and reporting on water traffic accident investigations includes:

[0071] S101. Use the water traffic accident text data to pre-train a language model so that the language model acquires semantic features and language patterns in the water traffic accident text data and obtains a keyword extraction model; use a conditional random field to annotate the output of the keyword extraction model.

[0072] S102. Based on the keyword extraction model, extract accident causes and risk points from the water traffic accident report to be analyzed; use a similarity algorithm to calculate the degree of association between each accident cause and between each risk point, and construct a coupling relationship matrix; use the coupling relationship matrix and a coupling degree threshold to screen out strong coupling relationship pairs of accident causes and strong coupling relationship pairs of risk points.

[0073] S103. According to the word frequency statistical analysis technique, perform word frequency analysis on the accident causes and risk points to obtain the importance of the accident causes and risk points.

[0074] S104. Use a topic model and combine it with the accident causes to classify the scenarios of the water traffic accident report to be analyzed and construct an accident scenario set.

[0075] In this alternative embodiment, before using the water traffic accident text data to pre-train the language model, it further includes:

[0076] Collect water traffic accident investigation reports and convert the water traffic accident investigation reports into text data in a unified format.

[0077] Perform a cleaning operation on the text data in the water traffic accident investigation report and remove noise data, irrelevant symbols and duplicate information.

[0078] Perform lexical and syntactic analysis on the text data in the water traffic accident investigation report and standardize professional terms to obtain water traffic accident text data.

[0079] In this alternative embodiment, using a conditional random field to annotate the output of the keyword extraction model includes:

[0080] Construct the feature function and transition function of the conditional random field; use the conditional random field and combine it with expert experience to annotate the output of the keyword extraction model.

[0081] In this alternative embodiment, using a similarity algorithm to calculate the degree of association between each accident cause and between each risk point and constructing a coupling relationship matrix includes:

[0082] Calculate the coupling degrees between each accident cause and each risk point through the cosine similarity calculation formula, and save the coupling degrees in the form of a matrix.

[0083] In this alternative embodiment, by using the coupling relationship matrix and the coupling degree threshold, the strongly coupled relationship pairs of accident causes and the strongly coupled relationship pairs of risk points are screened out, including:

[0084] Preset the coupling degree threshold, establish a coupling connection between two keywords that exceed the coupling degree threshold, display the coupling connection situation using a chord diagram, and count the number of coupling connections for each keyword, where the keywords include accident causes and risk points.

[0085] Based on the coupling connections established between the keywords, obtain the strongly coupled relationship pairs of accident causes and the strongly coupled relationship pairs of risk points.

[0086] In this alternative embodiment, according to the word frequency statistical analysis technique, perform word frequency analysis on accident causes and risk points to obtain the importance degrees of accident causes and risk points, including:

[0087] Through the accident cause list and the risk point list, perform word frequency statistics on accident causes and risk points to obtain the word frequency parameters of each accident cause and risk point.

[0088] Arrange them in descending order according to the word frequency parameters and display them using a bar chart; determine the coordinates of the corresponding keyword in the scatter plot according to the counted number of coupling connections for each keyword, and based on the coordinates of each keyword in the scatter plot, determine the importance degrees of the corresponding accident causes and risk points.

[0089] In this alternative embodiment, use the topic model and combine it with accident causes to classify the water traffic accident reports to be analyzed, and construct an accident scenario set, including:

[0090] Determine the parameters for training the topic model, including the word frequency amplification factor, the number of training topics, and the number of topic words to be displayed; amplify the word frequency of accident causes and remove the stop words.

[0091] Through the perplexity experiment, observe and select the required number of scenarios and the amplification factor; use the corpus to train the topic model to obtain a trained topic model.

[0092] Combine the number of scenarios, the amplification factor, and the trained topic model to obtain the accident scenario set.

[0093] In this alternative embodiment, when amplifying the word frequency of accident causes, the word frequency amplification formula is: ;

[0094] In the formula, represents the updated word The word frequency of represents its original word frequency; represents the word frequency increase coefficient of a certain word, represents the accident cause set, and k represents the words in the accident report.

[0095] In this alternative embodiment, after constructing the accident scenario set, it includes:

[0096] Classify and name each scenario in the accident scenario set; through the analysis of the relevance between the scenario, accident cause, and risk point, calculate the coupling degree between the accident cause and the risk point in a certain scenario, establish a coupling connection through a preset threshold, and display it through a Sankey diagram.

[0097] Figure 2 Fig. shows an embodiment of a water traffic accident investigation report processing and analysis system of the present invention.

[0098] In this alternative embodiment, the water traffic accident investigation report processing and analysis system includes:

[0099] A keyword extraction model training module 201, configured to pre-train a language model using water traffic accident text data, so that the language model obtains semantic features and language patterns in the water traffic accident text data, and obtain a keyword extraction model; use a conditional random field to label the output of the keyword extraction model.

[0100] A coupling analysis module 202, configured to extract accident causes and risk points from the water traffic accident report to be analyzed based on the keyword extraction model; use a similarity algorithm to calculate the degree of association between each accident cause and between each risk point, and construct a coupling relationship matrix; use the coupling relationship matrix and a coupling degree threshold to screen out strong coupling relationship pairs of accident causes and strong coupling relationship pairs of risk points.

[0101] An importance analysis module 203, configured to perform word frequency analysis on accident causes and risk points according to word frequency statistics analysis technology to obtain the importance of accident causes and risk points.

[0102] A scenario classification module 204, configured to use a topic model and combine accident causes to classify the water traffic accident report to be analyzed and construct an accident scenario set.

[0103] To facilitate the understanding of the above technical solution of the present invention, the above technical solution of the present invention will be further described from the perspectives of architecture and principle as follows:

[0104] Before constructing a framework for maritime accident analysis, it is necessary to define some basic concepts, including the maritime transport system, maritime accidents, accident causes, accident scenarios, hazards, etc. The Maritime Transport System (MTS) is a complex socio-technical environment (STE) system, mainly consisting of three elements: humans, ships, and the environment. If the shore-based operation management agencies and safety supervision agencies are included in the system scope, then the main elements of the maritime transport system can be divided into four subsystems: humans, ships, the environment, and management. The human subsystem consists of the crew on board the ship, generally divided into the deck department and the engineering department, including the captain, first mate / second mate / third mate, pilot, duty officer / helmsman, sailor, chief engineer, motorman, etc.; the ship subsystem mainly consists of the hull, main engine, steering gear, electrical system, air conditioning system, fire protection system, etc.; the environmental subsystem includes natural environmental elements such as wind, waves, currents, and visibility, as well as navigable environmental elements such as waterways, navigational aids, lighthouses, bridges, docks, and other ships; the management subsystem includes regulatory agencies, shipping companies, shipowners and other operating agencies, as well as shipyards and other ship design and production and construction agencies.

[0105] A maritime accident is a traffic accident that occurs to ships and floating facilities in the ocean, coastal waters, and inland navigable waters, such as collisions, groundings, flooding, sinkings, capsizings, hull damages, fires, explosions, main engine damages, cargo damages, crew casualties, marine pollution, etc. An accident is a special event formed by a combination of a series of conditions and events. Generally, Accident or Incident is used to distinguish it. This invention mainly analyzes based on the data that has entered the statistics and formed accident investigation reports, so Accident is more commonly used to represent it.

[0106] Accident Cause (AC) refers to a series of events and conditions that lead to an accident. According to Heinrich's accident cause theory, the occurrence of an accident is a chain reaction of a series of events. The causes of maritime accidents include unsafe acts of humans, unsafe states of ships, unsafe factors in the environment, and deficiencies in management.

[0107] An accident scenario, in a broad sense, refers to the entire life cycle process of an accident, including before, during, and after the accident. In this article, an accident scenario is defined as the situation that leads to the occurrence of an accident, that is, the set of accident causes: .

[0108] The term "Hazard" has multiple definitions in the field of safety science. Generally, it refers to the source of danger to people, property, and the environment. From the perspective of systems theory, this paper defines System Hazard (SH) as the set of hazard points (h) that may trigger system collapse (accidents): 。

[0109] The system hazards of the maritime transportation system include collision hazards, self-sinking hazards, grounding hazards, fire hazards, etc. according to different types of maritime accidents. Hazard points are the system entity components and potential factors that constitute system hazards, and can also be classified into four categories: human, ship, environment, and management. For example, in the case of collision hazards, the hazard points of people include entities such as ship drivers, captains, and watchkeeping sailors, as well as potential factors such as knowledge reserve, safety awareness, mental state, and physiological state; the hazard points of ships include entities such as engines, steering gears, navigation aids, and hulls; the hazard points of the environment include entities such as wind, waves, currents, and navigation obstacles; the hazard points of management include entities such as VTS centers, shipping companies, pilot stations, and various management regulations, as well as potential factors such as safety culture and training.

[0110] Hazard points are closely related to accident causes, and accident causes will necessarily include hazard points. For example, negligence in lookout is an unsafe behavior of people that leads to maritime accidents, and the crew responsible for lookout during ship navigation is one of the hazard points. Hazard points are interconnected and interact with each other, and are constantly changing complexly and dynamically in the maritime transportation system. A certain hazard point will trigger a risk event under specific space-time conditions, and even become an accident cause.

[0111] The present invention relates to a technical device for processing and analyzing reports on maritime traffic accidents using large language models, which is mainly applied to aspects such as maritime traffic safety management and risk prevention and control of shipping enterprises.

[0112] The device of the present invention mainly consists of a data collection and preprocessing module, a model training and optimization module, a risk coupling analysis module, a scenario classification and recognition module, and a result display and output module.

[0113] 1) Data collection and preprocessing module: Responsible for collecting reports on maritime traffic accidents from multiple channels such as maritime-related websites and shipping enterprise databases, and converting them into text data in a unified format. For the text content in the reports, cleaning operations are carried out to remove noise data, irrelevant symbols, and duplicate information. At the same time, lexical and syntactic analysis is performed, and professional terms are standardized to lay a foundation for subsequent analysis. For example, unifying the names of ship equipment expressed differently in different reports into standard terms to facilitate system recognition and processing.

[0114] 2) Model Training and Optimization Module: Adopt advanced large language models such as BERT (Bidirectional Encoder Representation from Transformers) as the core algorithm. First, use a large amount of water traffic accident text data to pre-train the BERT model so that it can learn the semantic features and language patterns in accident reports. On this basis, combined with expert experience and labeled data, fine-tune the model through methods such as Conditional Random Field (CRF) to further optimize the model's performance in accident cause and risk point extraction tasks. For example, for specific accident cause expressions such as "negligent lookout" and "equipment failure", adjust the model parameters through expert annotation and feedback to improve its recognition accuracy.

[0115] 3) Risk Coupling Analysis Module: Based on the trained model, conduct a coupling degree analysis on the extracted accident causes and risk points. Use the cosine similarity algorithm to calculate the correlation degree between various factors, construct a coupling relationship matrix, and screen out strongly coupled relationship pairs according to the set threshold. At the same time, combine the context information and statistical data in the accident report to evaluate the importance of risk points, and comprehensively consider factors such as word frequency, appearance position, and the number of associated factors to determine key risk points. For example, in a certain accident report, if "radar failure" and "ship collision" frequently appear simultaneously and are closely related in the text, it is identified as a high-coupling risk point, and its importance is evaluated according to its word frequency and influence degree in multiple reports.

[0116] 4) Scenario Classification and Recognition Module: Use an improved Latent Dirichlet Allocation (LDA) topic model to classify accident reports. Introduce a prior accident cause list during the word segmentation process, amplify the word frequency of relevant words, and enhance their weights in model training, so as to more accurately identify different accident scenarios. Through training, multiple topic models are obtained, each topic representing an accident scenario, and the topic words are the key accident causes and risk points in that scenario. Further analyze the distribution laws of accident causes and risk points under different scenarios, as well as the conversion relationships and possibilities between scenarios. For example, based on the analysis of a large number of accident reports, identify "collision scenarios caused by bad weather", "accident scenarios caused by ship equipment failures", etc., and analyze the action mechanisms and interrelationships of various risk factors under different scenarios.

[0117] 5) Result display and output module: Display the results of risk coupling analysis and scenario classification and recognition in an intuitive chart form, such as bar charts, pie charts, chord diagrams, scatter plots, and Sankey diagrams, etc. Bar charts are used to compare the word frequencies of different accident causes and risk points. Pie charts show the proportion of various factors in the whole. Chord diagrams present the coupling relationship network between risk points. Scatter plots combine word frequencies and the number of coupling connections to reflect the importance distribution of risk points. Sankey diagrams show the complex associations between scenarios, accident causes, and risk points. At the same time, generate a detailed analysis report, including accident overview, key causes, risk assessment results, and prevention suggestions, etc., to provide users with comprehensive decision-making references.

[0118] The connection relationships and interaction processes between each module are as follows: The data collection and preprocessing module transmits the processed text data to the model training and optimization module. The trained model analyzes and processes the input data, and sends the extracted accident cause and risk point information to the risk coupling analysis module and the scenario classification and recognition module respectively. After these two modules complete their respective analysis tasks, they summarize the results to the result display and output module, and finally present them to the user. The working process of the entire system is: First, the data collection and preprocessing module obtains and processes accident report data, then the model training and optimization module conducts model training and optimization, then uses the trained model to analyze the data, the risk coupling analysis module and the scenario classification and recognition module work in parallel, and finally the result display and output module displays the analysis results in an intuitive form and outputs them in a report form.

[0119] The specific analysis is as follows:

[0120] The extraction of accident causes and hazard points is a keyword extraction task. The extraction process mainly includes three steps: 1. Select the Chinese BERT pre-trained model; 2. Based on expert experience and CRF, perform sequence annotation on the training set, and on this basis, fine-tune the BERT pre-trained model to obtain a keyword extraction model.

[0121] 1. BERT-Embedding is an embedding tool based on the BERT model

[0122] The core of the BERT pre-trained model is to complete BERT-Embedding. For the sake of easy understanding, a description of the accident cause in the corpus is used to explain the process of BERT-Embedding: The input layer (Input) is the original text information, where [CLS] is a classification identifier and [SEP] is a discontinuous identifier. The sentence is randomly segmented into different word chunks; the embedding layer (Embedding) includes token embedding, segment embedding, and position embedding. Through Embedding, the original description "Both the captain and the sailor were on duty on the bridge, but failed to maintain continuous and regular lookouts" is transformed into a sequence of word tensors (Tensors) with three layers of embedding. Since the true semantic information of each word is related to its context, word order, and the semantic set it possesses, after word embedding, a large amount of text data training can be carried out to obtain word tensors with high accuracy. This part of the training task has been pre-completed, so the high training cost is saved, and fine-tuning training on specific corpora can be directly carried out on the basis of this pre-trained model to complete downstream tasks.

[0123] The present invention is directed to accident investigation report data. Such data is highly specialized and the data scale is not large. Therefore, a pre-trained Chinese BERT model (with 12 layers, 768 hidden layers, 12 heads, and 110M parameters) is selected as the pre-trained model. The pre-trained word embeddings can more accurately describe the semantic relationships between characters. On this basis, the downstream task of keyword extraction is completed.

[0124] 2. Data annotation based on CRF

[0125] CRF is a classic method for sequence annotation. Its core idea is that when performing sequence annotation, each character in the sequence is treated as a whole, considering the relationships before and after, rather than individual independent points. The annotation results of each point are dependent on each other, and training is carried out in units of paths. Therefore, through training, the model can, in addition to understanding the text, also understand the regularity knowledge of the output sequence and can better complete the task of extracting key information.

[0126] When using CRF for keyword extraction, it is necessary to define a feature function: ;

[0127] Among them, represents the input sequence, is a character i in the sequence, and represents that the output value is 1 if the prediction is correct, otherwise it is 0. Also known as the emission function or node function, it only judges the prediction result of a single character and does not involve the dependencies between labels. Given the input sequence , if the conditional probability distribution of the predicted sequence Y satisfies the Markov property of the random variable Y, then is only related to its previous label and jointly constitutes a conditional random field with the input X. Therefore, the following transition function needs to be added: ;

[0128] where X represents the observation sequence, i represents a character in the observation sequence, and represent the predicted labels of the i-th and (i - 1)-th characters respectively. If the prediction result of the current label conforms to the expected transition rule of the previous label, the eigenvalue is 1; otherwise, it is 0. The score of the sequence predicted by the CRF is calculated using the following formula: ;

[0129] where represents the calculation result of the predicted feature function of the i-th character in the sequence input to the CRF layer, represents the calculation result of the transition function of the character 1 after this character. Since there are multiple possibilities for the predicted sequence, and only one of them is the most correct, it is necessary to perform global normalization on all possible sequences to generate the probability from the correct sequence to the predicted sequence: ;

[0130] where represents the score of the predicted correct label sequence, while represents the sum of the scores of all possible predicted sequence results. Subsequently, during the training process, a loss function needs to be designed to calculate the maximum probability. The present invention adopts the Viterbi algorithm, which optimizes the time complexity from to .

[0131] Finally, after fine-tuning, a labeling model based on BERT-CRF is obtained. The label combination results are extracted from the corpus, and a hazard set and an accident cause set are constructed according to HZ (hazard) and AC (accident cause) in the BIOES labels respectively.

[0132] 3. Coupling degree analysis

[0133] Based on the BERT pre-training model, the word vectors fine-tuned by combining expert experience and CRF annotation are relatively accurate semantic expressions of hazard points. There is a correlation between hazard points. For example, "crew members" and "captain", "second officer" and "pilot", etc. often express similar semantics and have similar context. In the word vector space, this is represented as a relatively short distance, that is, the absolute value of the modulus is smaller. The cosine similarity is used to calculate the coupling degree between accident causes and between hazard points. The formula for cosine similarity is as follows: ;

[0134] Among them, vector A i represents the i-th vector A, vector Bi represents the i-th vector B, and n represents the total number of vectors. The calculated coupling degree is saved in the form of a matrix, and a threshold C is set. It is assumed that two keywords with a coupling degree exceeding C will establish a coupling connection, and a chord diagram is used to display the coupling connection diagram. After statistics, the coupling connection number (L) of each word can be obtained.

[0135] 4. Importance analysis

[0136] Word frequency statistical analysis is the most commonly used text analysis method. After the accident causes and hazard points are extracted from the accident corpus in the present invention, word frequency statistical analysis must be carried out first. It is necessary to conduct keyword word frequency statistics separately with reference to the accident cause list and the hazard point list, and return the word frequency parameters (TF) for the accident causes and hazard points. The following formula: ;

[0137] represents the word in the document the word frequency, the total length of. TF indicates that the higher the number of times a word appears in a certain document, the larger the value of TF. After that, it is sorted in descending order according to the word frequency parameter and displayed through a bar chart. Combining the statistical coupling connection number (L), as the coordinates (L, F) of the word, the importance of the word is displayed through a scatter plot. The word closer to the upper right corner is considered to have a higher importance.

[0138] 5. Scenario analysis

[0139] The purpose of scenario analysis is to construct an accident scenario set and find the accident causes and risk points included in different scenarios. The present invention uses an improved LDA topic model algorithm to perform accident scenario analysis. The accident scenario set is the trained LDA model, including topics and topic words. Among them, the topic is the scenario, and the topic word is the accident cause.

[0140] ;

[0141] ;

[0142] Among them, Denote the accident scenario set, Denote the accident scenario.

[0143] The number of topics in the LDA model determines its perplexity, and generally, there is an inverse relationship between the two. Therefore, when conducting scenario analysis experiments, the final number of scenarios to be trained is determined by observing the change in perplexity.

[0144] The main steps to construct the scenario set are as follows:

[0145] Step 1. Text preprocessing. It is necessary to preprocess the corpus, save 187 accident reports as 187 consecutive text segments, and remove irrelevant content such as punctuation marks, spaces, line breaks, etc.

[0146] Step 2. Parameter setting. Customize a set of parameters for training the LDA topic model, mainly including the word frequency amplification factor, the number of training topics, and the number of topic words to display.

[0147] Step 3. Set keywords and stop words. Keywords refer to words or phrases that need to be specially segmented. Since in accident reports, the word frequency of accident causes is not the highest, in order to obtain a more accurate LDA topic model, it is necessary to amplify the word frequency of these words. Therefore, all accident causes extracted by BERT are used as keywords for word frequency amplification. At the same time, it is necessary to remove some stop words with high word frequency but little meaning, such as accident, collision, safety, etc.

[0148] Step 4. Word frequency amplification. Amplify the word frequency of accident causes. As shown in the following formula: ;

[0149] In the above formula, Denote the updated word The word frequency of Denote its original word frequency, Denote the word frequency amplification factor (expand number) of a certain word, Denote the accident cause set (cause set). When a word belongs to , the word frequency increases Times, otherwise it does not increase. The value of

[0150] Step 5. Model training. Train the LDA topic model on the corpus to obtain the topic model.

[0151] Step 6. Perplexity calculation. Perplexity determines the accuracy of the topic model for topic classification. The lower the perplexity, the higher the discrimination of the topic model and the better the classification effect. In this paper, a lower perplexity represents a clearer accident scenario. According to different numbers of scenarios (number of topics) and amplification factors , the perplexity is calculated by combining experiments. Observe and select the most suitable combination of the number of scenarios and amplification factors, as well as the trained topic model, which is the constructed scenario set. The perplexity calculation formula is as follows: ;

[0152] where D represents the test set of the corpus, containing M documents; represents the number of words in document d, and represents the probability of generating word w in document d.

[0153] Step 7. Analyze the scenario set and classify and name the scenarios.

[0154] Step 8. Analysis of the relevance between scenarios - accident causes - hazard points. Calculate the coupling degree between accident causes and hazard points in a certain scenario, set a threshold to establish a coupling connection, and display it through a Sankey diagram.

[0155] The key technologies include:

[0156] (1) Select a suitable keyword extraction algorithm and choose a suitable corpus for training. Use the BERT pre - trained model and the CRF algorithm to pre - train the corpus. BERT is designed to pre - train unlabeled text by jointly considering the left and right contexts to obtain deep bidirectional representations. Therefore, with only an additional output layer, the pre - trained BERT model can be fine - tuned, which is very suitable for the use of this project. CRF stands for Conditional Random Field, which combines the characteristics of the maximum entropy model and the hidden Markov model. It has achieved good results in sequence labeling tasks such as word segmentation, part - of - speech tagging, and named entity recognition.

[0157] (2) Calculate the risk coupling degree through cosine similarity. Cosine similarity, also known as cosine similarity measure, evaluates the similarity between two vectors by calculating the cosine value of the angle between them. Cosine similarity plots vectors according to their coordinate values into a vector space, such as the most common two - dimensional space. Cosine similarity measures the size of the angle between two vectors, and the result is represented by the cosine value of the angle. Therefore, the cosine similarity between two vectors is: ;

[0158] where the numerator is the dot product of vector A and vector B, and the denominator is the product of their respective L2 norms, that is, taking the square root after summing the squares of all dimension values. The value of cosine similarity ranges from [-1, 1], and the larger the value, the more similar.

[0159] (3) Train using the LDA model to enable it to achieve the function of scene classification. LDA is the abbreviation of two commonly used models: Linear Discriminant Analysis and Latent Dirichlet Allocation. LDA plays a very important role in topic models and is often used for text classification. In the LDA model, a document is generated in the following way:

[0160] Sample the topic distribution of document i from the Dirichlet distribution. Sample the topic of a certain word in the document from the multinomial distribution of topics. Sample the word distribution corresponding to the topic from the Dirichlet distribution. Sample the words finally from the multinomial distribution of words.

[0161] (4) Use a Sankey diagram to show the connection between topics, topic words, and risk points. The Sankey diagram can very well show the many-to-many dependency relationship.

[0162] (5) Use a platform (such as the Qt platform) to build the software front-end. Qt is a cross-platform C++ graphical user interface library.

[0163] Technical route taken

[0164] 1. In the development of the software front-end interface. In the overall design, modular design is adopted, and the four functions of keyword extraction, risk coupling analysis, exhaustive analysis, and scene classification recognition are made into separate modules and inserted to reduce the coupling degree and adhesion degree. It provides convenience for future function addition and maintenance.

[0165] 2. Contact units such as shipyards and shipping companies to collect as much actual data as possible. Combine the open-source data sets on the Internet to build a relatively complete corpus.

[0166] 3. Combine the existing corpus to train the model. According to the characteristics of ship traffic accidents, adjust the parameters of the trained network.

[0167] The present invention includes four aspects of functions: (1) keyword extraction. (2) Coupling degree analysis. (3) Importance analysis. (4) Scene analysis. The set of accident causes and risk points extracted from the marine accident basic corpus (MABC) using the keyword extraction model trained based on prior knowledge is the basis for the subsequent three types of analysis. Coupling degree analysis is a method used to find the mutual correlation characteristics between accident causes and between dangerous points, and it is necessary to construct a coupling connection relationship based on the distance between word vectors; importance analysis is used to evaluate the importance of accident causes or dangerous points in a certain type of accident, depending on word frequency statistics and the number of coupling connections; scene analysis is used to identify and classify accident scenes and study the relationship between accident causes and risk points in a certain accident scene.

[0168] The initial Maritime Accident Prior Knowledge Base (MAKB) will be used as the prior knowledge of the present invention. The construction process mainly involves the expert group combining literature research to pre-comb out the possible accident causes and risk points of ship collision accidents, and then randomly selecting 20 accident investigation reports for analysis and demonstration to supplement, describe and review the accident causes and risk points. Finally, an expert list of accident causes and risk points is formed. Some of the accident causes and risk points in the expert list are shown in Table 1.

[0169] Table 1 List of Causes and Hazards of Expert Experience

[0170] Reason Hazard Category (c = crew, s = ship, e = environment, m = management) Neglect of lookout Captain c Failure to verify the bearing of the encountering ship Sailor c Failure to properly use the radar Pilot c Fatigue on duty Physical reaction c Unfamiliarity with system documents Attention c Insufficient navigation skills Judgment c Lack of experience Safety awareness c Poor visibility Rain e Large wind and waves Fog e Influence of shore lights at night Wind e High traffic density Wave e Incomplete crew complement Collision regulations m Weak monitoring Navigation rules m Illegal operation Inspection and visa m Employment of unlicensed personnel Duty schedule m Absence of management Handover specification m Main engine failure Light s Failure of reverse gear Sound signal s Failure to conduct routine maintenance of the main engine as planned High frequency s Failure of navigation aids Steam whistle s Overloading Radar s

[0171] For the ship collision accident set, the proposed Corpus Content-Aware Maritime Accident Analysis Model (CAM-MA model) of the present invention is applied to mine accident causes and risk points and analyze the characteristics of collision accidents. The training parameters of the keyword extraction model trained based on BERT-CRF are shown in Table 2. The parameters of the two types of keywords are the same. The accuracy (Accuracy) of directly using the BERT pre-trained model for extraction is 85.95%. After fine-tuning, the model accuracy is 98.19%, showing a significant improvement.

[0172] Table 2 Training Parameters of the Keyword Extraction Model Trained Based on BERT-CRF

[0173] Training parameters Cause extraction model Hazard extraction model Size of training set 6141 6141 Batch size 32 32 Number of batches 192 192 Epoch 4 4 Loss function Cross entropy Cross entropy Hardware environment Quadro RTX 6000 (architecture = 7.5)

[0174] In the accident cause set, a total of 139 accident causes are extracted and classified according to human, ship, environment and management. There are 90 human factors, accounting for about 45.39%, 14 ship factors, accounting for about 17.34%, 15 environmental factors, accounting for about 15.50%, 20 management factors, accounting for about 21.77%, and 2 other factors, accounting for about 0.6%. Through comparative analysis, it can be seen that since the main cause of ship collision accidents is human factors, the description of human factors in the accident investigation reports is more abundant and specific, determining that the largest number of accident causes can be extracted. According to the word frequency of risk points, the top 20 accident causes with the highest word frequency are shown in Table 3 and Figure 4 .

[0175] Table 3 List of Accident Causes (Top 20 / 139)

[0176] # Name Word frequency 1 Poor visibility 244 2 Neglect of lookout 92 3 Failure to maintain a proper lookout 80 4 Keep clear of other ships 54 5 Fatigue 26 6 Ship unseaworthy 24 7 Failure to sound the sound signal as required 20 8 High traffic density 16 9 Running along the outer edge 16 10 Subjective 12 11 Failure to apply good seamanship 10 12 Weak safety awareness 10 13 Failure to navigate at a safe speed 10 14 Sailing beyond the approved navigable area 10 15 Failure to equip as required 10 16 Safety inspection 8 17 Failure to detect in time 8 18 Carelessness 8 19 Competency assessment 8 20 Failure to take effective avoidance measures 6

[0177] The risk points are concentrated. When directly using the BERT pre-trained model for extraction, the accuracy is 85.95%. After fine-tuning, the model accuracy is 98.19%, showing a significant improvement. A total of 132 risk points are extracted and classified according to human, ship, environment, and pipeline. There are 33 risk points for humans, accounting for about 25%; 33 risk points for ships, accounting for about 25%; 27 risk points for the environment, accounting for about 20.45%; and 20 risk points for pipelines, accounting for about 29.55%. According to the word frequency of the risk points, the top 10 words with the highest frequency are International Regulations for Preventing Collisions at Sea, seafarers, regulations, shipmasters, companies, anchors, radars, ship managers, fishermen's safety awareness, and pilots. Most of them are related to human factors and management factors, as shown in Table 4 and Figure 5 .

[0178] Table 4 Risk List (Top 19 / 280)

[0179] # Name Word frequency 1 International Regulations for Preventing Collisions at Sea 1242 2 Crew 908 3 Regulations 846 4 Captain 784 5 Company 612 6 Anchor 378 7 Radar 318 8 Ship manager 306 9 Fishermen's safety awareness 288 10 Driver 282 11 Pilot station 266 12 Second officer 182 13 Vision 176 14 Seamanship 168 15 Bridge watch personnel 152 16 Maritime Traffic Safety Law 142 17 Watch officer 142 18 Pilot 142 19 Electronic chart 116

[0180] In terms of coupling degree, the coupling degree is calculated according to the cosine similarity, and the top 20 couplings after descending order are shown in Table 5. Selecting the threshold from low to high determines the scale of the overall coupling links. For example, when the threshold is 0.7, there are 1375 coupling links; when it is 0.8, there are 148 coupling links; and when it is 0.85, there are 30 coupling links.

[0181] Table 5 Risk Coupling Links (Top 20)

[0182]

[0183] The display of the coupling degree is as Figure 6 shown, Figure 6 where the horizontal axis represents the number of topics and the vertical axis represents the perplexity. At the same time, the accident scenarios are concentrated. In order to train a more accurate scenario analysis model, a perplexity calculation experiment is carried out, and the experiments are carried out according to e = (1, 10, 100, 500, 1000, 2000) respectively, and the maximum number of topics is set to 60. The results show that the value of e affects the upper limit of the perplexity. As e increases, the perplexity decreases significantly. This also reflects the role of improving LDA and increasing the amplification factor. When e exceeds 500, the decline of the upper limit of the perplexity slows down. In addition, although e affects the upper limit of the perplexity, the downward trend of the perplexity with the number of topics is basically the same. When the number of topics = 40, the decline rate of the perplexity begins to decrease significantly and even rebounds. When the number of topics and the amplification factor are too large, overfitting is likely to occur. Therefore, finally, the number of topics 40 and the amplification factor 500 are selected to train the topic model. The calculation results of the perplexity are as Figure 7 shown.

[0184] The present invention is summarized as follows:

[0185] 1. An accident analysis model CAM-MA was constructed, and a ship traffic accident risk and scenario analysis platform software was developed and applied to the ship collision accident text dataset. Combining expert prior knowledge, the BERT-CRF model was trained to extract 280 risk points and 320 accident causes. At the same time, word frequency analysis, coupling degree analysis based on cosine similarity, and scenario analysis based on the improved LDA topic model were carried out on the accident causes and risk points. The final results obtained a more comprehensive accident cause set, risk point set, coupling connection set, and accident scenario set. Based on the results, the characteristics of coastal ship collision accidents were analyzed, and suggestions for risk classification criteria for ship collision accidents were put forward.

[0186] 2. A ship navigation safety structure network model based on STAMP was constructed. Facing the ship collision accident set, its complex network characteristics were analyzed from the accident scenarios at different stages such as "collision danger", "close quarters situation", and "imminent danger", and the main risk characteristics existing in the ship navigation safety structure under the collision scenario were studied.

[0187] 3. A ship traffic accident risk quantitative assessment model based on risk entropy was constructed. The key risk nodes and edges were quantitatively evaluated, and risk prevention and control strategies were proposed.

[0188] In addition, in terms of model training, other deep learning-based language models such as the GPT (Generative Pretrained Transformer) series models can be tried, and combined with specific transfer learning techniques, optimized training can be carried out according to the characteristics of water traffic accident reports, and similar or better effects will be achieved in some aspects, but the model structure and training strategy need to be redesigned and adjusted. In terms of the risk coupling analysis algorithm, in addition to cosine similarity, methods based on graph neural networks can also be explored to construct a relationship graph of accident factors, and the coupling relationship can be analyzed through the feature learning of the nodes and edges of the graph, but this requires more complex algorithm design and computing resource support, and the effect in practical applications needs to be further verified and optimized.

[0189] The key points of the present invention are:

[0190] (1) A method and model structure for accurately extracting accident causes and risk points from water traffic accident reports by using large language models such as BERT in combination with expert experience and the CRF algorithm.

[0191] (2) A unique algorithm and parameter setting for coupling degree analysis and scenario classification recognition of accident causes and risk points based on cosine similarity and the improved LDA topic model.

[0192] (3) The close collaboration, efficient data transmission and processing flow among modules, as well as the overall architecture design of the system, ensure a comprehensive, rapid, and accurate analysis of accident reports.

[0193] The present invention solves the following technical problems:

[0194] First, it improves the analysis efficiency of water traffic accident investigation reports and realizes automated and intelligent processing; second, it enhances the accuracy of identifying accident causes and risk points and comprehensively extracts key information; third, it precisely evaluates the risk level and coupling relationship through quantitative analysis methods; fourth, it provides stable and reliable analysis results to provide strong support for water traffic safety management and decision-making.

[0195] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 3 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store static information and dynamic information data. The network interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0196] Those skilled in the art can understand that Figure 3 the structure shown in

[0197] is only a block diagram of part of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0198] In addition, the present invention also provides a computer device including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.

[0199] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0200] The present invention is not limited to the structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A method for processing and analyzing a water traffic accident investigation report, characterized in that: include: The language model is pre-trained using the water traffic accident text data, so that the language model can obtain the semantic features and language patterns in the water traffic accident text data, and a keyword extraction model is obtained; the output of the keyword extraction model is annotated using the conditional random field; Based on the keyword extraction model, the causes and risk points of the water traffic accident reports to be analyzed are extracted; the similarity algorithm is used to calculate the correlation between the causes of the accidents and the risk points, and a coupling relationship matrix is ​​constructed; Using the coupling relationship matrix and coupling degree threshold, the strong coupling relationship pairs of accident causes and risk points are screened out; According to the word frequency statistical analysis technology, the word frequency analysis of the accident causes and risk points is carried out to obtain the importance of the accident causes and risk points; Using the topic model and combining the causes of the accident, the water traffic accident reports to be analyzed are classified into scenarios and an accident scenario set is constructed, including: Determine the training parameters of the topic model, including the frequency amplification coefficient, the number of training topics, and the number of topic word displays; amplify the frequency of the cause of the accident and remove stop words; Through the perplexity experiment, the required number of scenes and the magnification factor are selected and observed; the topic model is trained using the corpus to obtain a trained topic model; Combining the number of scenes, the magnification factor and the trained topic model, the accident scene set is obtained.

2. A method for processing and analyzing water traffic accident investigation reports according to claim 1, characterized in that: Before the language model is pre-trained using the water traffic accident text data, the following steps are also included: Collect water traffic accident investigation reports and convert them into text data in a unified format; Clean the text data in the water traffic accident investigation report and remove noise data, irrelevant symbols and repeated information; The text data in the water traffic accident investigation report is subjected to lexical and syntactic analysis, and the professional terms are standardized to obtain the water traffic accident text data.

3. A method for processing and analyzing water traffic accident investigation reports according to claim 1, characterized in that: The method of labeling the output of the keyword extraction model using the conditional random field includes: Construct the characteristic function and transfer function of the conditional random field; use the conditional random field and combine it with expert experience to annotate the output of the keyword extraction model.

4. A method for processing and analyzing water traffic accident investigation reports according to claim 1, characterized in that: The similarity algorithm is used to calculate the correlation between the causes of accidents and the risk points, and to construct a coupling relationship matrix, including: The cosine similarity calculation formula is used to calculate the coupling degree between the causes of accidents and between the risk points, and the coupling degree is saved in the form of a matrix.

5. The method for processing and analyzing water traffic accident investigation reports according to claim 1, characterized in that: The coupling relationship matrix and the coupling degree threshold are used to screen out the accident cause strong coupling relationship pairs and the risk point strong coupling relationship pairs, including: A coupling degree threshold is set in advance, and a coupling connection is established between two keywords that exceed the coupling degree threshold. The coupling connection status is displayed using a chord diagram, and the number of coupling connections for each keyword is statistically obtained. The keywords include accident causes and risk points. According to the coupling connections established between the keywords, the strongly coupled relationship pairs of accident causes and the strongly coupled relationship pairs of risk points are obtained.

6. A method for processing and analyzing water traffic accident investigation reports according to claim 1, characterized in that: According to the word frequency statistical analysis technology, the word frequency analysis is performed on the accident causes and risk points to obtain the importance of the accident causes and risk points, including: Through the accident cause list and risk point list, the word frequency statistics of accident cause and risk point are carried out to obtain the word frequency parameters of each accident cause and risk point; Arrange in descending order according to the word frequency parameter and display it in a bar graph; determine the coordinates of the corresponding keyword in the scatter plot based on the statistical number of coupling connections for each keyword, and determine the importance of the corresponding accident causes and risk points based on the coordinates of each keyword in the scatter plot.

7. A method for processing and analyzing water traffic accident investigation reports according to claim 1, characterized in that: When the frequency of the word causing the accident is amplified, the frequency amplification formula is: ; In the formula, Indicates the updated word k The word frequency, represents its original word frequency; e represents the frequency increase coefficient of a certain word, C represents the set of accident causes, k Words that represent incident reports.

8. The method for processing and analyzing water traffic accident investigation reports according to claim 1, characterized in that: The construction of the accident scenario set includes: Classify and name each scene in the accident scene set; calculate the coupling degree between the accident cause and the risk point in a scene through the correlation analysis of scenes, accident causes and risk points, establish the coupling connection through the preset threshold, and display it through the Sankey diagram.

9. A water traffic accident investigation report processing and analysis system, characterized in that: include: The keyword extraction model training module is used to pre-train the language model using the water traffic accident text data, so that the language model can obtain the semantic features and language patterns in the water traffic accident text data, and obtain the keyword extraction model; the output of the keyword extraction model is annotated using the conditional random field; The coupling analysis module is used to extract the causes and risk points of the water traffic accident reports to be analyzed based on the keyword extraction model; the similarity algorithm is used to calculate the correlation between the causes of the accidents and the risk points, and to construct a coupling relationship matrix; Using the coupling relationship matrix and coupling degree threshold, the strong coupling relationship pairs of accident causes and risk points are screened out; Importance analysis module, used to perform word frequency analysis on accident causes and risk points based on word frequency statistical analysis technology to obtain the importance of accident causes and risk points; The scene classification module is used to classify the water traffic accident reports to be analyzed by using the topic model and combining the causes of the accident, and to construct an accident scene set, including: Determine the training parameters of the topic model, including the frequency amplification coefficient, the number of training topics, and the number of topic word displays; amplify the frequency of the cause of the accident and remove stop words; Through the perplexity experiment, observe and select the required number of scenes and the magnification factor; use the corpus to train the topic model to obtain a trained topic model; Combining the number of scenes, the magnification factor and the trained topic model, the accident scene set is obtained.

Citation Information

Patent Citations

  • Methods, systems and storage media for risk evolution analysis of water traffic accidents

    CN114936759A

  • Intelligent traffic text analysis method based on natural language processing

    CN115934936A