Civil aviation complex operation system risk factor extraction method and system and medium

By collecting and cleaning risk-related data in complex civil aviation operation systems, and using large language models and embedded models to extract and merge risk factors, the problem of data heterogeneity and interconnection is solved, and the accuracy and utilization rate of the risk factor library is improved.

CN120181576APending Publication Date: 2025-06-20THE SECOND RES INST OF CIVIL AVIATION ADMINISTRATION OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510254492.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

At present, the full operation safety of the complex civil aviation operation system is facing the problems of data heterogeneity, difficulty in data interconnection, and risk-related reports that have not formed effective data resource integration methods, resulting in idle data and failure to play its due value.

Method used

By collecting and cleaning various types of risk-related data, using pre-trained Chinese large language models and embedded models, risk factors are extracted, clustered and merged to form a risk factor library.

Benefits of technology

Automatic extraction and merging of risk factors has been realized, the redundancy of risk factors has been reduced, the accuracy of the risk factor library has been improved, the understanding and reasoning of risk factors in the later stage has been facilitated, the utilization rate of unsafe incident reports has been improved, and the labor consumption of auditors has been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181576A_ABST
    Figure CN120181576A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a civil aviation complex operation system risk factor extraction method and system and a storage medium. The method comprises the following steps: collecting data for cleaning to obtain paragraph data and a phrase set 1; processing the paragraph data by adopting a large language model to obtain a risk factor set and a phrase set 2; carrying out similar semantic discrimination on the risk elements described by the phrases in the phrase sets 1 and 2 by adopting an embedded model, and clustering to obtain the risk elements; and setting representative words for each class of risk elements, and storing the representative words to form a risk element library. The method has the advantages that the risk elements are extracted and concluded through the large language model, and the risk elements are merged based on the semantic similarity, so that the redundancy of the risk elements can be reduced, and the accuracy of a risk element library can be improved; the utilization rate of unsafe event reports can be improved, and manpower consumption of auditors is reduced; the pre-training model is directly adopted, extra data is not needed, generalization performance is achieved, and prompt words can be adjusted according to requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of civil aviation safety management, and particularly relates to a method, a system and a storage medium for extracting risk elements of a complex civil aviation operation system. Background Art

[0002] At present, the full-process operation safety of the complex civil aviation operation system faces the current situation of diverse safety guarantee methods and means, as well as complex data types and formats. On the one hand, the heterogeneity of data has led to difficulties in the integrated utilization of data by management units, making it difficult to achieve data interconnection; moreover, almost all the original risk-related reports are written in natural language, and in addition to manual review and reading, no effective data resource integration method has been formed, so the valuable data lies idle and fails to play its due value. On the other hand, since the safety risk problems of each operation unit often involve internal privacy, especially the detailed information of specific events, each unit has certain concerns about the sharing of risk information.

[0003] Currently, it is difficult to comprehensively summarize the causes of safety risks (i.e., risk elements) from multiple units such as air traffic control, airlines, and airports only through the text records within a single department or unit. There is a lack of effective utilization of multi-party data resources and no traceability support can be provided for improving the efficiency of safety risk mitigation and optimizing the solutions to safety incidents in the later stage.

[0004] Therefore, it is particularly important to establish a risk element library that fully and comprehensively utilizes the safety risk-related data of each unit. Summary of the Invention

[0005] Aiming at the technical problems mentioned in the background art, the purpose of the embodiments of the present invention is to provide a method, a system and a storage medium for extracting risk elements of a complex civil aviation operation system.

[0006] To achieve the above purpose, in the first aspect, the embodiments of the present invention provide a method for extracting risk elements of a complex civil aviation operation system, including:

[0007] Collect data, and clean the data to obtain first paragraph data, second paragraph data, third paragraph data and a first phrase set; the data includes long-text risk-related data and phrase-format risk-related data;

[0008] Use a pre-trained Chinese large language model to extract and process the first paragraph data, second paragraph data and third paragraph data to obtain a first risk element set, a second risk element set, a third risk element set and a second phrase set;

[0009] Use a pre-trained embedding model to perform similar semantic discrimination on the risk factors described by the phrases in the first phrase set and the second phrase set, and cluster the results of the similar semantic discrimination to obtain the target risk factors;

[0010] Set a risk factor representative word for each type of target risk factor, and store the target risk factors of all categories to form a risk factor library.

[0011] As a specific implementation manner of the present application, obtaining the first paragraph data, the second paragraph data, the third paragraph data, and the first phrase set, the specific process is as follows:

[0012] Collect a list of hazard sources; the descriptions of the hazard sources in the list of hazard sources are short paragraphs composed of natural language, named the first paragraph data;

[0013] Collect aviation accident investigation reports; the aviation accident investigation reports include cause analysis reports described in long texts, named the second paragraph data;

[0014] Collect unsafe event documents; the unsafe event documents include time, location, brief cause description, and casualty situation; the brief cause description is a short paragraph composed of natural language, named the third paragraph data;

[0015] Collect the airline risk list, air traffic control special situation disposal form, and academic literature materials integration with the risk source described in phrase format, and perform structured storage, named the first phrase set.

[0016] As a specific implementation manner of the present application, obtaining the first risk factor set, the second risk factor set, the third risk factor set, and the second phrase set, the specific process is as follows:

[0017] Load the relevant content of the hazard source description in the first paragraph data into the processing system and store it as the first risk factor set;

[0018] Load the content of the investigation conclusion chapter in the second paragraph data into the processing system and store it as the second risk factor set;

[0019] Load the content related to the brief cause description of the unsafe event in the third paragraph data into the processing system and store it as the third risk factor set;

[0020] Set prompt words according to the comprehensive risk factor data management, use convenience, risk factor classification requirements, and the definition of risk factors themselves;

[0021] Use a large language model to automatically extract risk factors from the first risk factor set, the second risk factor set, and the third risk factor set;

[0022] Store the extraction result as the second phrase set.

[0023] As a specific implementation manner of this application, obtaining the target risk factors, the specific process is as follows:

[0024] Load the first phrase set and the second phrase set;

[0025] Use the pre-trained embedding large model to perform v i = g(s i ) and matrix transformation on the first phrase set and the second phrase set to obtain the data set D; where s i is the current risk phrase element, g is the pre-trained embedding large model based on the contrastive learning loss function, and v i is a one-dimensional vector;

[0026] Cluster all the vectors in the data set D according to the DBSCAN rule to obtain the target risk factors.

[0027] As a preferred implementation manner of this application, after storing the target risk factors of all categories to form a risk factor library, the method further includes:

[0028] Have experts review the risk factor library to form the final risk factor library.

[0029] In a second aspect, the embodiment of this application further provides a risk factor extraction system for a civil aviation complex operation system, including:

[0030] The first unit is used to collect data and clean the data to obtain the first paragraph data, the second paragraph data, the third paragraph data and the first phrase set; the data includes long text risk-related data and phrase format risk-related data;

[0031] The second unit is used to perform extraction processing on the first paragraph data, the second paragraph data and the third paragraph data by using a pre-trained Chinese large language model to obtain the first risk factor set, the second risk factor set, the third risk factor set and the second phrase set;

[0032] The third unit is used to use a pre-trained embedding model to perform similar semantic discrimination on the risk factors described in the phrases in the first phrase set and the second phrase set, and cluster the similar semantic discrimination results to obtain the target risk factors;

[0033] The fourth unit is used to set risk factor representative words for each category of target risk factors and store the target risk factors of all categories to form a risk factor library.

[0034] In a third aspect, an embodiment of the present application further provides a risk factor extraction system for a civil aviation complex operation system, including a processor, an input device, an output device, and a memory. The processor, the input device, the output device, and the memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the method of the first aspect above.

[0035] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program therein. The computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method of the first aspect above.

[0036] Implementing the embodiments of the present invention has the following advantages:

[0037] 1. By extracting and summarizing risk factors through a large language model and merging risk factors based on semantic similarity, redundancy of risk factors can be reduced, which is beneficial to improving the accuracy of the risk factor library and facilitating the understanding and reasoning of risk factors in the later stage.

[0038] 2. Automatically reading reports and automatically extracting can improve the utilization rate of unsafe event reports, reduce the labor consumption of auditors, and increase readability and usability.

[0039] 3. Directly adopting a pre-trained model, without additional data, having generalization performance, and being able to adjust prompt words according to requirements. Description of the Drawings

[0040] In order to more clearly illustrate the specific implementation manners of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific implementation manners or the prior art.

[0041] Figure 1 is a flowchart of a method for extracting risk factors of a civil aviation complex operation system provided by an embodiment of the present invention;

[0042] Figure 2 is Figure 1 another flowchart of the method shown;

[0043] Figure 3 is a structural diagram of a risk factor extraction system for a civil aviation complex operation system provided by an embodiment of the present invention;

[0044] Figure 4 is Figure 3 another structural diagram of the system shown. Detailed Implementation Manner

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] It should be understood that when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0047] To solve the difficulty that the current safety risk management record text data composed of natural language is difficult to be integrated automatically through rule fusion, the present invention proposes a method for automatically extracting civil aviation safety risk elements and merging similar semantic elements based on a large language model. Risk element extraction is the ability to extract risk points expressed by different words or phrases from the original text of risk problems described in natural language. The present invention extracts and summarizes risk elements through a large language model, and merges similar semantic elements, so as to summarize the same type of risks and avoid redundancy caused by duplicate listings of risk points. While solving the problem of effective data utilization, the solution provided by the present invention desensitizes the relevant unit information and part of the privacy data in the summarized information, eliminates the concerns of each operating unit about risk information sharing, is conducive to the retrieval of the same type of risks, improves the safety management efficiency, promotes data interconnection and improves data value.

[0048] Please refer to Figure 1 and Figure 2 , the method for extracting risk elements of a civil aviation complex operation system provided by the embodiments of the present invention may include:

[0049] S1. Collect data, and clean the data to obtain first paragraph data, second paragraph data, third paragraph data, and a first phrase set.

[0050] Specifically, when implementing, two types of data, including long text risk-related data and phrase format risk-related data, are collected and data cleaning is performed, specifically including:

[0051] S1.1. Collect a list of hazard sources, which is a structured document containing content such as "hazard source name" - "hazard scenario and consequences" - "existing prevention and control mechanisms". The description of the hazard source in it is a short paragraph composed of natural language, and this type of data is named "first paragraph data".

[0052] S1.2 Collect the aviation accident investigation reports, which are cause analysis reports written in long paragraphs with a fixed directory format and named "Second Paragraph Data".

[0053] S1.3 Collect the unsafe event documents, which are structured documents containing "time", "location", "brief cause description", "casualty situation", etc. The "brief cause description" is a short paragraph composed of natural language and named "Third Paragraph Data".

[0054] S1.4 Collect the airline risk list, air traffic control special situation handling form, and academic literature materials integration with the risk source described in phrase format, and perform structured storage, named "First Phrase Set".

[0055] S2. Use a pre-trained Chinese large language model to extract and process the first paragraph data, second paragraph data, and third paragraph data to obtain the first risk factor set, second risk factor set, third risk factor set, and second phrase set.

[0056] Specifically, when implementing, use a pre-trained Chinese large language model for the "First Paragraph Data", "Second Paragraph Data", and "Third Paragraph Data", and use prompt words to extract the risk factors of the three reports and save the results as a structured table, specifically including:

[0057] S2.1 Load the content related to the hazard source description in the "First Paragraph Data" into the system and save it as the "First Risk Factor Set", where the elements are the description of a single hazard source or the comprehensive description of a series of hazard sources causing the current hazard. For example, "The tower is far from the runway, the tower command position is too low, the terminal building blocks part of the runway, the fourth apron and part of the third apron, and it is impossible to normally observe the aircraft status on the five sides and apron." etc.

[0058] S2.2 Load the "Investigation Conclusion" chapter in the "Second Paragraph Data" into the processing system as the "Second Risk Factor Set".

[0059] Considering that the input format of the transformer is "prompt word" + "text content", if the full text is too long, the requirements in the prompt word may be ignored. Therefore, using the "Second Paragraph Data" directory, only take the "Investigation Conclusion" chapter of each report as the loading item, and load the corresponding part of the "Second Paragraph Data" as the "Second Risk Factor Set", where each element is the investigation conclusion of a single report, such as "Based on the above analysis, it is possible to rule out reasons such as in-flight shutdown and control system failure caused by aircraft mechanical failures. The main reasons may be the following two: one is the low-altitude stall caused by the pilot's improper operation at low altitude or local local airflows; the other is that the pilot loses control of the aircraft due to physical illness, in-flight disability, or loss of situational awareness." etc.

[0060] S2.3 Load the part related to the brief description of the causes of unsafe events in the "Third Paragraph Data" as the "Third Risk Factor Set", where the elements are short paragraphs describing the causes of specific unsafe events, such as "On XX day of X month, XX aircraft carried out individual training flights. During the full stop landing process, due to the pilot's operational error, excessive speed, insufficient judgment of the landing distance, it hit the fence at the end of the runway, resulting in the scrapping of the propeller, damage to the engine compartment, scrapping of the nose landing gear, damage to the left wing due to collision, severe damage and scrapping of the right wing, deformation of the right main landing gear, damage to the right cabin door, and no casualties." etc.

[0061] S2.4 Considering the requirements of the convenience of risk factor data management and use, risk factor classification needs, and the definition of risk factors themselves, use the few-shot learning method to set the prompt as: Extract all risk factors from the following text and list them in the following format with 1, 2, 3, 4; for example, 1. **Insufficient English ability of the crew affects communication**; 2. **Insufficient monitoring ability of the crew**; 3. **Excessive wind speed**.... It is necessary to read the entire article and list all risk factors. Each risk factor in the **** needs to completely summarize a situation, point out the subject, and cannot be a phrase such as 'excessive speed'. Add the subject 'aircraft speed is excessive', not 'improper execution of procedures' but 'crew improper execution of procedures'; do not use specific numbers, for example, do not list 'the aircraft cockpit altitude rises abnormally to 8000 feet and then 14000 feet, triggering the warning mechanism', it should be expressed as 'the aircraft cockpit altitude rises abnormally'; list the reasons instead of the consequences, for example, 'constituting a serious aviation incident' should not be listed because this is a consequence, not a reason.

[0062] S2.5 Use the large language model and the prompt in S2.4 to automatically extract risk factors from the "First Risk Factor Set", "Second Risk Factor Set", and "Third Risk Factor Set", and store the extracted risk factor information in a structured manner as the "Second Phrase Set", where each element is a phrase description of a single risk factor extracted by the large model, such as "Improper execution of flight plan".

[0063] S3. Use the pre-trained embedding model to perform similar semantic discrimination on the risk factors described by the phrases in the First Phrase Set and the Second Phrase Set, and cluster the results of the similar semantic discrimination to obtain the target risk factors.

[0064] Specifically in implementation, use the pre-trained embedding model to perform similar semantic discrimination on the risk factors described by the phrases in the "First Phrase Set" and the "Second Phrase Set", and then use the clustering algorithm to set an appropriate similar semantic threshold, merge multiple risk factors below the similar semantic threshold, and set the first phrase after clustering to replace all phrases within the group, specifically including:

[0065] S3.1 Load the "first phrase set" and "second phrase set". Denote the current risk phrase element as s i , and use the pre-trained embedding large model to perform v i = g(s i ), where g is the embedding large model pre-trained based on the contrastive learning loss function , v i is a one-dimensional vector, specifically a 1*384 vector. Understandably, this length can be adjusted according to specific requirements. In the pre-trained model formula, h i , h j represent samples i and j, L i,j is the loss function value of the pre-trained sample h i and the pre-trained sample h j , M is the total number of samples, τ is the temperature parameter used to adjust the smoothness of the loss function, and sim(·,·) is the similarity function. After using v i = g(s i ) to operate on the total N risk factors of the "first phrase set" and "second phrase set" in sequence, it is converted into an N*384 embedding matrix, denoted as dataset D.

[0066] S3.2 For risk factor 1), the number of clustering categories cannot be determined in advance; 2) the semantic changes of risk factors are relatively large, and there are a large number of risk factors with significant semantic differences from other factors, so they should not be merged into the same category, that is, there are many "outliers" defined in clustering. The commonly used K-means clustering method requires setting the number of categories in advance and is sensitive to outliers. Therefore, in this embodiment, the density method DBSCAN is used to cluster the embedded risk factors. For the vector v i , according to the predefined neighborhood threshold parameter eps, the minimum number of samples parameter MinPts for forming a cluster, and the distance function dist(·,·), define the neighborhood N(v i ) of v i as the set {j∈D|dist(v i ,v j )≤eps}. When the number of points in the neighborhood of v i is not less than MinPts, v i is called a core point. Subsequently, cluster all the vectors in dataset D according to the rules of DBSCAN.

[0067] S3.3 Select two important parameters eps and MinPts in semantic clustering. eps defines the range of the neighborhood around the data points, and this value is obtained through multiple experiments. Since there are individual risk factors with large distinguishability from other factors and do not need to be forced into the same category, the value of MinPts is fixed at 1.

[0068] S3.4 When selecting the distance metric, this embodiment uses the cosine distance as the clustering index, and its calculation formula is where v i and v j represent vectors, the dot product represents the inner product of vectors, and |v i | and |v j | respectively represent the norms of the vectors. Based on this distance metric, DBSCAN clustering analysis is performed on the risk factors.

[0069] S4. Set risk factor representative words for each type of target risk factor, and store the target risk factors of all categories to form a risk factor library.

[0070] S5. Have experts review the risk factor library to form the final risk factor library.

[0071] Implementing the method for extracting risk factors of a civil aviation complex operation system provided by the embodiments of the present invention has the following advantages:

[0072] 1. Extract and summarize risk factors through a large language model, and merge risk factors based on semantic similarity, which can reduce risk factor redundancy, is beneficial to improving the accuracy of the risk factor library, and is convenient for the later understanding and reasoning of risk factors.

[0073] 2. Automatically read reports and extract automatically, which can improve the utilization rate of unsafe event reports, reduce the labor consumption of reviewers, and increase readability and usability.

[0074] 3. Directly adopt a pre-trained model without additional data, has generalization performance, and can adjust the prompt words according to requirements.

[0075] Based on the same inventive concept, the embodiments of the present invention also provide a system for extracting risk factors of a civil aviation complex operation system, as Figure 3 shown, including:

[0076] The first unit is used to collect data and clean the data to obtain the first paragraph data, the second paragraph data, the third paragraph data, and the first phrase set; the data includes long text risk-related data and phrase format risk-related data;

[0077] The second unit is used to perform extraction processing on the first paragraph data, the second paragraph data, and the third paragraph data by using a pre-trained Chinese large language model to obtain the first risk factor set, the second risk factor set, the third risk factor set, and the second phrase set;

[0078] A third unit, configured to perform similar semantic discrimination on the risk factors described by the phrases in the first phrase set and the second phrase set by using a pre-trained embedding model, and cluster the results of the similar semantic discrimination to obtain target risk factors;

[0079] A fourth unit, configured to set representative words for each type of target risk factor and store all categories of target risk factors to form a risk factor library.

[0080] Among them, the first unit is specifically configured to:

[0081] Collect a list of hazard sources; the descriptions of the hazard sources in the list of hazard sources are short paragraphs composed of natural language, named the first paragraph data;

[0082] Collect aviation accident investigation reports; the aviation accident investigation reports include cause analysis reports described in long paragraphs, named the second paragraph data;

[0083] Collect unsafe event documents; the unsafe event documents include time, location, brief cause description and casualty situation; the brief cause description is a short paragraph composed of natural language, named the third paragraph data;

[0084] Collect airline risk lists, air traffic control special situation disposal forms and academic literature materials integration with risk source descriptions in phrase format, and perform structured storage, named the first phrase set.

[0085] Furthermore, the second unit is specifically configured to:

[0086] Load the relevant content of the hazard source description in the first paragraph data into the processing system and store it as the first risk factor set;

[0087] Load the content of the investigation conclusion chapter in the second paragraph data into the processing system and store it as the second risk factor set;

[0088] Load the content related to the brief cause description of the unsafe event in the third paragraph data into the processing system and store it as the third risk factor set;

[0089] Set prompt words according to the comprehensive risk factor data management, use convenience, risk factor classification requirements and the definition of risk factors themselves;

[0090] Use a large language model to automatically extract risk factors from the first risk factor set, the second risk factor set and the third risk factor set;

[0091] Store the extraction results as the second phrase set.

[0092] Furthermore, the third unit is specifically configured to:

[0093] Load the first phrase set and the second phrase set;

[0094] Use a pre-trained embedding large model to perform v on the first phrase set and the second phrase set i = g(s i ) and matrix transformation to obtain the data set D; where s i is the current risk phrase element, g is the pre-trained embedding large model based on the contrastive learning loss function, and v i is a 1*384 vector;

[0095] Cluster all the vectors in the data set D according to the rules of DBSCAN to obtain the target risk factors.

[0096] Among them, during the clustering process of all the vectors in the data set D, the selected clustering parameters are eps and MinPts, and the clustering metric is the cosine distance.

[0097] Furthermore, the fourth unit is further configured to:

[0098] Have experts review the risk factor library to form the final risk factor library.

[0099] Furthermore, another embodiment of the present invention also provides a risk factor extraction system for a civil aviation complex operation system. As Figure 4 shown, the system may include: one or more processors 101, one or more input devices 102, one or more output devices 103, and a memory 104. The above-mentioned processors 101, input devices 102, output devices 103, and memory 104 are interconnected through a bus 105. The memory 104 is used to store a computer program, the computer program includes program instructions, and the processor 101 is configured to call the program instructions to execute the methods in the method embodiments part above.

[0100] It should be understood that in the embodiments of the present invention, the so-called processor 101 may be a central processing unit (Central Processing Unit, CPU), and this processor may also be other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), field-programmable gate arrays (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0101] The input device 102 may include a keyboard, etc., and the output device 103 may include a display (such as an LCD), a speaker, etc.

[0102] The memory 104 may include a read-only memory and a random access memory, and provide instructions and data to the processor 101. A part of the memory 104 may also include a non-volatile random access memory. For example, the memory 104 may also store information about the device type.

[0103] In a specific implementation, the processor 101, the input device 102, and the output device 103 described in the embodiments of the present invention may implement the implementation manners described in the embodiments of the method for extracting risk factors of a civil aviation complex operation system provided by the embodiments of the present invention, which will not be elaborated herein.

[0104] It should be noted that for the specific working process of this embodiment, please refer to the foregoing method embodiment part, which will not be elaborated herein.

[0105] Correspondingly, the embodiments of the present invention provide a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, the following is implemented: the above-mentioned method for extracting risk factors of a civil aviation complex operation system.

[0106] The computer-readable storage medium may be an internal storage unit of the system described in any of the foregoing embodiments, such as the hard disk or memory of the system. The computer-readable storage medium may also be an external storage device of the system, such as a plug-in hard disk equipped on the system, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the system. The computer-readable storage medium is used to store the computer program and other programs and data required by the system. The computer-readable storage medium may also be used to temporarily store data that has been output or will be output.

[0107] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0108] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections between each other can be indirect couplings or communication connections through some interfaces, devices or units, and can also be electrical, mechanical or other forms of connection.

[0109] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present invention.

[0110] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0111] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical discs that can store program codes.

[0112] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for extracting risk factors of complex civil aviation operation systems, characterized in that: include: Collecting data and cleaning the data to obtain first paragraph data, second paragraph data, third paragraph data and a first phrase set; the data includes long text risk-related data and phrase format risk-related data; The first paragraph data, the second paragraph data and the third paragraph data are extracted and processed by using a pre-trained Chinese large language model to obtain a first risk factor set, a second risk factor set, a third risk factor set and a second phrase set; Using a pre-trained embedding model to perform similar semantic discrimination on the risk factors described by the phrases in the first phrase set and the second phrase set, and clustering the similar semantic discrimination results to obtain target risk factors; A risk factor representative word is set for each type of target risk factor, and target risk factors of all categories are stored to form a risk factor library.

2. The method according to claim 1, characterized in that The first paragraph data, the second paragraph data, the third paragraph data and the first phrase set are obtained. The specific process is: Collect a list of hazard sources; the hazard source description in the hazard source list is a short paragraph composed of natural language, which is named as the first paragraph data; Collect aviation accident investigation reports; The aviation accident investigation report includes a cause analysis report described in a long paragraph, named as the second paragraph data; Collect unsafe incident files; the unsafe incident files include time, location, brief description of cause and casualties; the brief description of cause is a short paragraph composed of natural language, named as the third paragraph data; The airline risk list, air traffic control special situation handling form and academic literature with risk source description in phrase format are collected and integrated, and structured storage is performed, which is named the first phrase set.

3. The method according to claim 2, characterized in that The first risk factor set, the second risk factor set, the third risk factor set and the second phrase set are obtained. The specific process is: Loading relevant contents of the hazard source description in the first paragraph of data into a processing system and storing them as a first risk factor set; Loading the contents of the investigation conclusion section in the second paragraph of data into the processing system and storing it as a second risk factor set; Loading the content related to the brief description of the cause of the unsafe event in the third paragraph data into the processing system and storing it as a third risk factor set; Set prompt words based on comprehensive risk factor data management, ease of use, risk factor classification requirements and risk factor definitions; Using a large language model to automatically extract risk factors from the first risk factor set, the second risk factor set, and the third risk factor set; The extraction result is stored as a second phrase set.

4. The method according to claim 3, characterized in that Get the target risk factor. The specific process is as follows: Loading the first phrase set and the second phrase set; The first phrase set and the second phrase set are v-tested using a pre-trained embedding model. i =g(s i ) and matrix transformation to obtain the data set D; where s i is the current risk phrase element, g is the embedding model pre-trained based on the contrastive learning loss function, and v i is a one-dimensional vector; All vectors in the data set D are clustered according to the DBSCAN rules to obtain the target risk factors.

5. The method according to claim 4, characterized in that In the process of clustering all vectors in the data set D, the selected clustering parameters are eps and MinPts, and the clustering index is cosine distance.

6. The method according to any one of claims 1 to 5, characterized in that: After storing all categories of target risk factors to form a risk factor library, the method further includes: Experts review the risk factor library to form a final risk factor library.

7. A risk factor extraction system for complex civil aviation operation systems, characterized in that: include: The first unit is used to collect data and clean the data to obtain first paragraph data, second paragraph data, third paragraph data and a first phrase set; the data includes long text risk related data and phrase format risk related data; The second unit is used to extract and process the first paragraph data, the second paragraph data and the third paragraph data using a pre-trained Chinese large language model to obtain a first risk factor set, a second risk factor set, a third risk factor set and a second phrase set; The third unit is used to use the pre-trained embedding model to perform similar semantic discrimination on the risk factors described by the phrases in the first phrase set and the second phrase set, and cluster the similar semantic discrimination results to obtain the target risk factors; The fourth unit is used to set a risk factor representative word for each type of target risk factor and store all types of target risk factors to form a risk factor library.

8. A risk factor extraction system for complex civil aviation operation systems, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method as claimed in claim 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, the processor is caused to perform the method according to claim 6.