Civil aviation risk cause network association analysis method based on semantic enhancement model
By integrating the enhanced topic model of BERT and LDA with the improved Apriori algorithm and combining it with the Logistic regression model, the semantic understanding and dynamic mining problems in the causal analysis of civil aviation risks were solved, enabling accurate analysis of civil aviation risks and the formulation of control strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2026-01-08
- Publication Date
- 2026-05-15
AI Technical Summary
Existing methods for analyzing the causes of civil aviation risks are insufficient in terms of semantic understanding depth, dynamic mining of association rules, and accuracy of risk quantification, making it difficult to effectively support deep semantic understanding and dynamic evolution analysis.
We employ an enhanced topic model that integrates BERT and LDA for deep semantic topic mining. By combining an improved Apriori algorithm and a Logistic regression model, we dynamically capture strongly correlated causal rules and construct a risk causal network. We improve the accuracy of analysis through adaptive thresholding and multi-source feature fusion.
It achieves deep semantic understanding and dynamic association rule mining of civil aviation risk events, accurately identifies core risk nodes and key propagation paths, and provides precise risk quantification indicators and network structure characteristics to support civil aviation safety management decisions.
Smart Images

Figure CN122046292A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of civil aviation safety management and data mining technology, specifically to a method for civil aviation risk causal network correlation analysis based on semantic enhancement models. Background Technology
[0002] Civil aviation risk causes are characterized by multi-source, hidden, and correlated features. However, the massive amount of unstructured text data poses a severe challenge to the quantitative mining of causal correlation patterns, making it difficult to directly support deep semantic understanding and dynamic evolutionary analysis. Traditional safety management methods mainly rely on expert experience and statistical analysis, which have limitations in terms of semantic depth, implicit correlation mining, and dynamic evolutionary analysis. Existing data mining and machine learning technologies have provided new ideas for civil aviation risk cause analysis. LDA topic models and BERT models have shown effectiveness in text mining, the Apriori algorithm has been applied to association rule mining, and complex network methods are used for accident cause modeling. However, existing research mostly uses BERT for text classification tasks, with less emphasis on deep integration with topic models for in-depth mining of civil aviation professional semantics; the Apriori algorithm uses a fixed threshold, which is difficult to adapt to the volatility and sparsity of civil aviation risk data; network analysis still has room for improvement in risk quantification, multi-source feature fusion, and dynamic threshold optimization; and there is a lack of a systematic method that combines deep semantic analysis, dynamic rule mining, and network risk quantification. Summary of the Invention
[0003] Purpose of the invention: The purpose of this invention is to provide a network association analysis method for civil aviation risk causation based on a semantic enhancement model, which solves the shortcomings of existing civil aviation risk causation analysis methods in terms of semantic understanding depth, dynamic mining of association rules, and accuracy of risk quantification.
[0004] Technical solution: The present invention provides a method for analyzing the network correlation of civil aviation risk causes based on a semantic enhancement model, comprising the following steps:
[0005] S1. Obtain textual data on civil aviation transport accidents and incidents, and construct a civil aviation risk event dataset;
[0006] S2. Construct an enhanced BERT-LDA topic model that integrates BERT semantic representation and quality adaptive adjustment to perform deep semantic topic mining on text in the civil aviation risk event dataset and identify the set of basic causal types.
[0007] S3. An improved Apriori algorithm based on an adaptive threshold mechanism for the number of factors is used to mine association rules in the basic causal type set and dynamically capture strongly associated causal rule sets.
[0008] S4. Construct an optimized Logistic regression model that integrates multi-source features, and quantify the risk contribution of each causative type based on the basic causative type and the causative rule set;
[0009] S5. Based on the output of the optimized Logistic regression model and the set of causal rules, construct a civil aviation risk causal association network, identify the core risk nodes and key risk propagation paths in the network, and output the basis for risk management and control decisions.
[0010] Further, step S1 is as follows: load the civil aviation professional dictionary and stop word list, perform customized word segmentation, retain content words through part-of-speech tagging, and deduplicate and standardize the word segmentation results.
[0011] Furthermore, in step S2, the enhanced BERT-LDA topic model includes: a BERT semantic encoding module for encoding documents into document-level semantic vectors; a semantic space mapping module for mapping the continuous semantic space of BERT into a discrete topic probability distribution compatible with the LDA model; a quality adaptive weight adjustment module for dynamically adjusting the weights of the BERT semantic distribution and the LDA topic distribution during fusion based on the consistency scores of each topic; a fusion module for performing linear weighted fusion of the BERT semantic distribution and the LDA topic distribution with sparsity compensation to generate the final topic distribution; and a topic optimization module for filtering generalized words based on knowledge in the civil aviation field and enhancing the weights of professional terms, outputting a set of basic causal types.
[0012] Furthermore, in step S3, the improved Apriori algorithm includes: an adaptive threshold mechanism for the number of factors used to dynamically calculate the minimum support threshold based on the number of factors contained in the current mined itemset, wherein a lower support requirement is applied to multi-factor itemsets compared to single-factor itemsets; a multi-level constraint rule screening mechanism used to screen strong association rules from candidate rules that simultaneously satisfy support, confidence, and lift constraints based on dynamically adjusted confidence thresholds and preset lift thresholds; and a rule risk scoring function used to calculate a comprehensive risk score for the screened association rules.
[0013] Furthermore, in step S4, optimizing the Logistic regression model includes: an extended feature space construction module for constructing a multi-source fusion feature space containing basic causal features, inter-causal interaction features, and association rule triggering features; a model construction and optimization module for constructing a Logistic regression model with the extended feature space as input, and introducing loss function optimization and L2 regularization to prevent overfitting; and a risk contribution quantification module for outputting single-factor risk confidence and synergistic enhancement effect values of multi-factor coupling.
[0014] Furthermore, in step S5, the construction and analysis of the civil aviation risk causal association network includes: a network construction module, used to construct an undirected weighted network with basic causal types as nodes and association strength calculated based on risk-weighted co-occurrence values as edge weights; a node importance comprehensive evaluation module, used to integrate network topology centrality indicators, Logistic regression model coefficients, and single-factor risk confidence scores to calculate the comprehensive importance index of each node; an adaptive edge screening module, used to dynamically adjust edge weight thresholds according to sample size to screen out key association edges; and a core node and path identification module, used to identify core risk nodes based on the comprehensive importance index and identify key risk propagation paths based on edge weights and node importance.
[0015] Furthermore, the final outputs are as follows: risk causal network topology, list of core risk nodes, high-risk association rules and their synergistic enhancement effect values, which are used to support the formulation of differentiated civil aviation safety management and control strategies.
[0016] An electronic device according to the present invention includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the steps of any of the methods described herein. The present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the methods described herein.
[0017] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: First, by integrating BERT and LDA, the invention achieves deep semantic understanding of professional terms and implicit causes in civil aviation accident and incident reports, overcoming the problem of shallow semantic understanding in traditional methods. Second, the invention introduces an adaptive threshold mechanism for the number of factors, enabling association rule mining to adapt to dynamic changes in data features, thus improving the accuracy and applicability of rule mining. Furthermore, the invention constructs an extended feature space containing basic features, interaction features, and rule features, and achieves accurate quantification of risk contribution by optimizing the Logistic model. Simultaneously, the invention introduces an adaptive threshold mechanism and a comprehensive importance index for network analysis, accurately identifying core risk nodes and key propagation paths, avoiding bias caused by fixed thresholds. The risk quantification indicators and network structure characteristics output by the method of this invention can directly support civil aviation safety management decisions, possessing strong engineering application value and practicality, and contributing to the development of civil aviation safety risk assessment technology. Attached Figure Description
[0018] Figure 1 This is a flowchart of the present invention;
[0019] Figure 2 A schematic diagram illustrating the algorithm performance of this invention. Detailed Implementation
[0020] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0021] like Figure 1 As shown, this embodiment of the invention provides a method for civil aviation risk causal network correlation analysis based on a semantic enhancement model, including the following steps:
[0022] S1. Obtain civil aviation transport accident and incident text data to construct a civil aviation risk event dataset; the civil aviation risk event dataset includes: accident and incident report text, event date, event type, aircraft information, event description, cause analysis, and other information.
[0023] The text data undergoes preprocessing, including loading a civil aviation professional dictionary and a stop word list, followed by customized processing using the jieba word segmentation tool. The civil aviation professional dictionary contains specific terminology within the civil aviation field, ensuring the completeness of professional terminology during word segmentation; the stop word list includes modal particles and other words that contribute little to semantic analysis. Nouns, verbs, and adjectives are retained through part-of-speech tagging, and the segmentation results are deduplicated and standardized to unify the expression of synonyms.
[0024] S2. Construct an EnhancedBertLdaModel model that integrates BERT semantic representation and quality adaptive adjustment to perform topic mining on civil aviation risk event texts and identify underlying cause types; including the following steps:
[0025] S2-1. The BERT model is used to semantically encode documents, capturing contextual semantics through a multi-layer Transformer self-attention mechanism. BERT achieves this through a self-attention mechanism:
[0026]
[0027] Where Q is the query matrix, K is the key matrix, and V is the value matrix. This is the scaling factor for the key vector. For a document containing m words, the document semantic vector is calculated as follows:
[0028]
[0029] in, Indicator The vector representation obtained after the BERT embedding layer is typically 768-dimensional. `Transformer` represents a multi-layer self-attention transformation, and `Pooling` represents the pooling operation. This results in the final document-level semantic vector. Cosine similarity is used to reflect the strength of semantic associations between documents:
[0030]
[0031] S2-2. Mapping BERT's continuous semantic space to a discrete probability distribution based on semantic similarity. This mapping is a key bridge connecting deep learning models and probabilistic topic models. For each word w in the vocabulary, calculate its probability distribution under each topic:
[0032]
[0033] in, The center vector of topic k is obtained by weighted summation of the topic distributions of all documents:
[0034]
[0035] in This represents the probability that document d belongs to topic k; Let w be the BERT embedding vector of word w; The cosine similarity between the topic center vector and the word vector measures the semantic relevance between the word and the topic. The similarity scaling factor controls the steepness of the distribution; softmax normalization ensures that the sum of probabilities is 1. This mapping mechanism allows semantically similar words to receive higher probabilities under the same topic.
[0036] S2-3. Design an adaptive weight adjustment mechanism based on topic quality to dynamically adjust the fusion weights of BERT and LDA. Utilize topic consistency scores. This measure assesses the semantic coherence of words within a topic, with a value ranging from 0 to 1. A higher score indicates better topic quality. The weighting adjustment formula is as follows:
[0037]
[0038] Able to ensure When the topic consistency score is low, Larger A score close to 0.3 indicates a high BERT weight, leveraging BERT's semantic capabilities to compensate for LDA's shortcomings; when the topic consistency score is high, A value close to 0.1 indicates a higher LDA weight, preserving the topic structure mined by LDA. This adaptive mechanism enables differentiated processing of topics of varying quality.
[0039] S2-4. Construct a fusion function with sparsity compensation to solve the dimensionality mismatch problem between the continuous distribution of BERT and the discrete distribution of LDA. Through linear weighting, an organic combination of the semantic distribution of BERT and the topic distribution of LDA is achieved, which retains both the deep semantic understanding capability of BERT and the topic structure mining capability of LDA.
[0040]
[0041] The first term is the semantic information contributed by BERT, the second term is the topic structure contributed by LDA, and the third term is the sparsity compensation term. As a smoothing factor, it compensates for a certain probability quality when the LDA probability is too small, thus avoiding the generation of zero-probability words.
[0042]
[0043] Normalization ensures the validity of the probability distribution.
[0044] S2-5. In the topic optimization phase, civil aviation domain knowledge is embedded for refined processing. By setting a domain terminology whitelist and a generalized word blacklist, overly generalized high-frequency words are filtered out. These words appear in almost all documents and contribute little to topic differentiation. At the same time, the weight of civil aviation professional terms in topic distribution is strengthened to improve professional semantic representation capabilities. After optimization, a basic causal type set F is output, where each causal type corresponds to a topic, including keywords and probability distributions related to that type.
[0045] S3. Based on the improved Apriori algorithm, association rule mining is performed on the identified causal types to dynamically capture strongly associated causal rules; including the following steps:
[0046] S3-1. Using the fundamental factors identified by LDA as an itemset F, construct a transaction for each incident record. Transaction It is the collection of all causal factors appearing in the incident record. Transaction Encoder is used to encode transaction data into a Boolean matrix. Where N is the total number of accident records, This represents the total number of factors. The elements in the matrix... This indicates that the i-th record contains the j-th factor. This indicates that it does not include.
[0047] S3-2. Introducing an adaptive threshold mechanism based on the number of factors, the minimum support is dynamically calculated according to the sample set size and the number of factors. Traditional Apriori uses a fixed support threshold, but civil aviation data suffers from sample sparsity and diverse factor combinations; a fixed threshold easily misses important but low-frequency rules or generates a large number of noisy rules. The formula for calculating the minimum support is:
[0048]
[0049] Where n is the number of factors, i.e. the size of the itemset; For the sample set size, As the baseline support, For adaptive coefficients, This is the floor function. This mechanism ensures that single-factor rules use a higher threshold to filter out overly generalized rules, while multi-factor rules use a lower threshold to capture sparse but important coupling patterns.
[0050] To further adjust the threshold for the multi-factor rule, a scaling factor is introduced:
[0051]
[0052] in This is a scaling factor that decreases as the number of factors increases, ensuring that more complex rules receive more lenient support requirements. The generation conditions for frequent itemset families are:
[0053]
[0054] in This represents the set of k frequent itemsets, where support(X) is the frequency of itemset X in the dataset. Multiple frequent itemsets are generated progressively by scanning layer by layer, starting with single itemsets.
[0055] S3-3, Generate candidate association rules from frequent itemsets. For frequent itemsets... Rules can be generated This indicates that when factor set A occurs, factor set B is also likely to occur. Define the reliability of the confidence measure rule:
[0056]
[0057] Confidence level indicates the proportion of incidents containing A that also contain B. Higher confidence level indicates a more reliable rule. Lifting degree measures the significance of a rule:
[0058]
[0059] Lift measures the degree to which a correlation is stronger than a random correlation. When lift > 1, it means that the correlation between A and B is stronger than in the random case; lift = 1 means that A and B are independent; lift < 1 means that A and B are negatively correlated.
[0060] Introduce an adaptive confidence threshold, which is dynamically adjusted based on the number of factors:
[0061]
[0062] in As the baseline confidence level, These are adaptive coefficients. A three-layer constraint mechanism is established, requiring the following rules to be satisfied simultaneously:
[0063] 1. Support constraint:
[0064] 2. Confidence constraint:
[0065] 3. Lift constraint:
[0066] in The corresponding threshold parameter is used. Only rules that meet all three conditions are retained, ensuring that the mined rules have statistical significance and practical meaning.
[0067] S3-4. Introduce a risk scoring function to directly quantify the risk intensity of rules:
[0068]
[0069] Risk scoring comprehensively considers the frequency and reliability of rules. High support indicates that the rule occurs frequently, while high confidence indicates that the rule is highly reliable. The product of the two reflects the overall risk level of the rule. The score provides a quantitative basis for calculating risk weights in subsequent network analysis.
[0070] S3-5. Identify high-frequency coupling patterns. Civil aviation risks are often caused by the synergistic effects of multiple factors; identifying high-frequency coupling patterns is crucial for risk management. Co-occurrence frequency of statistical factor pairs:
[0071]
[0072] Where R is the rule set. For the indicator function, when the antecedent of rule r contains both and The value is 1 if the condition is met, and 0 otherwise. It is used to count the number of times two factors co-occur in the antecedent of a rule.
[0073] Calculate the average risk of the factor pair:
[0074]
[0075] in The antecedent also includes and The set of rules, The size of this set is given. Average risk reflects the average strength of the factor at the time the risk is triggered. The core coupling patterns and their correlation strengths are output and integrated to form the global rule set R.
[0076] S4. Construct an optimized Logistic regression model and quantify the risk contribution of each contributing factor through multi-source feature fusion; including the following steps:
[0077] S4-1. An extended feature space is constructed by fusing the basic factor set F identified by BERT-LDA with the association rule set R output by Apriori through multi-source feature fusion. The basic factor set F contains all identified causal types, and each factor is treated as a binary feature, indicating whether the factor occurred in the accident.
[0078] Identifying highly collaborative factors for constructing interaction features Based on the core coupling patterns identified in step three, interaction features are constructed for factors with high co-occurrence frequency and high average risk:
[0079]
[0080] in The binary eigenvalues of the basic factors and the interaction features capture the additional risk effect that occurs when two factors occur simultaneously.
[0081] Constructing rule features Define rule triggering indicator functions for the association rules discovered by Apriori. When the incident record contains all the antecedent factors of rule k, ;otherwise The rule features reflect the triggering of multi-factor coupling patterns.
[0082] Forming an extended feature space:
[0083]
[0084] This feature space integrates single-factor information, inter-factor synergistic effects, and rule-triggered information, providing rich inputs for the model.
[0085] S4-2. Construct a Logistic Regression Model to quantify the contribution of each factor to accident risk. The decision function integrates basic features, interaction features, and rule features:
[0086]
[0087]
[0088] in Basic factors The weighting coefficient reflects the direct contribution of this factor to risk; The weighting coefficients of the interaction features reflect the strength of the synergistic effect of the factor pair; y = y * b, where b is the weight coefficient of the rule feature, reflecting the additional risk when the rule is triggered; b is the bias term. The decision value z is mapped to a risk probability using the Sigmoid function, where y = 1 indicates that an accident or symptom has occurred, and y = 0 indicates that it has not occurred. Output probability. This represents the probability of an accident occurring under a given combination of factors.
[0089] S4-3. To address the class imbalance problem commonly found in civil aviation accident data, majority class F1 optimization and L2 regularization are introduced. The loss function is defined as:
[0090]
[0091] The first term is cross-entropy loss, which measures the predicted probability. With real labels The difference is addressed by adjusting sample weights to optimize the F1 score of the majority class, preventing the model from becoming overly biased towards the majority class. The second term is the L2 regularization term, where C is the regularization parameter. These are the model parameters. Regularization controls model complexity, prevents overfitting, and improves generalization ability by penalizing excessively large parameter values. The optimal regularization parameter C is determined through grid search, eliminating class imbalance bias while maintaining generalization ability.
[0092] S4-4. Based on the optimized model, predict the probability output of single-factor risk confidence:
[0093]
[0094] in Including factors All sample sets, This is the size of the sample set. Indicator Factors When triggered, the average accident probability of the sample reflects the inherent risk level of the factor.
[0095] Output the synergistic enhancement effect value of the rules and external factors, quantifying the additional risks generated by the coupling of multiple factors:
[0096]
[0097] in, For rule r and external factors Simultaneously triggered sample set, This represents the average accident probability for this sample set. For external factors only Triggered baseline sample set, Baseline average accident probability. Enhancement effect value. A positive value indicates that the triggering of rule r significantly increases the factor. The risk level reflects the synergistic effect of multiple coupled factors. Based on this enhancement effect value, combined with the support and lift of the rules, the core risk enhancement rules are identified.
[0098] S5. Based on the optimized Logistic model output and association rules, construct a civil aviation risk causal association network to identify core risk nodes and key risk paths, and output the basis for risk management decision-making. This includes the following steps:
[0099] S5-1. Constructing an Undirected Weighted Graph The model characterizes the correlation mechanism between factors. The node set V consists of the causative factor set F, with each node representing a causative factor. The edge set E consists of factor pairs whose co-occurrence frequency exceeds a preset threshold; an edge between two factors indicates that they frequently co-occur in the incident and are correlated. The edge weight set W reflects the correlation strength and risk level of the factor pairs. The network graph transforms the raw data into a network topology, facilitating quantitative analysis using graph theory methods.
[0100] S5-2. Using the correlation model to predict probabilities and factor co-occurrence frequencies, a risk-weighted co-occurrence matrix is constructed. Traditional co-occurrence matrices only count frequencies and do not consider differences in risk levels. The risk-weighted co-occurrence matrix constructed in this invention combines frequency and risk probability:
[0101]
[0102]
[0103] in For the accident sample set, As an indicator function, when sample x simultaneously contains factors and The value is 1. This represents the accident probability for the sample. The formula sums the accident probabilities for each accident sample containing a factor pair to obtain the risk-weighted co-occurrence value. (Integrated edge weights) The first item reflects the weighted co-occurrence strength of the two factors in the accident, and the second item reflects the average risk level of each factor; it can simultaneously reflect the dual information of correlation strength and risk level.
[0104] S5-3. Calculate network topology metrics to quantify the importance of nodes in the network. Degree centrality measures the direct connectivity of a node:
[0105]
[0106] in For nodes The degree, i.e., the number of connecting edges; This represents the total number of network nodes. High degree centrality indicates that this factor is correlated with many other factors.
[0107] Tight centrality measures the degree to which a node is central in a network:
[0108]
[0109] in For nodes To the node The shortest path length. High compact centrality indicates that the node is located at the center of the network because the distances to other nodes are short.
[0110] Betweenness centrality measures the bridging role of nodes in information transmission:
[0111]
[0112] in For nodes arrive The total number of shortest paths, For the nodes The number of shortest paths. High betweenness centrality indicates that the node acts as a "bridge" connecting different groups of nodes.
[0113] By combining network topology metrics, risk confidence scores, and model coefficients, a comprehensive importance index is constructed:
[0114]
[0115] in This represents the absolute value of the coefficient of factor f in the Logistic model, reflecting the contribution of that factor to the model's predictions. The weighting coefficients are determined empirically or through optimization methods. The comprehensive importance index integrates three dimensions: network structure characteristics, risk intensity, and model importance, achieving a comprehensive quantitative assessment of node importance.
[0116] S5-4. Introduce an adaptive threshold mechanism to dynamically adjust the edge filtering threshold based on the sample size. A fixed threshold may lead to an overly sparse network when the sample size is small, or an overly dense network when the sample size is large. The dynamic threshold calculation formula is as follows:
[0117]
[0118] in For the sample set size, Let k be the minimum threshold and k be the scaling factor. The criteria for filtering edges are: Only edges with weights exceeding a dynamic threshold are retained in the network to avoid interference from noisy edges and ensure the reliability of the network structure.
[0119] S5-5, Using comprehensive importance indicators Sorting and identifying core risk nodes. Based on edge weights. Based on node importance, key risk propagation paths are identified, and critical propagation chains from source nodes to target nodes are identified. The system outputs the network topology of risk causative networks, a list of core nodes, high-risk association rules, and synergistic enhancement effect values, providing quantitative decision-making basis for formulating differentiated security management strategies and implementing precise risk prevention and control measures.
[0120] like Figure 2 As shown, the optimized model significantly improves the accuracy of risk sample identification compared to the basic model: the AUC reaches 0.916, an improvement of 1.5% over the basic model; the false positive rate (MAE) decreases to 0.067, a decrease of 67.0% over the basic model. The optimized model performs well across various metrics, enabling efficient prediction even under imbalanced data conditions, providing a reliable foundation for subsequent analysis.
Claims
1. A method for analyzing the causal network of civil aviation risks based on a semantic enhancement model, characterized in that, Includes the following steps: S1. Obtain textual data on civil aviation transport accidents and incidents, and construct a civil aviation risk event dataset; S2. Construct an enhanced BERT-LDA topic model that integrates BERT semantic representation and quality adaptive adjustment to perform deep semantic topic mining on text in the civil aviation risk event dataset and identify the set of basic causal types. S3. An improved Apriori algorithm based on an adaptive threshold mechanism for the number of factors is used to mine association rules in the basic causal type set and dynamically capture strongly associated causal rule sets. S4. Construct an optimized Logistic regression model that integrates multi-source features, and quantify the risk contribution of each causative type based on the basic causative type and the causative rule set; S5. Based on the output of the optimized Logistic regression model and the set of causal rules, construct a civil aviation risk causal association network, identify the core risk nodes and key risk propagation paths in the network, and output the basis for risk management and control decisions.
2. The method for civil aviation risk causal network correlation analysis based on semantic enhancement model according to claim 1, characterized in that, Step S1 is as follows: Load the civil aviation professional dictionary and stop word list, perform customized word segmentation, retain content words through part-of-speech tagging, and deduplicate and standardize the word segmentation results.
3. The method for civil aviation risk causal network correlation analysis based on semantic enhancement model according to claim 1, characterized in that, In step S2, the enhanced BERT-LDA topic model includes: a BERT semantic encoding module for encoding documents into document-level semantic vectors; a semantic space mapping module for mapping the continuous semantic space of BERT into a discrete topic probability distribution compatible with the LDA model; a quality adaptive weight adjustment module for dynamically adjusting the weights of the BERT semantic distribution and the LDA topic distribution during fusion based on the consistency scores of each topic; a fusion module for performing linear weighted fusion of the BERT semantic distribution and the LDA topic distribution with sparsity compensation to generate the final topic distribution; and a topic optimization module for filtering generalized words based on knowledge in the civil aviation field and enhancing the weights of professional terms, outputting a set of basic causal types.
4. The method for civil aviation risk causal network correlation analysis based on semantic enhancement model according to claim 1, characterized in that, In step S3, the improved Apriori algorithm includes: an adaptive threshold mechanism for the number of factors used to dynamically calculate the minimum support threshold based on the number of factors contained in the current mining itemset, wherein a lower support requirement is applied to multi-factor itemsets compared to single-factor itemsets; a multi-level constraint rule selection mechanism used to select strong association rules from candidate rules that simultaneously satisfy support, confidence, and lift constraints based on dynamically adjusted confidence thresholds and preset lift thresholds; and a rule risk scoring function used to calculate a comprehensive risk score for the selected association rules.
5. The method for civil aviation risk causal network correlation analysis based on semantic enhancement model according to claim 1, characterized in that, In step S4, optimizing the Logistic regression model includes: an extended feature space construction module for constructing a multi-source fusion feature space containing basic causal features, inter-causal interaction features, and association rule triggering features; a model construction and optimization module for constructing a Logistic regression model with the extended feature space as input, and introducing loss function optimization and L2 regularization to prevent overfitting; and a risk contribution quantification module for outputting single-factor risk confidence and synergistic enhancement effect values of multi-factor coupling.
6. The method for civil aviation risk causal network correlation analysis based on semantic enhancement model according to claim 1, characterized in that, In step S5, the construction and analysis of the civil aviation risk causal association network includes: a network construction module, used to construct an undirected weighted network with basic causal types as nodes and association strength calculated based on risk-weighted co-occurrence values as edge weights; a node importance comprehensive evaluation module, used to integrate network topology centrality indicators, Logistic regression model coefficients, and single-factor risk confidence scores to calculate the comprehensive importance index of each node; an adaptive edge screening module, used to dynamically adjust edge weight thresholds according to sample size to screen out key association edges; and a core node and path identification module, used to identify core risk nodes based on the comprehensive importance index and identify key risk propagation paths based on edge weights and node importance.
7. The method for civil aviation risk causal network correlation analysis based on semantic enhancement model according to claim 1, characterized in that, The final outputs are as follows: risk causal network topology, list of core risk nodes, high-risk association rules and their synergistic enhancement effect values, which are used to support the formulation of differentiated civil aviation safety management and control strategies.
8. An electronic device, characterized in that, It includes a memory and a processor, the memory storing a computer program, and the processor executing the program to implement the steps of the method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The device contains a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-7.