Mountain highway open shed tunnel safety risk intelligent assessment method based on KG-SW-EWM-FCE coupling principle

By constructing a knowledge graph and combining it with a multi-algorithm fusion model based on the KG-SW-EWM-FCE coupling principle, an intelligent assessment method is adopted to solve the problems of subjective dependence and ambiguity in traditional assessment techniques, and to achieve accurate, interpretable assessment and effective prevention and control of landslide disasters in open-air tunnels.

CN121544036APending Publication Date: 2026-02-17JIANGXI HIGHWAY RES & DESIGN INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511718413.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Traditional open-cut tunnel landslide disaster assessment techniques suffer from strong subjective dependence, blind selection of risk factors, simplistic weight calculation, fuzzy handling that does not conform to engineering practice, and lack of closed-loop optimization mechanisms. This results in large deviations between assessment results and actual disaster conditions or insufficient targeted prevention and control recommendations, making it difficult to meet the objective, accurate, and implementable engineering needs.

Method used

An intelligent assessment method based on the KG-SW-EWM-FCE coupling principle is adopted. A structured graph of entities, relationships, and attributes is constructed through knowledge graph. Combined with a multi-algorithm fusion model, data-driven risk factor screening and weight calculation are performed to build a closed-loop optimization mechanism, thereby achieving the accuracy and interpretability of the assessment results.

Benefits of technology

It improves the objectivity and accuracy of the assessment, enhances the adaptability and interpretability of the model, ensures that the assessment results are consistent with the actual risks, and makes the prevention and control measures highly targeted, which can effectively reduce the risk of landslides and collapses in open tunnels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121544036A_ABST
    Figure CN121544036A_ABST
Patent Text Reader

Abstract

The invention provides a mountain highway open shed tunnel safety risk intelligent assessment method based on a KG-SW-EWM-FCE coupling principle. The mountain highway open shed tunnel safety risk intelligent assessment method comprises the following steps: S1, full-amount acquisition of open shed tunnel anti-collapse disaster related data; s2, preprocessing and standardizing the collected data; s3, constructing an anti-collapse and anti-slip disaster knowledge graph of the open shed tunnel; s4, screening key risk factors based on the knowledge graph; s5, a fusion algorithm model based on KG-SW-EWM-FCE is constructed; s6, performing fusion model training and parameter optimization; and S7, intelligently evaluating the safety risk of collapse and slip disaster resistance of the open shed tunnel. Through the innovative design of mapping knowledge domain-multi-algorithm fusion-closed loop optimization, the core pain point of the traditional technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of knowledge graph and intelligent evaluation of open shed tunnel anti-landslide disaster, in particular to a mountainous highway open shed tunnel safety risk intelligent evaluation method based on a KG-SW-EWM-FCE coupling principle. BACKGROUND

[0002] In mountainous highways, mountainous highways have a large risk of landslide disasters due to multiple classifications and long distances, while railways have large investments (not single shed tunnels as the only defense line) and may choose to cross the line when encountering risk, so mountainous highways face a larger risk of landslide disasters. The open shed tunnel is the core protective structure for mountainous highways to resist slope collapse and landslide disasters--it directly bears the impact load when the slope is unstable, and the structural safety not only determines the continuity of the line operation, but also relates to the safety of life and property of passing vehicles and personnel. As mountainous highway facilities extend to complex geological areas, the open shed tunnel faces increasingly diverse landslide disaster scenarios, and higher requirements are placed on the objectivity, accuracy and adaptability of its anti-landslide disaster safety risk evaluation. However, the current mainstream evaluation technology still has many core problems that are not in line with engineering actual needs, which are specifically reflected in the following aspects: First, the problem of strong subjective dependence has not been solved. Traditional evaluation relies heavily on expert experience to determine index weights, such as the weight distribution of material compressive strength and slope gradient. Different experts often have different opinions due to their different understandings of disaster mechanisms--some experts based on past structural design experience may give higher weights to material strength, while experts who have been engaged in geological survey for a long time may focus more on the influence of slope gradient, resulting in a difference of 1-2 risk levels in the evaluation results of the same open shed tunnel. Subsequent prevention and control measures lack unified and reliable basis, and even contradictory conclusions may appear in multiple evaluation reports for the same project.

[0003] Second, the risk factor screening is blind and cannot be adapted to the disaster mechanism. Existing methods mostly use statistical methods to screen factors, but ignore the multi-factor coupling effect of landslide disasters. For example, in the analysis of rock and soil cohesion and slope gradient, a factor may be excluded due to its weak correlation with disaster results, but in engineering practice, the coupling of the two will significantly affect the stability of the slope--when the cohesion is less than 20 kPa and the slope is more than 50°, the probability of disaster occurrence will increase dramatically. This evaluation dimension loss caused by the screening logic deviating from the mechanism directly reduces the adaptability of the model to complex disaster scenarios.

[0004] Thirdly, the weight calculation method is relatively single, and cannot balance technical logic and data rules. Either it completely relies on subjective experience, such as historical data in a certain area shows that the annual average rainfall has a greater impact on disaster results than seismic parameters, but experts may give high weight to seismic parameters due to past earthquake disaster project experience, resulting in a disconnection between weight and actual disaster rules. Or it only relies on data statistics, only focusing on the discreteness of parameters, but ignoring the engineering common sense that the risk will suddenly change when the slope gradient exceeds 60°--such as the slope data in a certain area has low discreteness, and the entropy weight method will give it a low weight, but the slope in this area is generally close to 70°, which is actually a core risk factor. This single calculation leads to insufficient reliability of the weight.

[0005] Fourthly, the fuzzy processing of risk level does not conform to the actual engineering. The relationship between the risk level of a clear shed tunnel and the quantized value of the index is inherently fuzzy, but traditional methods often use absolute threshold division--such as setting the standardized value of 0.4 for the slope gradient as the dividing line between relatively safe and generally safe. In actual engineering, the risk difference between clear shed tunnels with slope gradients of 0.39 and 0.41 is minimal. This black-and-white division not only leads to distorted evaluation results, but also may mislead prevention and control decisions: for example, a clear shed tunnel with a slope gradient of 0.41 is determined to be generally safe and requires a large amount of investment for reinforcement, while in reality its risk is not significantly different from the relatively safe level, resulting in excessive protection.

[0006] Fifthly, there is a lack of closed-loop optimization mechanism, and the model has poor adaptability. After the traditional evaluation model is trained, the parameters are fixed, but the geological conditions and climate characteristics of clear shed tunnels in different regions differ greatly--such as the rock-soil bulk density in the southwest mountainous area is generally higher than that in the north China mountainous area, and the annual rainfall is 3-5 times that of north China. Using the same model parameters to evaluate clear shed tunnels in both regions will result in underestimation of the risk level in the southwest mountainous area and overestimation in the north China mountainous area. The model cannot be continuously optimized as the engineering scene expands, and the evaluation accuracy will gradually decrease over time.

[0007] Sixthly, the evaluation process has poor interpretability and is difficult to support engineering decisions. Traditional methods can only output the final risk level, but cannot trace the source of the risk--engineers cannot determine whether the high risk is caused by insufficient material compressive strength, excessive slope displacement, or excessive rainfall. Subsequent prevention and control measures can only be a comprehensive approach, which lacks targetedness and causes waste of manpower and resources.

[0008] These problems, when combined, make it difficult for traditional evaluation techniques to meet the current engineering needs of objective, accurate, and practical clear shed tunnel anti-landslide disaster evaluation--either the evaluation results deviate greatly from the actual disaster situation, leading to safety accidents or excessive protection; or the prevention and control recommendations lack targetedness and cannot effectively reduce the risk. Therefore, there is an urgent need for an intelligent evaluation method that can integrate disaster mechanisms, data rules, and engineering practices, and has objectivity, accuracy, and interpretability, to solve the pain points of traditional techniques. SUMMARY

[0009] The application provides a mountainous highway open shed tunnel safety risk intelligent evaluation method based on a KG-SW-EWM-FCE coupling principle, solves the core pain points of traditional technologies through innovative design of knowledge graph-multi-algorithm fusion-closed loop optimization, and achieves the above purposes by adopting the following technical solutions. The mountainous highway open shed tunnel safety risk intelligent evaluation method based on the KG-SW-EWM-FCE coupling principle comprises the following steps: S1: Collecting all the structure parameters of the open shed tunnel, the landslide disaster environment parameters and the historical disaster data to form an original data set, wherein the historical disaster data comprises the corresponding relationship between the structure parameters, the environment parameters and the disaster results, and the disaster results are no damage, slight damage, moderate damage, severe damage or complete destruction; S2: Performing cleaning, completion and standardization processing on the original data set to form a standardized data set, wherein the standardized data set comprises a structure parameter standardized subset, an environment parameter standardized subset and a historical disaster standardized subset, the historical disaster standardized subset comprises the corresponding relationship between the key factor standardized data and the disaster result quantitative value, and the disaster result quantitative value is obtained by converting the disaster result; S3: Based on the standardized data set, a knowledge graph comprising an open shed tunnel entity, a landslide disaster entity, a risk factor entity, an evaluation index entity and a disaster result entity is constructed, the knowledge graph comprises the attributes of each entity and the association relationship between the entities, and the knowledge graph is outputted; S4: Based on the association relationship of the knowledge graph and the historical disaster standardized subset, the association strength between each risk factor and the disaster result quantitative value is calculated, the redundant factors are removed, and the key risk factors and the corresponding key factor standardized data are selected; S5: Based on the key risk factors, an evaluation index system is constructed, a fusion algorithm model comprising a KG (Knowledge Graph)-SW (Semantic Weight) data weight model, an EWM (Entropy Weight Method) objective weight model and a FCE (Fuzzy Comprehensive Evaluation) fuzzy comprehensive evaluation model is constructed, the data weight of the KG-SW data weight model is calculated based on the semantic association of the knowledge graph and the correlation coefficient between the key factors and the disaster result quantitative value, the objective weight of the EWM objective weight model is calculated based on the discreteness of the key factor standardized data, the combined weight is obtained by balancing the data weight and the objective weight, the risk grade evaluation is realized based on the combined weight and the membership function through the FCE fuzzy comprehensive evaluation model, and the fusion algorithm model is outputted; S6: Using the historical disaster standardized subset as training data, optimizing the balance coefficient and membership function parameters of the fusion algorithm model to obtain an optimized fusion model, and outputting the optimized fusion model; S7: Collecting the original data of key risk factors of the open shed tunnel to be evaluated, obtaining the standardized data of the key factors to be evaluated through the preprocessing process of step S2, inputting the data into the optimized fusion model, and obtaining the open shed tunnel anti-landslide disaster safety risk evaluation result, which includes the risk level and the comprehensive membership vector, and the risk level corresponds to the disaster result one by one; S8: Setting up monitoring points for the evaluated open shed tunnel, collecting structural deformation, slope displacement and real-time environmental parameters to form monitoring data, or collecting actual damage data of the evaluated open shed tunnel after landslide disaster to form actual disaster result data; converting the monitoring data into actual risk levels according to the specification limits, or directly corresponding the actual damage degree to the actual risk level, and the actual risk level constitutes the verification data; comparing the risk level in the evaluation result with the verification data, and determining the samples inconsistent with each other as deviation samples, which include the standardized data of the key factors to be evaluated and the corresponding actual risk level; supplementing the deviation samples to the training data of step S6, and re-executing step S6 to obtain an updated fusion model.

[0010] In the specification, the fusion process of the data weight and the objective weight in step S5 is: combining the KG-SW data weight and the EWM objective weight through a linear weighting formula, the balance coefficient in the linear weighting formula ranges from 0 to 1, normalizing the result after fusion to obtain a combined weight that satisfies the sum of weights being 1, and the combined weight is directly input into the FCE fuzzy comprehensive evaluation model and interacts with the membership function.

[0011] In the specification, the interaction process of the FCE fuzzy comprehensive evaluation model and the combined weight in step S5 is: inputting the standardized data of the key factors to be evaluated into the calibrated trapezoidal membership function to obtain the membership of each index to different risk levels and form a fuzzy evaluation matrix, weighting and summing the combined weight and the fuzzy evaluation matrix through fuzzy matrix multiplication to obtain a comprehensive membership vector, and determining the final risk level based on the maximum membership principle.

[0012] In the specification, the screening process of the key risk factors in step S4 is: setting the correlation strength threshold to 0.3, screening out candidate key factors with a correlation strength not lower than the threshold; calculating the redundancy between the candidate key factors by using the mutual information method, setting the redundancy threshold to 0.8, removing factors with a redundancy higher than the redundancy threshold, and retaining factors with higher correlation strength to obtain a final set of key risk factors.

[0013] In this specification, steps S6, S7, and S8 constitute the closed-loop optimization process of the model. The closed-loop optimization process is as follows: optimize the parameters of the fusion model using standardized historical disaster data; verify the accuracy of the assessment using validation data converted from monitoring data or actual disaster result data; supplement the biased samples into the training data to repeat model optimization, thereby achieving continuous iterative updates of the fusion model. The closing condition of the closed-loop optimization process is: the assessment accuracy corresponding to the validation data is not less than 85%, the prediction accuracy for extremely dangerous and relatively dangerous levels is not less than 90%, the change in accuracy during model iteration does not exceed 1%, and the mean square error is stable below 0.1.

[0014] In this specification, the construction process of the KG-SW data weight model in step S5 is as follows: based on the knowledge graph, the number of associations between each key risk factor and the disaster result entity is counted, and the semantic importance is calculated; combined with the absolute value of the Pearson correlation coefficient of the standardized data of the key factors and the quantitative value of the disaster result, the comprehensive importance is obtained; the comprehensive importance is normalized to obtain the data weight corresponding to each key risk factor.

[0015] In this manual, the method for converting monitoring data into actual risk level in step S8 is as follows: if the structural deformation and slope displacement exceed the limits specified in the highway tunnel design code, the risk level of the corresponding assessment result is upgraded by one level; if all monitoring data are within the limits specified in the code, the actual risk level is determined to be consistent with the risk level of the assessment result.

[0016] In this manual, the composition of the deviation sample in step S8 is clearly defined as follows: standardized data of key factors of the tunnel to be evaluated, risk level in the evaluation results, and actual risk level in the validation data. These three correspond to form a complete deviation sample, which is then added to the training data to re-optimize the model parameters. Key risk factors include material compressive strength, slope gradient, soil-rock cohesion, average annual rainfall, foundation depth, peak ground acceleration, buffer layer thickness, and soil-rock unit weight.

[0017] In summary, this invention offers at least the following advantages: Significantly enhanced objectivity in assessment: Based on the semantic associations of the knowledge graph and historical data, a data-driven KG-SW data weighting model is constructed to replace the traditional expert scoring method, eliminating subjective bias and providing clear data support and semantic basis for weight calculation. Precise key factor selection: Combining the semantic associations of risk factors and disaster outcomes in the knowledge graph with statistical correlations, the selected key factors not only conform to the disaster action mechanism but also effectively distinguish risk levels, avoiding factor redundancy or omission. Significantly enhanced weight reliability: Through the fusion calculation driven by KG-SW semantic associations and EWM data discreteness, both semantic logic and data patterns are considered. The combined weights reflect both the qualitative judgment of indicator importance and the objective characteristics of the data, resulting in stronger robustness. More realistic fuzzy handling: Employing the FCE fuzzy comprehensive evaluation model, the trapezoidal membership function quantifies the fuzzy correspondence between indicators and risk levels, avoiding absolute classification and improving the accuracy and rationality of risk level assessment. Adaptive and Continuous Model Optimization: A closed-loop mechanism of evaluation, verification, feedback, and optimization is constructed. Model parameters are continuously updated using on-site monitoring data and actual disaster results, enhancing the model's adaptability to different regions and types of open-cut tunnels. Evaluation accuracy continuously improves as application scenarios expand. Highly Interpretable Evaluation Process: The knowledge graph clearly depicts the semantic relationships between entities, relationships, and attributes. Evaluation results can be traced back to the weight percentage and membership contribution of specific key factors, providing clear logical support for the formulation of prevention and control measures and improving the scientific nature of engineering decisions. Highly Practical for Engineering: The evaluation process is highly standardized and automated, with smooth integration from data collection to report generation. Early warning levels are clearly defined, and prevention and control recommendations are highly targeted, allowing for direct application in engineering practice and effectively reducing the risk of landslides and collapses in open-cut tunnels. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of an intelligent assessment method for safety risks of open-air tunnels on mountainous highways based on the coupling principle of KG-SW-EWM-FCE.

[0019] Figure 2 This is a schematic diagram of the entire evaluation process involved in this invention.

[0020] Figure 3 This is a schematic diagram of the closed-loop optimization of the model involved in this invention.

[0021] Figure 4 This is a schematic diagram of the interaction of the fusion algorithm involved in this invention. Detailed Implementation

[0022] like Figure 1 and Figure 2As shown, this embodiment provides an intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle, including: S1: Full collection of data related to landslide and collapse disasters in the open tunnel; focusing on three core dimensions—the tunnel's own structure, landslide and collapse disaster environment, and historical disasters—a full collection of data covering the original parameters required by the S5 evaluation index system is conducted to provide a basic data source for subsequent analysis.

[0023] 1. Data type collected: Structural parameters of the open-air tunnel: span L (m) Net height H (m) Vertical support structure cross-sectional dimensions d ( ), radius of curvature of the vault r (m) Material compressive strength (MPa), tensile strength of material (MPa), tensile strength of steel bars (MPa), reinforcement ratio (%) Foundation depth h (m), type of material for the tunnel roof buffer layer (sand / pebbles / EPS / rubber), and thickness of the buffer layer t (m). Among these, , h t directly corresponds to the S5 secondary indicator , , The remaining parameters provide a candidate basis for subsequent key factor screening. Landslide disaster environmental parameters: slope gradient. (°), slope height (m), Soil and rock type (cohesive soil / sandy soil / gravelly soil / rock), Soil and rock unit weight (kN / ), Angle of friction within the soil and rock (°), Cohesion of soil and rock mass c (kPa), average annual rainfall P (mm), rainfall intensity I (mm / h), peak ground acceleration a (g) Depth of groundwater level (m). Among them, Directly corresponds to the S5 secondary indicator The remaining parameters are used to construct a complete disaster environment data dimension to support semantic association analysis. Historical disaster data: the occurrence time of landslides and mudslides in the region over the past 50 years, disaster type (collapse / landslide / debris flow / clastic flow), disaster scale (volume V, ), Scope of influence R(m) The degree of damage to the open-air tunnel (no damage / minor damage / moderate damage / severe damage / complete destruction), and the complete environmental parameters and structural parameters of the open-air tunnel at the time of the disaster. This type of data must ensure a one-to-one correspondence between structural parameters, environmental parameters, and disaster results, providing core training samples for semantic association calculation of the KG-SW model in S5, entropy analysis of the EWM model, and membership function calibration of the FCE model.

[0024] 2. Data Sources and Acquisition: Structural parameters of the open-plan tunnel: Original data should be extracted primarily from the tunnel's design drawings and construction archives, such as material testing reports and support construction records. This ensures the reliability of the parameters. For missing key parameters (such as material compressive strength)... The thickness of the buffer layer (t) was supplemented by on-site measurements—the material strength was measured using a rebound hammer (accuracy ±0.1MPa). Measurement points were selected based on a uniform distribution principle, with at least 30 samples for each parameter, and the average value was taken as the final data. Landslide environmental parameters: slope gradient. Slope height A digital elevation model (DEM) is generated through UAV aerial surveying and modeling (resolution 0.05m / pixel), and then accurate data is extracted using GIS tools; soil and rock parameters ( The data was obtained through borehole sampling tests, with a borehole depth of no less than 1 / 3 of the slope height. Three parallel tests were conducted at each sampling point, and the average value was used to eliminate experimental errors. Rainfall... P Rainfall intensity I Statistical calculations were performed based on nearly 30 years of daily monitoring data from regional meteorological stations to ensure data continuity over time; peak ground acceleration (PGA) was measured. a Directly reference parameters for the corresponding region from the National Seismic Zoning Map (GB18306-2015); groundwater level depth. Water level observation wells were deployed on-site (one every 100m), and stable values ​​were obtained after 30 consecutive days of monitoring. Historical disaster data: systematically collected from the natural resources department's disaster survey database, relevant academic literature, and engineering survey reports. The authenticity of each record was verified using the data traceability principle, and disaster cases without clear parameter support were eliminated. The final historical dataset must contain at least 100 complete samples to meet the sample size requirements for training the S5 fusion algorithm model.

[0025] 3. Data Output: Original Data Set ,in This is a structural parameter dataset (containing the original values ​​of the parameters and measured / extracted annotations). This is an environmental parameter dataset (containing the original values ​​of the parameters and the annotations of how they were obtained). This is a historical disaster dataset (containing a complete record of parameter-result correspondences). This dataset will be directly fed into S2 for preprocessing to ensure that the data format and accuracy meet the requirements of subsequent knowledge graph construction and S5 algorithm modeling.

[0026] S2: Data preprocessing and standardization; The raw data output from S1 undergoes a three-stage process of cleaning, completion, and standardization to eliminate data noise, fill in missing values, and standardize data scale, generating a high-quality standardized dataset. This dataset not only provides standardized entity attribute data for S3 knowledge graph construction, but also directly adapts to the semantic association calculation of the KG-SW model, the entropy analysis of the EWM model, and the membership function input of the FCE model in S5, ensuring the accuracy and efficiency of subsequent algorithm calculations.

[0027] 1. Data cleaning: Outlier removal: using 3... σ Criteria (applicable to normally distributed data) for identifying outliers -- First, calculate the mean of each parameter. and standard deviation It will exceed Data within a specified range is marked as potentially anomaly. For suspected anomaly data, manual verification is performed based on the data source (e.g., checking construction records and measurement logs). Anomalies confirmed to be caused by measurement errors or recording mistakes are directly removed. For reasonable anomalies caused by special disaster scenarios (e.g., extreme rainfall), the special scenario data is retained and labeled to provide comprehensive scenario coverage for S5 model training. Duplicate data removal: A dual deduplication method of parameter fingerprinting and result matching is used. Records with identical structural and environmental parameters and the same disaster results are considered duplicate data, and only one complete record is retained. Records with similar parameters (difference ≤ 5%) but identical results are merged into one record, and the average parameter value is taken as the unified value to avoid duplicate data affecting the efficiency and accuracy of S5 model training.

[0028] 2. Data Completion: Low Missing Rate Parameter Completion (Missing Rate ≤ 5%): Using... K Nearest Neighbor (KNN) interpolation method K The value is determined based on the sample size (when the sample size is ≥100). K =5, when sample size <100 K =3). Based on the dataset to which the missing parameter belongs (e.g., ...). Based on this, calculate the Euclidean distance between the missing sample and other samples (based on all complete parameters), and select the closest sample. KFor each sample, the mean value of this parameter is used as the missing value to ensure that the completed data conforms to the overall data pattern. For high missing rate parameter completion (missing rate > 5%): reasonable estimation is performed based on open-cut tunnel design specifications (such as JTG / TD30-2015 Highway Tunnel Design Specification) and regional geological survey reports—for example, the missing soil-rock cohesion. c The estimated value should be adjusted by referring to the average cohesion value of soil and rock masses of the same region and type, and combined with the slope stability analysis results. The estimated data should be clearly marked with an estimation label and recorded separately in the subsequent S5 model training to analyze the impact of the estimated data on the model accuracy.

[0029] 3. Data Standardization: Quantitative Parameter Standardization: For all secondary indicators in the S5 evaluation indicator system, the corresponding quantitative parameters (such as...) , (e.g., h), after min-max standardization, the conversion formula is: ;in, Here, x represents the standardized parameter value (range [0,1]), and x represents the original parameter value. For this parameter in The minimum value in, For this parameter in The maximum value in the standardization is the maximum value in the standardization process. The core purpose of standardization is to eliminate the differences in parameter dimensions (such as the difference between MPa and °, m), so that all parameters can directly participate in the weight calculation of KG-SW, the entropy analysis of EWM, and the membership function mapping of FCE in S5. Qualitative parameter standardization: For a small number of qualitative parameters (such as the material type of the cave roof buffer layer), unique thermal coding is used to ensure that they can be integrated into the quantitative calculation process. For example, the material type of the cave roof buffer layer is divided into sandy soil coding [0,1,0,0,0], pebble coding [0,0,1,0,0], EPS coding [0,0,0,1,0], rubber coding [0,0,0,0,1], and the soil and rock type is divided into cohesive soil coding [1,0,0,0], sandy soil coding [0,1,0,0], gravelly soil coding [0,0,1,0], and rock coding [0,0,0,1]. The coding results are directly used as entity attribute data for the construction of the S3 knowledge graph.

[0030] 4. Data Output: Standardized Dataset ,in Includes standardized values ​​of structural parameters (quantitative parameters are values ​​in the [0,1] interval, and qualitative parameters are uniquely encoded). Includes standardized values ​​of environmental parameters. This dataset contains standardized parameters for historical disaster samples and corresponding quantitative values ​​of disaster outcomes (no damage = 0, minor damage = 1, moderate damage = 2, severe damage = 3, complete destruction = 4). This dataset will be directly input into S3 for knowledge graph construction. As the core data foundation for training the S5 fusion algorithm model, it ensures that the data can be seamlessly connected to subsequent steps.

[0031] S3: Construction of a knowledge graph for landslide and collapse prevention in the Mingpeng Tunnel; based on the standardized dataset output from S2. Construct a structured knowledge graph containing entities, relationships, and attributes. G This map accurately depicts the semantic relationships between open-cut tunnels, landslide hazards, risk factors, assessment indicators, and disaster outcomes. It not only provides semantic support for the selection of key risk factors in S4, but also directly provides core data for calculating the semantic importance of the KG-SW model in S5 and analyzing the correlation strength between risk factors and disaster outcomes. It is a crucial link in achieving a data-driven and semantically supported integrated assessment.

[0032] 1. Knowledge Graph Ontology Definition: Core Entity Type (Strongly Correlated with S5 Indicator System): Mingpengdong Entity Landslide disaster entities Risk factor entities Evaluation Indicators Entities Disaster Result Entity Among them, risk factor entities The attributes completely correspond to the original parameters in S1, and the evaluation index entity The attributes correspond one-to-one with the first and second-level indicators of S5, ensuring a seamless connection between the ontology definition and subsequent evaluation logic. Entity attribute definition: Attributes: Direct association Standardized parameters in, such as { Buffer layer material code}, the attribute value is the result after S2 normalization (quantitative parameter [0,1] interval value, qualitative parameter unique thermal code); Attributes: Direct association Standardized parameters in, such as Attribute value format and Consistent; Attributes: Includes standardized values ​​and classification labels (structural / environmental) of all potential risk factors, such as { (Structure class, ), (Environmental category) ), ...}; Attributes: Fully corresponds to the S5 evaluation indicator system; the primary indicator attributes are { (Structure class) (Environmental category)}, secondary indicator attributes are ; Attributes: Includes disaster outcome quantification value (0-4), disaster type, and standardized disaster scale value. The attribute value comes directly from Extracted from [the relevant source]. Entity relationship definition (supporting S5 semantic association calculation): and Inclusion relationship: This indicates that the structural parameters of the tunnel belong to the structural risk factors, such as ( ,Include, factor); and Derivative relationships: Parameters representing landslide environments are derived into environmental risk factors, such as ( ,derivative, factor); and The correspondence: indicates that the risk factor directly corresponds to the S5 assessment indicator, such as ( Factor, Correspondence index); and Impact relationship: Indicates that the assessment indicators have a significant impact on the disaster outcome, such as ( Indicators, impacts, severe damage outcomes); and The bearing relationship: indicates the result of the collapse and landslide disaster that the tunnel bears, such as ( (Suffering, moderate damage result).

[0033] 2. Entity and Relation Extraction (ensuring extraction results meet S5 model requirements): Entity Extraction: Employing a BERT-based Named Entity Recognition (NER) model, input... All parameter names, numerical labels, and classification information (such as material compressive strength) are included. (Structured class) The model outputs an accurate set of entities. During the extraction process, special attention should be paid to ensuring the entities corresponding to the S5 secondary indicators (such as...). The key evaluation dimensions are accurately identified, avoiding omissions. Relationship extraction: An attention-based relationship extraction model is used, with input entity pairs and corresponding contextual data (such as...). Indicators (slope) (This involves) historical data association records with severe damage outcomes, and the model automatically identifies and outputs the relationships between entities. For influence relationships, the association strength (such as the number of times a certain indicator is associated with a certain outcome) needs to be recorded simultaneously to provide direct data support for the S5KG-SW model to calculate semantic importance.

[0034] 3. Knowledge Graph Construction and Quality Verification (Ensuring the graph's support for S5): Graph Storage: The knowledge graph is stored using the graph database Neo4j, with triples (head entity, relation, tail entity) as the basic storage unit, for example ( Indicators, Corresponding to factor),( The system includes indicators, impacts, and severe damage outcomes. Attributes (such as association frequency and data source) are added to each triple for easy S5 model invocation. Quality verification includes: Logical consistency check: Irrational relationships are eliminated using a custom rule base (e.g., structural indicators cannot derive environmental factors, and assessment indicators must correspond to at least one risk factor) to ensure the semantic logic of the graph is consistent with the S5 assessment system; Data integrity check: The system verifies whether the entities corresponding to the eight secondary indicators of S5 have sufficient impact relationship data (each indicator is associated with no fewer than 20 disaster outcomes) to ensure it meets the requirements for semantic importance calculation in the S5KG-SW model; Data accuracy check: 10% of the triples are randomly selected and compared with… The original data is compared to ensure that the relationship description is consistent with the data facts, and the accuracy rate must reach more than 95%.

[0035] 4. Data Output: Knowledge Graph of Landslide and Collapse Resistance in Mingpeng Tunnel G =( E , R , A E is the set of entities (containing 5 types of core entities and sub-entities corresponding to the S5 indicator). R It is a set of relationships (including 5 core relationship types and association strength attributes). A For entity attribute set (and (The standardized parameters correspond one-to-one). This graph is directly input into S4 for key risk factor screening, and also serves as the core input of the KG-SW model in S5, providing structured support for semantic importance calculation and correlation analysis between risk factors and disaster outcomes.

[0036] S4: Key risk factor screening based on knowledge graph; knowledge graph constructed using S3. G For semantic support, combined with the output of S2 Data was analyzed using a two-step method of correlation strength calculation and redundancy removal to identify key factors that significantly impact the safety risk of landslides and collapses in open-cut tunnels from the initial risk factor set. The screening results are directly used as the core content of the S5 evaluation indicator system, which not only ensures the relevance and effectiveness of the indicator system, but also reduces the computational dimension of the S5 fusion algorithm and improves evaluation efficiency.

[0037] 1. Initial risk factor set generation (comprehensive coverage of potential influencing factors): from S3's knowledge graph G Extracting risk factor entities All corresponding attributes form the initial risk factor set. Based on the collected data from S1, the initial number of factors... n =21 (11 structure classes:) L , H , d, r , , , h , , 1. Type of material for the tunnel roof buffer layer; t; 10 environmental categories: , Types of rock and soil , c, P, I, a ), to ensure that no potential influencing factors are overlooked.

[0038] 2. Factor Association Strength Calculation (Combining semantics and data to quantify the degree of influence): Association Strength It is a core indicator for measuring the impact of risk factors on disaster outcomes. Taking into account both the semantic relationships within the knowledge graph and the statistical correlations of historical data, the formula is: ; : No. i The correlation strength of the initial risk factors, with values ​​ranging from [0,1], indicates that the larger the value, the more significant the impact on the disaster outcome; The first in the S3 knowledge graph i Individual risk factors and disaster outcome entities Number of influence relationships (e.g.) The number of times the factor is associated with outcomes such as minor injury and severe injury. All initial risk factors and The total number of influence relationships (i.e.) ); : No. i Standardized values ​​of each risk factor (from The Pearson correlation coefficient between the disaster outcome quantification value Y(0-4) and the disaster outcome quantification value Y(0-4) ranges from -1 to 1. A larger absolute value indicates a stronger statistical correlation; a positive value is used here (because the degree of influence is unrelated to the direction of the correlation). Calculation process: First, from the knowledge graph of S3... G Extract the number of influence relationships for each factor, and then based on... The Pearson correlation coefficient was calculated from the data, and the correlation strength of each initial factor was finally obtained through a formula. This ensures that the calculation results simultaneously reflect semantic association patterns and data statistical patterns.

[0039] 3. Key Factor Screening (Eliminating Redundancy and Focusing on Core Indicators): Step 1: Setting a Threshold for Correlation Strength .based on Statistical analysis of the data, combined with the computational efficiency requirements of the S5 fusion algorithm, determines --Filter out The factors form a set of candidate key factors. By using this threshold screening, factors with a negligible impact on the disaster outcome (such as the radius of curvature r of the arch and the depth of the groundwater level) are initially eliminated. This reduces subsequent computational load. The second step is redundancy removal. The mutual information method is used to calculate the redundancy among candidate factors. A higher mutual information value indicates a higher degree of information overlap (stronger redundancy) between factors. A redundancy threshold is set. If two candidate factors Then retain the correlation strength. Larger factors are used to avoid information redundancy that could distort the S5 model's evaluation results. Final selection results: After two steps of selection, a set of key risk factors was obtained. There are a total of 8 factors, which correspond exactly to the 8 secondary indicators of S5 -- 3 of which are structural: (Material compressive strength) (Foundation depth) (Buffer layer thickness); 5 environmental categories: (Slope) (Cohesion of soil and rock mass) (Annual average rainfall) Peak ground acceleration (PGA) (Unit weight of rock and soil).

[0040] 4. Data Output: Set of Key Risk Factors (Includes the names and classification labels of 8 factors) and standardized data for each factor. .in From S2 The standardized values ​​of the corresponding factors are extracted, and the data format is a range of [0,1] values ​​(quantitative factors) or one-hot encoding (qualitative factors; in this scheme, all 8 key factors are quantitative factors). This output is directly fed into S5 as the core data for constructing the evaluation index system, calculating KG-SW data weights and EWM objective weights, ensuring that the index system of S5 is completely consistent with the screening results.

[0041] S5: Construction of a fusion algorithm model based on KG-SW-EWM-FCE; A three-level fusion algorithm system is constructed, consisting of a knowledge graph-based semantic association data weight model (KG-SW), an entropy weight method objective weight model (EWM), and a fuzzy comprehensive evaluation model (FCE). KG-SW mines the semantic associations between risk factors and disaster outcomes in the knowledge graph to generate data-driven data weights; EWM calculates objective weights based on the discreteness of historical data, reflecting the information value of the data itself; by combining the weights and integrating the advantages of both, and then combining FCE to handle the fuzziness and uncertainty in risk assessment, a precise, objective, and interpretable assessment of the safety risk of landslides and collapses in open-cut tunnels is ultimately achieved. This fusion model ensures the data support for weight calculations and solves the fuzziness problem in risk level classification, serving as the core computational unit of the entire assessment method.

[0042] 5.1 Construction of the evaluation indicator system (based on the set of key risk factors) The construction of an evaluation indicator system is the foundation of risk assessment. It must strictly adhere to the principles of clear hierarchy, full factor coverage, and close logical connections, and rely entirely on the set of key risk factors output by S4. This ensures that the indicators correspond one-to-one with the key factors.

[0043] 1. Logic of indicator hierarchy: First-level indicators: Based on the attribute category of key risk factors, they are divided into structural key factor indicators. Key environmental factors and indicators The classification is based on the following: structural factors directly determine the tunnel's resistance to landslides, while environmental factors are external triggers for landslide disasters. Both types of factors work together to influence the disaster outcome, and this hierarchical classification aligns with the internal-external cause mechanism of disaster action. Secondary indicators: These fully inherit the eight key risk factors output from S4, with each secondary indicator strictly corresponding to one key risk factor to ensure the completeness and relevance of the indicator system.

[0044] 2. Indicator System (Quantitative Indicators): Primary Indicators Structural key factor indicators: secondary indicators Material compressive strength index, corresponding key risk factors (Material compressive strength), source S2 output middle Level 2 Foundation burial depth index, corresponding key risk factors =h (foundation depth), sourced from S2 output. middle Secondary indicators The thickness of the buffer layer corresponds to key risk factors. (Buffer layer thickness), source S2 output middle Primary indicators Key environmental factors and indicators: secondary indicators Slope gradient index, corresponding to key risk factors (Slope gradient), source: S2 output middle Secondary indicators The cohesion index of soil and rock mass corresponds to key risk factors. =c (cohesion of soil and rock mass), sourced from the output of S2. middle Secondary indicators Annual average rainfall index, corresponding to key risk factors =P (average annual rainfall), source: output of S2 middle Secondary indicators Peak ground acceleration (PGA) index, corresponding to key risk factors =a (peak ground acceleration), sourced from S2 output. middle Secondary indicators The unit weight index of soil and rock mass, corresponding to key risk factors. (Unit weight of soil and rock mass), sourced from S2 output. middle .

[0045] 3. Core function of the indicator system: to clarify the specific dimensions of the assessment object, transform the abstract risk of landslide and collapse into eight quantifiable and calculable specific indicators, provide clear calculation objects for subsequent weight calculation and fuzzy evaluation, and ensure the pertinence and operability of the assessment process.

[0046] 5.2KG-SW (Semantic Relationship Data Weighting Model Based on Knowledge Graph) 1. Core Model Positioning: To replace the traditional expert scoring method and construct a data-driven data weight calculation model. The traditional expert scoring method suffers from subjective bias and individual experience differences, while the KG-SW model generates objective data weights based on the semantic relationship between risk factors and disaster outcomes in the knowledge graph constructed in S3 and the quantitative analysis results in S4. This retains the qualitative judgment logic of data weights on the importance of factors and eliminates human bias through data support.

[0047] 2. Model Construction Principle: The core logic of the model is the importance of semantic association strength as a determinant of factors—in the knowledge graph, this is related to the disaster outcome entity. The closer the association and the greater the semantic contribution of a risk factor, the higher its importance to risk assessment, and the greater its corresponding data weight. The construction process relies on the triplet data (head entity-relationship-tail entity) of the knowledge graph to extract risk factor entities. -Impact-Disaster Result Entities The correlation is used to calculate the weights by combining the results of quantitative analysis.

[0048] 3. Detailed Model Training Process: Step 1: Relationship Extraction and Quantification. Input the knowledge graph G=(E,R,A) of S3, traverse all triples of risk factor entity-impact-disaster outcome entity, and statistically analyze each key risk factor. (correspond (attributes) and Number of direct relationships --The higher the number, the more frequently the factor influences disaster outcomes in historical data and semantic logic. Simultaneously, the Pearson correlation coefficients between the key factors already calculated in S4 and disaster outcomes are extracted. (Y represents the quantified disaster outcome: no damage = 0, minor damage = 1, moderate damage = 2, severe damage = 3, complete destruction = 4). This coefficient reflects the degree of linear correlation between the factor and the disaster outcome, ranging from -1 to 1. The larger the absolute value, the stronger the correlation. Second step: Semantic importance. Calculation. Semantic importance is an indicator that measures the coreness of a factor in a knowledge graph, equal to the sum of the individual key factors and their relative importance. The number of associations, divided by the total number of key factors and The total number of associations is calculated using the following formula: ; : No. i The semantic importance of each key factor, with a value range of [0,1]. The closer it is to 1, the more central the factor is in semantic association. The first in the S3 knowledge graph i Key factors and disaster outcome entities The number of influence relationships (e.g., material compressive strength) (Number of associations with disaster outcomes such as severe damage and complete destruction). All 8 key factors and The total number of influence relationships (i.e.) Step 3: Training data calibration. This involves calibrating the semantic importance... Correlation coefficient with Pearson Perform linear fusion to generate a semantic-data dual-dimensional importance index. (The absolute value is taken because the sign of the correlation coefficient only indicates the direction of the association and does not affect the magnitude of the importance.) This ensures that the trained model considers both semantic association and historical data patterns.

[0049] 4. Model Application (Data Weight Calculation): Based on the training results KG-SW data weights are obtained through normalization. The formula is: ; : No. i The KG-SW data weights of each secondary indicator, with values ​​ranging from [0,1], satisfy the following conditions: .

[0050] 5. Core role and contribution: It replaces the traditional expert scoring method, and generates objective and interpretable data weights through semantic association of knowledge graphs and quantitative analysis of historical data, avoiding human experience bias; The weight calculation process is strongly correlated with the S3 knowledge graph and the S4 key factor screening results, ensuring the logical coherence of the entire evaluation method. At the same time, the semantic association frequency and data correlation strength are considered so that the data weights not only conform to the disaster action mechanism, but also fit the actual data patterns.

[0051] 5.3EWM (Entropy Weight Method Objective Weight Model) 1. Core Model Positioning: Based on information entropy theory, this model calculates objective weights by analyzing the dispersion of standardized data for key risk factors. It does not rely on subjective judgment but is entirely determined by the informational value of the data itself—the greater the data dispersion, the more significant the difference in the factor across different disaster scenarios, the higher its discriminative power for risk assessment, and the greater its corresponding objective weight.

[0052] 2. Model Construction Principle: Information entropy is an indicator that measures the uncertainty of data. If all sample data for a certain indicator tend to be consistent (low dispersion), its information entropy is high, indicating that the indicator provides less effective information and should have a smaller weight. Conversely, if the data has a high degree of dispersion (significant differences in sample values), its information entropy is low, it provides more effective information, and should have a larger weight. The model calculates the information entropy of each indicator, converts it into a difference coefficient, and then normalizes it to obtain the objective weights.

[0053] 3. Detailed Model Training Process: Step 1: Input Data Preparation. Input the standardized dataset of key factors from the S4 output. (8 indicators, each containing N=100 historical sample data points, where N is the number of historical data points in S2.) The sample size was chosen to be 100 records to ensure the stability of the entropy calculation. The data format is as follows: A matrix (N=100, m=8) is denoted as ,in Indicates the first j The first sample i Standardized values ​​of key factors ( i=1,2,...,8; j=1,2,...,100). Step 2: Probability distribution calculation. Since the standardized data ranges from [0,1], directly calculating the entropy value will result in... The data is meaningless, so we first perform a data shift (to ensure...). >0), then calculate the first i Probability distribution of the j-th sample of each indicator The formula is: ; In the EWM model, the first i The first indicator j The probability value of each sample, taking values ​​in the range (0,1), satisfies... ; Minimum value (take) ), used to avoid =0 leads to calculation anomalies. Step 3: Entropy calibration. Verify the rationality of the probability distribution using historical sample data. If a certain indicator... If all values ​​tend to be 1 / N (i.e., the data is highly concentrated), then the indicator is marked as a low information value indicator, and its difference coefficient should be given special attention in subsequent weight calculations.

[0054] 4. Model Application (Objective Weight Calculation): The objective weight calculation is completed in three steps, each relying on the probability distribution data obtained from training: Step 1: Calculate the information entropy of each indicator. : ; In the EWM model, the first i The information entropy of each indicator ranges from [0,1]. When all... When they are equal ( ), =1, at which point the indicator has no discriminating power; when a certain sample... =1, and all others are 0. =0, at which point the indicator's discriminative power is strongest. Second step: Calculate the coefficient of difference. : ; : No. i The difference coefficients of the indicators range from [0,1]. The larger the value, the greater the dispersion of the indicator data, and the more effective information it provides. Step 3: Normalization to obtain objective weights. : ; : No. i The objective weights of the EWM for each secondary indicator, with values ​​ranging from [0,1], satisfy the following conditions: .

[0055] 5. Core Role and Contribution: Weights are calculated entirely based on the discrete nature of historical data, reflecting the objective laws inherent in the data itself, and complementing the semantic association data weights of KG-SW; it can identify the indicators that contribute most to distinguishing risk levels (such as...). of The largest value indicates the greatest dispersion of the slope gradient, and thus the highest discrimination for risk assessment.

[0056] 5.4 Combined Weight Fusion (Interactive Fusion of KG-SW and EWM) 1. Fusion Logic: KG-SW weights reflect the objective importance of semantic associations and data relevance, while EWM weights reflect the objective importance of data dispersion. Both have limitations when used alone (e.g., KG-SW may be affected by the number of associations in the knowledge graph, and EWM may be affected by outlier samples). By combining and fusing these weights, a triple coupling of semantic associations, data features, and objective laws can be achieved, improving the reliability and robustness of the weights.

[0057] 2. Fusion Process (Direct Interaction between KG-SW and EWM): Step 1: Introducing the Balancing Coefficient . Used to adjust the proportion of KG-SW data weights and EWM objective weights, with a value range of... : Only KG-SW weights are used at this time. Only EWM weights are used, with initial values ​​set to (This indicates that in the initial state, data weights and objective weights are equally important). Step 2: Combined weight calculation. The two types of weights are combined using a linear weighting formula to obtain the preliminary combined weights. : ; : No. i The initial combined weights of the secondary indicators, with values ​​ranging from [0,1]; : Balance coefficient (initial value 0.5, subsequently optimized through S6 model training). Third step: Normalization. Since the sum of the initial combined weights may be affected by... If the values ​​deviate from 1 due to different values, normalization calibration is required to obtain the final combined weights. : ; : No. i The final combined weights of the secondary indicators satisfy... =1 is the core input parameter of the FCE model.

[0058] 4. Core Role and Contribution: Achieving deep interaction between KG-SW and EWM through the balancing coefficient. The proportions of the two types of weights are adjusted to avoid the limitations of a single weight; the combined weights retain the semantic association logic of the knowledge graph while incorporating the objective discrete characteristics of the data, providing a more reliable weight input for the subsequent FCE model; normalization ensures that the weights meet mathematical constraints, guaranteeing the calculation effectiveness of the subsequent fuzzy comprehensive evaluation.

[0059] 5.5FCE (Fuzzy Comprehensive Evaluation Model) 1. Core Model Positioning: To resolve the fuzzy correspondence between quantitative indicators and risk levels in risk assessment. The classification of landslide and collapse risk levels in open-pit tunnels is not absolutely clear (e.g., a slope of 0.4 may be considered either relatively safe or generally safe). FCE constructs a fuzzy membership function to transform quantitative indicators into the degree of membership to each risk level, and then combines combined weights for comprehensive evaluation, ultimately achieving accurate risk level classification.

[0060] 2. Model Construction Principle: Based on fuzzy set theory, each secondary indicator is regarded as a fuzzy subset, and each risk level is regarded as an element of the fuzzy set. The degree of the indicator value belonging to a certain risk level (between 0 and 1) is described by the membership function. Then, the membership degree of the indicator and the combined weight are fused by fuzzy matrix multiplication to obtain the comprehensive membership vector. Finally, the risk level is determined based on the principle of maximum membership degree.

[0061] 3. Detailed Model Training Process: Step 1: Determine the Risk Assessment Level Set Based on the actual impact of landslides and collapses on the open-air tunnel, a five-level risk level is defined: ; Extremely safe (the structure of the open tunnel is intact, the probability of landslides is less than 5%, and there is no risk of damage); Relatively safe (the open-air tunnel structure is stable, with a 5%–20% probability of landslides, and only minor damage is possible). Generally safe (the structure of the open-air tunnel is basically stable, with a 20% to 50% probability of landslides, which may result in moderate damage). : Relatively dangerous (the open-air tunnel structure has insufficient stability, and the probability of collapse and landslide disasters is 50% to 80%, which may result in serious damage); Extremely dangerous (the open-air tunnel structure is on the verge of instability, with a landslide probability >80%, potentially leading to complete destruction). Step 2: Construct the membership function. A trapezoidal membership function is used (balancing computational simplicity and practical relevance), based on S2. The correspondence between standardized values ​​of key factors and disaster outcomes is established, and the inflection point of the membership function for each secondary indicator is calibrated. The following are the complete membership functions (standardized values) for all eight secondary indicators. ): (1) (Material compressive strength; the higher the standardized value, the safer it is): ; ; ; ; ; (2) (For foundation burial depth, a larger standardized value indicates greater safety): ; ; ; ; ; (3) (Buffer layer thickness; the larger the standardized value, the safer): ; ; ; ; ; (4) (Slope gradient; the larger the standardized value, the more dangerous it is): ; ; ; ; ; (5) (Cohesion of soil and rock mass; the higher the standardized value, the safer it is): ; ; ; ; ; (6) (Annual average rainfall; the higher the standardized value, the more dangerous it is): ; ; ; ; ; (7) (Peak ground acceleration; the larger the normalized value, the more dangerous it is): ; ; ; ; ; (8) (The higher the standardized value of the unit weight of soil and rock mass, the more dangerous it is): ; ; ; ; ; The inflection points of all membership functions pass through S2. Data calibration (e.g., standardizing the range of each indicator corresponding to the statistical extreme safety level, determining the inflection point position) ensures that the function closely matches the actual disaster patterns. Third step: Membership function verification. Input the standardized values ​​of key factors from 100 historical samples into the membership function, calculate the membership degree of each sample to each risk level, and compare it with the actual disaster results (actual risk level). If the consistency rate between the risk level corresponding to the maximum membership degree and the actual level is ≥80%, the function is valid; otherwise, adjust the inflection point position and repeat the verification until the requirements are met.

[0062] 4. Model Application: The fuzzy comprehensive evaluation is completed in three steps. The core is to perform matrix interaction calculations of the combined weights and fuzzy membership degrees: Step 1: Construct the fuzzy evaluation matrix R. For the Mingpengdong Cave to be evaluated, standardize the values ​​of its eight secondary indicators. (From S7) Input the trained membership function to obtain the membership degree of each indicator to the five risk levels. ( i =1~8, k =1~5), forming an 8×5 fuzzy evaluation matrix: Among them, the elements of the fuzzy evaluation matrix Indicates the first i The second-level indicator belongs to the first kThe membership degree of each risk level is determined, with values ​​ranging from [0,1]. The second step is fuzzy matrix multiplication (the interaction between combined weights and membership degrees). The combined weight vector... Performing fuzzy multiplication with the fuzzy evaluation matrix R yields the comprehensive evaluation vector B: ; "Indicates fuzzy matrix multiplication (weighted summation operation); As a comprehensive evaluation vector, The Mingpengdong Cave, to be evaluated, belongs to the category of... k The comprehensive membership degree of each risk level satisfies (Since the combined weights have been normalized). Step 3: Determine the risk level. The maximum membership principle is adopted: find the risk level corresponding to the maximum value in the comprehensive evaluation vector B, which is the final risk level of the tunnel to be evaluated. If two or more membership degrees are equal and both are the maximum values, the lower risk level is taken to ensure that the evaluation result is conservative and conforms to the principle of engineering safety.

[0063] 5. Core Role and Contribution: Addressing the ambiguity in risk assessment by quantifying the fuzzy correspondence between quantitative indicators and qualitative risk levels, avoiding absolute judgments of black-and-white nature; deeply interacting with combined weights, integrating indicator importance and membership degrees through fuzzy matrix multiplication to ensure assessment results consider both indicator weights and ambiguity; outputting clear risk levels based on the principle of maximum membership degree while retaining the comprehensive membership vector, providing a quantitative basis for subsequent risk warnings and prevention recommendations (e.g., although relatively safe, the risk of transitioning to general safety needs to be monitored). The fusion algorithm interaction is as follows... Figure 4 As shown.

[0064] 5.6 Data Output; The fusion algorithm model M=(KG-SW,EWM,FCE) contains the following core outputs, which are directly passed to S6 for training and optimization: 1. KG-SW model output: data weights of 8 secondary indicators. semantic importance Pearson correlation coefficient 2. EWM Model Output: Objective weights of 8 secondary indicators Information entropy Coefficient of difference 3. Combination weight formula: (Including initial balance coefficient) ); 4. FCE model output: complete trapezoidal membership functions for 8 secondary indicators (including all inflection point parameters), and a set of risk assessment levels. Fuzzy matrix multiplication formula and the maximum membership degree principle determination rule.

[0065] S6: Fusion Model Training and Parameter Optimization; Based on the fusion algorithm model M=(KG-SW,EWM,FCE) constructed in S5, the output of S2 is used... and S4 output Using the training data, the model parameters are iteratively optimized through a process of balancing coefficient optimization, membership function calibration, and model validation to obtain an optimized model with higher accuracy and stronger robustness. This model will be directly used for intelligent evaluation of S7, ensuring that the accuracy of the evaluation results meets the actual needs of engineering projects.

[0066] 1. Training dataset construction: Input S2, output (Includes over 100 complete historical disaster samples) and S4 output (Standardized values ​​of 8 key factors), training samples are extracted according to the correspondence between key factors and disaster outcomes. The structure of each sample is as follows: ,in The standardized values ​​of the 8 key factors for the j-th historical sample (j=1,2,...,N, where N is the number of samples). The actual risk level of this sample (based on disaster outcome quantification conversion: 0 → 1→ 、2→ 3→ 4→ Dataset partitioning: The dataset is divided into training and validation sets in an 8:2 ratio. The training set contains 80% of the samples (for optimizing model parameters), and the validation set contains 20% of the samples (for validating model accuracy). The partitioning process uses stratified sampling to ensure that the proportion of samples of each risk level is consistent in the training and validation sets, thus avoiding overfitting of the model due to uneven sample distribution.

[0067] 2. Model training process: Step 1: Balancing coefficients Optimization. Balance coefficient This is a key parameter for S5 combined weight fusion, used to adjust the proportion of KG-SW data weights and EWM objective weights. Optimization is achieved using a grid search method. The search range is set to [0,1], with a step size of 0.05, resulting in 21 candidate values. For each candidate... The training set samples are input into the fusion model M of S5 to calculate the predicted risk level. And calculate the accuracy of the training set. (Number of correctly predicted samples / Total number of training samples); Select to make The largest As the optimal balance coefficient For example, when When the training set accuracy reaches its highest level (92%), then... Step 2: FCE membership function calibration. (Fixed) Optimize the trapezoidal membership function parameters (i.e., inflection point values) of the FCE model in S5. Employ cross-validation (5-fold cross-validation): divide the training set into 5 subsets, using 4 subsets to train the model and 1 subset to validate it, calculating the mean squared error (MSE) for each validation. ;in, To verify the sample size of the subset, For the first k Quantitative value of predicted risk level for each sample ( =0, =1, =2, =3, =4), This is the actual quantized value of the sample. This is achieved by adjusting the inflection point value of the membership function (e.g., ...). of The inflection point of the ranking was adjusted from 0.2 / 0.6 to 0.25 / 0.55, and the MSE was minimized to ensure that the membership function better fits the historical data patterns. The third step is model iterative optimization. The above two steps are repeated. After each optimization, the accuracy of the training set and the MSE of cross-validation are calculated. When the accuracy no longer improves (change ≤ 1%) and the MSE stabilizes below 0.1, the iteration is stopped, and a preliminary optimized model is obtained.

[0068] 3. Model Validation: Input the validation set samples into the initially optimized model and calculate the validation set accuracy. (Number of correctly predicted samples / Total number of validation samples). Set validation criteria: If the criterion is met, the model optimization is effective; if not, the optimization strategy should be readjusted (e.g., expanding the...). Adjusting the search step size and membership function type), repeating the training process until the validation set accuracy meets the target. Additional validation: Using a confusion matrix to analyze the model's predictive performance at various risk levels, with a focus on high-risk levels (…). , The prediction accuracy rate (≥90%) should be improved to avoid engineering safety hazards caused by misjudgment of high-risk levels.

[0069] 4. Data Output: Optimized Fusion Model It includes the core content: optimal balance coefficient. (e.g., 0.55), optimized KG-SW data weights (based on The reverse calibration shows that the original KG-SW weights are determined by the knowledge graph association strength × semantic importance. During optimization, the weighting coefficients of association strength in the KG-SW calculation are adjusted based on feedback (e.g., fine-tuning the weighting percentage of the association strength formula in S4), ultimately outputting... Ensure that its weights are consistent with EWM weights. After fusion, the model evaluation error is minimized (e.g., reducing risk level prediction bias) and the optimized EWM objective weights are obtained. (exist The calibration results of the original EWM weights under constraints; the original EWM weights are calculated from the entropy (dispersion) of the indicator data. During optimization, the sample weights used in the entropy calculation are adjusted (e.g., assigning higher weights to samples from recent disasters), so that... and pass After fusion, the data better reflects the correspondence between data characteristics and risk outcomes in actual disasters, avoiding weight distortion caused by historical data distribution biases; the calibrated FCE membership function (including all inflection point parameters); and the final combined weights. (Calculated using the combined weight formula of S5). This model is directly input into S7 for intelligent risk assessment of landslide and collapse hazards in open-cut tunnels, ensuring that the assessment accuracy meets the actual engineering requirements.

[0070] S7: Intelligent assessment of safety risks of landslides and collapses in open-air tunnels; The key risk factor data of the open-air tunnel to be assessed are processed according to the preprocessing standards of S2, and then input into the optimized fusion model of S6. The risk level of the tunnel to be evaluated is calculated using the fuzzy comprehensive evaluation process defined in S5. and the comprehensive membership vector The evaluation process relies entirely on the optimization algorithm of S5 and the precise parameters of S6 to ensure the objectivity, accuracy, and interpretability of the evaluation results, providing a core basis for subsequent verification feedback and report generation.

[0071] 1. Data Collection and Processing to be Evaluated: Data Collection: The set of key risk factors output by S4. (8 factors) Collect raw data of the open-air tunnel to be evaluated. --The data collection method is consistent with S1 (prioritizing the extraction of design / construction archive data, and conducting on-site measurements for missing parameters) to ensure the accuracy and reliability of the original data. For example, collecting material compressive strength data. At the same time, the on-site measurement sample size shall not be less than 30, and the average value shall be taken as the raw data; slope gradient shall be collected. At that time, precise values ​​were extracted through UAV aerial survey modeling. Data preprocessing for evaluation: The process strictly followed the S2 cleaning-complete-standardization workflow. Cleaning: using 3 σThe criteria include: removing outliers to ensure the data is free of significant errors; data completion: for missing values ​​(KNN interpolation for missing values ​​≤5%, and estimation using normed methods for missing values ​​>5%); standardization: quantitative parameters are standardized using min-max to convert them to values ​​in the [0,1] interval, using the same formula as S2. The standardized data of the key factors to be evaluated are then output. Data format and S4 Completely identical, ensuring that the fusion model can be directly input into S5.

[0072] 2. Model Application and Risk Calculation: Step 1: Constructing the Fuzzy Evaluation Matrix .Will Each standardized value in (i=1~8), input the optimized S6 Membership function: Calculates the membership degree of each secondary indicator to the five risk levels. ( k =1~5, corresponding to For example, the slope of the slope to be evaluated. =0.5, input optimized The membership function of the index is obtained. =0、 =0.5、 =0.5、 =0、 =0, that is =0、 =0.5、 =0.5、 =0、 =0. After calculating the membership degrees of all indicators, an 8×5 fuzzy evaluation matrix is ​​formed. Step 2: Calculate the comprehensive evaluation vector Calling the output of S6 (Final combined weights), according to the fuzzy matrix multiplication formula in S5, combine the weight vector with the fuzzy evaluation matrix. Perform interactive computing: ; For the first i The optimal combination weights of the secondary indicators, for The elements, the calculation result ,satisfy , The Mingpengdong Cave, to be evaluated, belongs to the category of... k The comprehensive membership degree of each risk level. Step 3: Determine the risk level. Using the maximum membership principle defined in S5: find... The risk level corresponding to the maximum value is the final risk level of the tunnel to be assessed. If multiple maximum values ​​occur (e.g....), If so, then the lower risk level (e.g.) will be selected. (To ensure relative safety), the assessment results are conservative and in line with the principle of prioritizing engineering safety.

[0073] 3. Data Output: Risk assessment results of the open-air tunnel to be evaluated. ,in For the final risk level ( Extremely safe~ Extremely dangerous) This is a comprehensive membership vector (containing membership values ​​at each level). This is the standardized result of the data to be evaluated (used for subsequent verification and traceability). This result is directly passed to S8 for verification feedback.

[0074] S8: Assessment Result Verification and Model Feedback Optimization; Verify the accuracy of the S7 assessment results using on-site monitoring data or actual disaster data, forming a closed-loop mechanism of assessment-verification-optimization. If the assessment results deviate from the actual situation, supplement the training set with verification data, re-optimize the S5 fusion model, continuously improve the model's assessment accuracy and adaptability, and ensure that subsequent risk assessment results for open-pit tunnels are more consistent with engineering realities. The model closed-loop optimization process is as follows: Figure 3 As shown.

[0075] 1. Verification Data Acquisition: On-site monitoring data collection: For the completed open-air tunnels, a long-term monitoring system will be deployed for continuous monitoring for one year. Monitoring indicators are directly related to key factors in S5, such as: structural deformation of the open-air tunnel (e.g., wall settlement, arch displacement, accuracy ±0.01mm), slope displacement (horizontal and vertical displacement, accuracy ±0.01mm), and surrounding environmental parameters (real-time rainfall, slope stress, accuracy ±0.1mm and ±0.1kPa respectively). The monitoring frequency is once every 24 hours daily, and increased to once every hour during extreme weather (heavy rain, earthquakes) to ensure the capture of key change data. Actual disaster result data collection: If a landslide occurs around the open-air tunnel after the assessment, the actual impact data of the disaster will be collected, including the degree of damage to the open-air tunnel (determined by on-site investigation: no damage / minor damage / moderate damage / severe damage / complete destruction), and environmental parameters at the time of the disaster (e.g., rainfall, ground acceleration), forming actual disaster result data. (Convert to risk level: e.g., moderate injury →) Data processing for verification: Transforming monitoring data into verification indicators—for example, structural deformation exceeding the standard limit (e.g., wall settlement > 5mm) corresponds to an upgrade in risk level (e.g., ...). → The assessment results are valid if the slope displacement is stable (<2mm / year). A final validation dataset is then generated. = , monitoring indicator data, actual parameter data}, where The actual risk level serves as the core verification basis.

[0076] 2. Verification of the accuracy of evaluation results: Consistency determination: The evaluation results output by S7 will be verified. With verification data Compare them; if they match (e.g.) and If the results are inconsistent (e.g., ...), then the evaluation results are considered valid; if they are inconsistent (e.g., ...), then the evaluation results are considered valid. but If the deviation is not found, it is considered a deviation, and the deviation type (underestimation / overestimation) is recorded. Deviation cause analysis: For deviation cases, the causes are traced from three dimensions: Data level: Verify S7's... Are there any data acquisition errors (such as insufficient accuracy of measured parameters) or preprocessing errors (such as those used during standardization)? , (Unreasonable); Model level: Analyze whether the parameters of the S5 fusion model are suitable for the characteristics of the open-pit tunnels in this area, such as whether the membership function inflection point does not consider the characteristics of the regional soil and rock mass; Environmental level: Are there any unforeseen factors that were not considered during the assessment, such as extreme rainstorms exceeding the range of historical data? Consistency statistics: Statistically analyze the verification results of all assessed open-pit tunnels and calculate the consistency rate. = Number of consistent samples / Total number of validation samples, with a consistency rate of ≥90% required; if the consistency rate is <90%, the model feedback optimization process will be initiated.

[0077] 3. Model Feedback Optimization: Training Set Supplementation: Adding biased samples... The corresponding relationships are added to the training dataset T of S6 to expand the sample coverage of the training set, especially for samples from special scenarios, ensuring that the model can learn more patterns from real-world scenarios. Model re-optimization: Based on the supplemented training set, the model training process of S6 is re-executed—the balancing coefficients are re-optimized. 1. Calibrate the FCE membership function parameters and update the combined weights. The updated fusion model is obtained. Model replacement and validation: Replace the original This will be used for subsequent risk assessment of open-pit tunnels; at the same time, 10 new validation samples will be selected to test the evaluation accuracy of the updated model, ensuring that the accuracy is improved by ≥5%, and it will be officially put into use after the validation is passed.

[0078] 4. Data Output: Verification Report and the updated fusion model, in which , Includes consistency statistics, deviation cause analysis, and monitoring data summary. It includes all optimized parameters. It will serve as the core model for subsequent S7 assessments, continuously improving assessment accuracy.

Claims

1. A method for intelligent safety risk assessment of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle, characterized in that, include: S1: Collect the structural parameters of the open-air tunnel itself, the environmental parameters of landslide disasters, and historical disaster data to form a raw dataset. The historical disaster data includes the correspondence between structural parameters, environmental parameters, and disaster results. S2: Clean, complete, and standardize the original dataset to form a standardized dataset; S3: Based on a standardized dataset, construct a knowledge graph containing entities such as open-cut tunnels, landslide disasters, risk factors, assessment indicators, and disaster results. The knowledge graph includes the attributes of each entity and the relationships between entities. S4: Based on the association relationships of the knowledge graph and the standardized subset of historical disasters in the standardized dataset, calculate the correlation strength between each risk factor and the disaster outcome, remove redundant factors, and select key risk factors and their corresponding standardized data. S5: Construct an evaluation index system based on key risk factors, and build a fusion algorithm model that includes a knowledge graph-based data weight calculation model, an entropy weight method objective weight model, and a fuzzy comprehensive evaluation model. The knowledge graph-based data weight calculation model calculates data weights based on the semantic association of the knowledge graph and the correlation coefficient between key factors and disaster results. The entropy weight method objective weight model calculates objective weights based on the discreteness of standardized data of key factors. The combined weights are obtained by merging data weights and objective weights through a balance coefficient. The fuzzy comprehensive evaluation model realizes risk level evaluation based on the combined weights and membership functions. S6: Using standardized subsets of historical disasters and standardized data of key factors as training data, optimize the balance coefficients and membership function parameters of the fusion algorithm model to obtain the optimized fusion model; S7: Collect the raw data of the key risk factors of the tunnel to be evaluated. After the preprocessing process in step S2, obtain the standardized data of the key factors to be evaluated. Input the standardized data into the optimized fusion model to obtain the safety risk assessment results of the tunnel against landslide disasters. The assessment results include the risk level and the comprehensive membership vector.

2. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 1, is characterized in that... S8: For the assessed open-air tunnel, set up monitoring points to collect structural deformation, slope displacement, and real-time environmental parameters to form monitoring data, or collect data on the actual damage level of landslides that occur after assessment to form actual disaster result data; convert the monitoring data into actual risk levels according to the standard limits, or directly correspond the actual damage level to the actual risk level, and the actual risk level constitutes the verification data; compare the risk level in the assessment results with the verification data, and determine the samples that are inconsistent between the two as deviation samples, which contain standardized data of the key factors to be assessed and the corresponding actual risk level. The biased samples are added to the training data in step S6, and step S6 is repeated to obtain the updated fusion model.

3. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 1, is characterized in that... The process of fusion of data weights and objective weights in step S5 is as follows: the data weights and objective weights are combined through a linear weighting formula. The balance coefficient in the linear weighting formula ranges from 0 to 1. After fusion, the result is normalized to obtain a combined weight that satisfies the condition that the weights sum to 1. The combined weights are directly input into the fuzzy comprehensive evaluation model and interact with the membership function.

4. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 1, characterized in that, The interaction process between the fuzzy comprehensive evaluation model and the combined weights in step S5 is as follows: the standardized data of the key factors to be evaluated are input into the calibrated trapezoidal membership function to obtain the membership degree of each indicator to different risk levels and form a fuzzy evaluation matrix. The combined weights and the fuzzy evaluation matrix are weighted and summed by fuzzy matrix multiplication to obtain the comprehensive membership vector. The final risk level is determined based on the principle of maximum membership degree.

5. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 1, characterized in that, The key risk factor screening process in step S4 is as follows: set the correlation strength threshold to 0.3, screen out candidate key factors with correlation strength not lower than the threshold; use the mutual information method to calculate the redundancy between candidate key factors, set the redundancy threshold to 0.8, remove factors with redundancy higher than the redundancy threshold, retain factors with higher correlation strength, and obtain the final set of key risk factors.

6. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 2, is characterized in that... Steps S6, S7, and S8 constitute the closed-loop optimization process of the model. The closed-loop optimization process is as follows: optimize the parameters of the fusion model through standardized historical disaster data, verify the accuracy of the evaluation by using the verification data converted from monitoring data or actual disaster result data, supplement the deviation samples into the training data to repeat the model optimization, and realize the continuous iterative update of the fusion model.

7. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 6, is characterized in that... The closing condition for the closed-loop optimization process is: the evaluation accuracy corresponding to the validation data is not less than 85%, the prediction accuracy of extremely dangerous and relatively dangerous levels is not less than 90%, the change in accuracy during model iteration does not exceed 1%, and the mean square error is stable below 0.

1.

8. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 1, characterized in that, The construction process of the knowledge graph-based data weight calculation model in step S5 is as follows: based on the knowledge graph, the number of associations between each key risk factor and disaster outcome entity is counted, and the semantic importance is calculated; combined with the Pearson correlation coefficient absolute value of the standardized data of key factors and the quantitative value of disaster outcome, the comprehensive importance is obtained. The overall importance is normalized to obtain the data weights corresponding to each key risk factor.

9. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 1, characterized in that, In step S5, the membership function of the fuzzy comprehensive evaluation model is a trapezoidal membership function. Based on the correspondence between the risk level and the standardized value of key factors in the standardized historical disaster data, the inflection point parameter of the membership function is determined. The inflection point parameter is calibrated by cross-validation to ensure that the membership function conforms to the actual disaster pattern.

10. The intelligent assessment method for safety risks of open-air tunnels in mountainous highways based on the KG-SW-EWM-FCE coupling principle as described in claim 2, characterized in that, The deviation sample in step S8 consists of: standardized data of key factors of the tunnel to be evaluated, risk level in the evaluation results, and actual risk level in the validation data. These three correspond to form a complete deviation sample, which is then added to the training data and the model parameters are re-optimized.