A multi-source weighted training set construction method for deformation prediction of engineering tunnels during tunneling

Through the multi-source weighted training set construction method, tunnel data of multiple sources and different characteristics are processed, and structured training sets are generated, which solves the problem of low data utilization efficiency in the existing technology and realizes efficient tunnel deformation prediction.

CN115374570BActive Publication Date: 2025-06-13BEIJING URBAN CONSTRUCTION DESIGN & DEVELOPMENT GROUP CO LIMITED
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211229156.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-29
Filing Date
2022-10-09
Publication Date
2025-06-13
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize data from multiple sources and with different characteristics, resulting in machine learning algorithms that cannot simply assume that each learning case has the same value in tunnel deformation prediction.

Method used

Through the multi-source weighted training set construction method, design data and tunnel deformation are obtained, divided into numerical and non-numerical data, One-Hot encoding and weight assignment are performed, and combined with information source weights and adaptive weights, a training set that can be directly used in machine learning algorithms is generated.

Benefits of technology

The structured processing of multi-source data is realized, which avoids the addedability loss of non-numerical data, enhances the availability of the training set, and provides efficient prediction support for machine learning algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115374570B_ABST
    Figure CN115374570B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a multi-source weighted training set for predicting the deformation of a tunneling project, including: Step S1, obtaining the design data of the tunneling project and the corresponding tunnel deformation; Step S2, processing the design data to obtain an initial input set DatasetV1; Step S3, determining the information source weights according to the source of the design data; Step S4, determining the adaptability weights of different data according to the values of the key features of the data; Step S5, combining the weights in Steps S3 and S4 and using them as the features of the input set; Step S6, using the input set in Step S5 as the input training set and the tunnel deformation value as the output training set to form a training set. By structuring and weighting the multi-source engineering information, the present invention obtains a training set that can reflect the characteristics of the tunneling project and is conducive to the use of machine learning algorithms, effectively supporting the intelligent prediction of tunnel deformation and providing a quantitative basis for relevant project evaluations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tunnel crossing engineering, and particularly relates to a method for constructing a multi-source weighted training set for predicting tunnel deformation in crossing engineering. Background Art

[0002] With the continuous development of computer processing power and intelligent algorithms, machine learning has gradually become an effective means for the industrial community to solve practical problems. Supervised learning in machine learning methods can directly establish the potential relationship between input and output based on existing cases, thus avoiding complex intermediate links, and can simply, quickly, and accurately predict the required quantity. It is a general tool for realizing intelligence and automation in various fields. And this mode of supervised learning highly depends on the training set used by the learning algorithm. Using a structured and clearly characterized training set can help the learning algorithm achieve the learning goal more efficiently.

[0003] For the tunnel deformation of crossing engineering, a large number of engineering practices and academic researches have accumulated a lot of valuable data, providing excellent conditions for machine learning to solve the deformation prediction problem. However, these data come from diverse sources, including data with different characteristics obtained by various research methods such as on-site measurement, numerical simulation, theoretical analysis, and model tests. At the same time, there are also many unstructured data such as design drawings and cloud maps. These data are difficult to be directly utilized by relevant learning algorithms. On the other hand, the credibility of diverse information sources varies, and at the same time, deformation prediction is highly correlated with the attributes of the project itself. The degree of fit of cases has an important impact on prediction. Therefore, when applying machine learning algorithms in this field, it cannot be simply considered that each learning case has the same value, but the training set should be adjusted according to requirements and prior knowledge.

[0004] Therefore, a training set generation method that can start from existing data and adapt to the usage requirements of relevant learning algorithms is needed to solve the defects existing in the prior art. Summary of the Invention

[0005] In order to solve the above technical problems, the present invention provides a method for constructing a multi-source weighted training set for predicting tunnel deformation in crossing engineering, aiming to optimize and modify the feature content in the design data, reshape the multi-source data from the perspective of feature engineering, and perform weight processing on the reshaped data according to the characteristics of crossing engineering to obtain a training set that can be directly applied to machine learning algorithms, providing strong support for efficiently predicting tunnel deformation. In addition, a method and system for predicting tunnel deformation in crossing engineering based on machine learning algorithms are also provided.

[0006] In order to achieve the above object, the technical solutions adopted by the present invention include:

[0007] A multi-source weighted training set construction method for predicting the deformation of a tunnel in a crossing project, characterized by comprising the following steps:

[0008] Step S1: Obtain the design data of the crossing project and the corresponding tunnel deformation. The design data is the input part of the training set, including engineering geological information, tunnel structure information, and construction information. The tunnel deformation is the output part of the training set;

[0009] Step S2: According to the characteristics of the design data in Step S1, divide it into numerical data and non-numerical data. Form a data set ASet from the numerical data, process the non-numerical data to obtain a data set Bset, and combine Aset and Bset to obtain an initial input set DatasetV1;

[0010] Step S3: According to the source of the design data in the initial input set DatasetV1, adopt a certain measurement strategy to determine the information source weight;

[0011] Step S4: According to the value-taking situation of the key features of the design data in the initial input set DatasetV1, determine the adaptation weights of different data;

[0012] Step S5: Combine the weights in Steps S3 and S4 and use them as the features of the input set to obtain a weighted input set DatasetV2;

[0013] Step S6: Use the weighted input set DatasetV2 in Step S5 as the input training set and the tunnel deformation value as the output training set to form a standard training set TrainDataset.

[0014] According to the implementation scheme of the present invention, in Step S2, the processing of non-numerical data includes encoding the non-numerical data using the One-Hot method. The encoding process includes the following steps:

[0015] Step S2.1: Screen the features that are non-numerical data and form a data set FeatureSet. The total number of features in the data set FeatureSet is denoted as m, and the data set corresponding to each feature is denoted as FeatureSet i (i = 1, 2,..., m);

[0016] Step S2.3: For each feature data set FeatureSet i , divide out n i state features, and according to the value taken by each case, obtain a data set ConditionSet i under the state feature representation, where n i is the number of all discrete values under this feature; and

[0017] Step S2.4: Combine the datasets ConditionSet under m state feature representations i to obtain the non-numerical data state feature representation dataset Bset.

[0018] According to the embodiments of the present invention, the design data sources in step S3 include theoretical analysis, numerical simulation, model tests, and on-site monitoring.

[0019] According to the embodiments of the present invention, the information source weight in step S3 is an index for measuring the importance of different data sources. The specific measurement strategies may include:

[0020] (1) Expert scoring method: directly determine the scores of each data source according to expert evaluations, so that the data of the same source have the same weight;

[0021] (2) Interval random generation method: stipulate the weight intervals of each data source, and randomly generate weights for each case within the weight intervals according to its source;

[0022] (3) Ranking and comparison method: rank the importance of data sources, draw multiple cases from the dataset without replacement and arrange them according to importance, and assign weights from high to low according to the ranking until all cases are drawn; and

[0023] (4) Comprehensive method: comprehensively obtain the final weights by using the weights obtained by two or more methods respectively.

[0024] According to the embodiments of the present invention, the adaptability weight in step S4 is to measure the degree of closeness between the data and the typical working conditions, including the following steps:

[0025] Step S4.1: Determine the important judgment feature values y i (i = 1, 2,..., n), where n is the number of important judgment features (i.e., key features);

[0026] Step S4.2: Calculate the Jaccard distance D of the discrete features in the important judgment features 1 , and the formula is:

[0027]

[0028] In the above formula, x and y are respectively the sets of discrete feature values of the case and the typical working condition;

[0029] Step S4.3: Calculate the relative Euclidean distance D of the continuous features in the important judgment features 2 , and the formula is:

[0030]

[0031] In the above formula, x i and y i are the values of the case and the typical working condition for the i-th continuous feature respectively, and l is the number of continuous features among the important judgment features; and

[0032] Step S4.4, combining the distances D 1 and D 2 , calculate the adaptability weight of each case, and the formula is:

[0033] ω fit = f(D 1 ) + g(D 2 )

[0034] In the formula, both f(x) and g(x) are distance conversion functions, and an inverse proportional function with a positive inverse proportional coefficient, a linear function with a negative slope and a positive intercept, or other monotonically decreasing functions with a positive value range can be selected.

[0035] According to the implementation scheme of the present invention, the weight combination in step S5 is the comprehensive weight of the case obtained through combination strategies such as scaling and distribution according to the information source weight and adaptability weight of each case. The specific combination strategies include:

[0036] (1) Cumulative addition method, stipulate the proportion of the item weight in the combined weight, and take the result of cumulative addition according to the proportion as the combined weight;

[0037] (2) Cumulative multiplication method, multiply all item weights, and take the product result as the combined weight;

[0038] (3) Minimum selection method, select the minimum value among all item weights as the combined weight;

[0039] (4) Maximum selection method, select the maximum value among all item weights as the combined weight; and

[0040] (5) Random method, take the minimum value among all item weights as the lower bound of the interval, the maximum value as the upper bound of the interval, and randomly generate a weight within the interval as the combined weight.

[0041] In addition, the present invention also provides a deformation prediction method and system for a cross-engineering tunnel based on a machine learning algorithm.

[0042] Compared with the prior art, the present invention has the following advantages:

[0043] (1) Starting from readily available engineering data, the present invention uses digital coding conversion means to make multi-source data have a unified structured format, and performs One-hot processing on non-numerical features, avoiding the problem of additivity loss of non-numerical data, greatly enhancing the usability of the training set, thus supporting the use of relevant machine learning algorithms and providing a basis for rapid prediction.

[0044] (2) In view of the characteristics of the crossing project, the present invention performs weighted processing on the samples in the training set on the premise of meeting the requirements of the learning algorithm, thereby incorporating the key characteristics of the crossing project into the data level in a targeted manner, providing a large number of key learning features for the relevant learning algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 FIG. is a schematic flow chart of a method for constructing a multi-source weighted training set for tunnel deformation prediction in a crossing project according to an embodiment of the present invention; and

[0046] Figure 2 FIG. is a schematic diagram of the generation process of the initial input set DatasetV1 in the method for constructing a multi-source weighted training set for tunnel deformation prediction in a crossing project according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be further described in detail below with reference to the drawings and through specific embodiments. The content shown is used to fully elaborate the content of the present invention and does not limit the present invention.

[0048] Taking the prediction of the vertical deformation of an existing tunnel under crossing construction as an example, the method for constructing a multi-source weighted training set for tunnel deformation prediction in a crossing project according to the present invention will be described in detail. As shown in the attached Figure 1 The solution of the present invention may include the following steps:

[0049] Step S1, obtaining the design data of the crossing project and the corresponding tunnel deformation. The design data is the input part of the training set and may include engineering geological information, tunnel structure information, construction information, etc. The tunnel deformation is the output part of the training set:

[0050] For this embodiment, multi-channel data such as engineering design books, publicly published academic literatures, and relevant monitoring reports are collected. After sorting and screening, a total of 187 design data cases and the corresponding tunnel vertical deformation values are obtained.

[0051] Step S2, classifying the design data in step S1 into numerical data and non-numerical data according to the characteristics of the design data, forming a data set ASet for the numerical data, processing the non-numerical data to obtain a data set BSet, and combining ASet and BSet to obtain an initial input set DatasetV1 (see the attachedFigure 2 ):

[0052] Based on the case information, the common or convertible design data features are summarized, including 30 engineering geological features, 10 tunnel structure features, and 5 construction information features, a total of 45 features. The above data is divided into numerical and non-numerical data:

[0053] The steps of forming the data set ASet from numerical data include screening the features that are numerical data and forming the data set ASet. In the data of this embodiment, 42 features are numerical data, which together form the data set ASet.

[0054] The One-Hot method is used to encode non-numerical data, and the encoding process includes:

[0055] Step S2.1, screening the features that are non-numerical data and forming the data set FeatureSet. The total number of features in the data set FeatureSet is denoted as m, and the data set corresponding to each feature is denoted as FeatureSet i (i = 1, 2, ……, m). In the data of this embodiment, there are 3 features of "lining type", "construction method", and "control measures" that are non-numerical data, thus forming the FeatureSet data set with the total number of features m = 3. The data sets of the three features respectively correspond to FeatureSet 1 、FeatureSet 2 、FeatureSet 3 , as shown in Table 1.

[0056] Table 1: FeatureSet data set

[0057] Case number <![CDATA[FeatureSet 1 > <![CDATA[FeatureSet 2 > <![CDATA[FeatureSet 3 > 1 Segment Shield tunneling method Micro-disturbance grouting 2 Shotcrete NATM None …… 187 Segment TBM method Face reinforcement

[0058] Step S2.2, for each feature data set FeatureSet i , divide out n i state features, and according to the values of each case, obtain the data set ConditionSet i under the state feature representation, where n i is the number of all discrete values under this feature. In this embodiment, taking the FeatureSet 2 formed by "construction method" as an example, the feature values in 187 cases include four discrete values of "shield method", "TBM method", "mining method", and "open cut method". Accordingly, n 2 = 4 state features are divided and respectively correspond to the original discrete values. When the case belongs to this discrete value, the feature value is 1, and when it does not belong, the feature value is 0, thus generating the data set ConditionSet under the state feature representation.2 , as shown in Table 2 specifically.

[0059] Table 2: ConditionSet 2 Data set

[0060]

[0061] Step S2.3: Combine the data sets ConditionSet under m state feature representations i to obtain the non-numerical data state feature representation data set Bset.

[0062] After that, further processing is performed on the initial input set DatasetV1.

[0063] Step S3: According to the source of the design data, adopt a certain measurement strategy to determine the information source weight; the information source weight is an index to measure the importance of different data sources, and methods such as the expert scoring method, interval random generation method, ranking comparison method, and comprehensive method can be used to determine it.

[0064] The sources of design data can include theoretical analysis, numerical simulation, model tests, and on-site monitoring, etc. For this embodiment, after statistics, there are 15 cases based on theoretical analysis, 136 cases based on numerical simulation, 7 cases based on model tests, and 29 cases based on on-site monitoring. In this embodiment, according to the expert scoring method, the importance weights of theoretical analysis, numerical simulation, model tests, and on-site detection are 0.07, 0.24, 0.12, and 0.57 respectively after weighing.

[0065] Step S4: According to the value conditions of the key features of the design data, determine the adaptability weights of different data. The adaptability weight is an index to measure the degree of closeness of the data to the typical working conditions:

[0066] In this embodiment, a shield tunnel in soft soil areas is used as a typical case, and the constructed training set is mainly aimed at predicting the settlement in soft soil areas, while referring to the potential influence laws of other working conditions. Therefore, the adaptability weights are calculated according to the following specific steps:

[0067] Step S4.1: Determine the important judgment feature value y of the typical working condition i(i = 1, 2, ……, n), where n is the number of important judgment features (key features). In this embodiment, two features, namely "soil layer modulus" and "construction method", are used as key features, and 3000 kPa and "shield method" are used as the values of the key features. It should be noted that the "construction method" feature has been converted in step S2, and the converted state feature is used as the key feature here. Among them, the typical working conditions are one or more representative cases in the crossing project, and there are one or more important judgment feature values in some features, which can reflect the characteristics of this type of working condition; the key feature refers to one or more features that distinguish different types of cases, and the type of the prediction target can be used as the selection basis, and the engineering experience and classification metrics commonly recognized in the industry are used as the values of the key features, which are well-known to those of ordinary skill in the art;

[0068] Step S4.2: Calculate the Jaccard distance D of the discrete features in the important judgment features 1 , and the formula is:

[0069]

[0070] In the above formula, x and y are respectively the sets of the discrete feature values of the case and the typical working condition.

[0071] Step S4.3: Calculate the relative Euclidean distance D of the continuous features in the important judgment features 2 , and the formula is:

[0072]

[0073] In the above formula, x i and y i are respectively the values of the case and the typical feature in the i-th continuous feature, and l is the number of continuous features in the important judgment features.

[0074] Step S4.4: Combine the distance D 1 and D 2 , and calculate the adaptability weight of each case. The formula is:

[0075] ω fit = f(D 1 ) + g(D 2 )

[0076] In the formula, both f(x) and g(x) are distance conversion functions. In this embodiment, both distance conversion functions take inverse proportional functions with an inverse proportional coefficient equal to 1.

[0077] Step S5: Combine the weights in steps S3 and S4 and use them as the features of the input set to obtain the weighted input set DatasetV2:

[0078] More specifically, in this embodiment, the information source weights are determined according to the four types of sources in S3, and the adaptability weights are determined according to the soft soil shield working conditions in S4. Then, the multiplication method is adopted, that is, all sub-item weights are multiplied, and the product result is used as the combined weight, and the weight is added as a new feature to the obtained dataset.

[0079] Step S6: Use the weighted input set DatasetV2 in step S5 as the input training set, and use the tunnel deformation value as the output training set to form the standard training set TrainDataset.

[0080] In this embodiment, by using the input training set obtained in the previous steps and combining the tunnel vertical deformation values collected in the materials, a training set for machine learning algorithms can be obtained.

[0081] Furthermore, the present invention also provides a method for predicting the deformation of a tunneling project tunnel based on a machine learning algorithm, which may include constructing a multi-source weighted training set by using the method of the present invention; then using the multi-source weighted training set to train the machine learning algorithm. The machine learning algorithm can be, for example, algorithms such as neural networks and decision trees, and more specifically, it can include ANN, CNN, XGboost, lightGBM, etc. After training, the machine learning algorithm is used to predict the deformation of the tunneling project tunnel.

[0082] In addition, the present invention also provides a system for predicting the deformation of a tunneling project tunnel based on a machine learning algorithm. The system may include one or more processors and a memory. The memory stores instructions executable by the one or more processors, and the instructions enable the system to execute each method step according to the present invention.

[0083] The above description of the embodiments is for those of ordinary skill in the art in this technical field to understand and apply the present invention. It is obvious that those skilled in the art can easily make various modifications to these embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the embodiments here, and the improvements and modifications made by those skilled in the art without departing from the scope of the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A method for constructing a multi-source weighted training set for predicting the deformation of a tunneling engineering tunnel, comprising the following steps: Step S1: Obtain the design data of the tunneling project and the corresponding tunnel deformation data, where the design data is the input part of the training set, including engineering geological information, tunnel structure information, and construction information, and the tunnel deformation is the output part of the training set; Step S2: Divide the design data into numerical data and non-numerical data according to the characteristics of the design data in Step S1, form the numerical data into a data set ASet, process the non-numerical data to obtain a data set Bset, and combine Aset and Bset to obtain an initial input set DatasetV1; Step S3: Determine the information source weights according to the source of the design data in the initial input set DatasetV1 using a certain measurement strategy; Step S4: Determine the adaptability weights of different data according to the value of the key features of the design data in the initial input set DatasetV1; Step S5: Combine the weights in Steps S3 and S4 and use them as the features of the input set to obtain a weighted input set DatasetV2; Step S6: Use the weighted input set DatasetV2 in Step S5 as the input training set and the tunnel deformation data as the output training set to form a standard training set TrainDataset; Among them, the information source weight in Step S3 is an index for measuring the importance of different data sources, and the measurement method is selected from one or more of the expert scoring method, the interval random generation method, and the ranking comparison method; Among them, the adaptability weight in Step S4 is a measure of the degree of proximity of the data to the typical working conditions, and its determination includes the following steps: Step S4.1: Determine the key characteristic value y of the typical working condition i , where i = 1, 2, ……, n, and n is the number of key characteristics; Step S4.2, calculate the Jaccard distance D of the discrete features among the key features 1 ; Step S4.3, calculate the relative Euclidean distance D of the continuous features among the key features 2 ; and Step S4.4, Combine distance D 1 and D 2 , calculate the adaptability weight of each case, and the calculation formula is: ω fit = f(D 1 ) + g(D 2 ) In the formula, both f(x) and g(x) are distance conversion functions.

2. A method for constructing a multi-source weighted training set for predicting the deformation of a tunneling engineering tunnel according to claim 1, characterized in that: In Step S2, the processing of the non-numerical data includes encoding the non-numerical data using the One-Hot method, and the encoding process includes the following steps: Step S2.1: Screen the data with non-numerical design features and form a data set FeatureSet. The total number of features in the data set FeatureSet is denoted as m, and the data set corresponding to each feature is denoted as FeatureSet i , where i = 1, 2, ……, m; Step S2.

2. For each feature dataset FeatureSet i , divide into n i state features, and obtain the dataset ConditionSet i under the state feature representation according to the values of each case, where n i is the number of all discrete values under this feature; and Step S2.3: Combine the datasets ConditionSet under m state feature representations i to obtain the non-numerical data state feature representation dataset Bset.

3. A method for constructing a multi-source weighted training set for predicting the deformation of a tunneling engineering tunnel according to claim 1, characterized in that: The sources of the design data in Step S3 include theoretical analysis, numerical simulation, model tests, and on-site monitoring.

4. A method for constructing a multi-source weighted training set for predicting the deformation of a tunneling engineering tunnel according to claim 1, characterized in that: The expert scoring method directly determines the scores of each data source according to expert evaluations, so that the data of the same source has the same weight; The interval random generation method stipulates the weight interval of each data source, and randomly generates weights within the weight interval for each case according to its source; The ranking comparison method ranks the importance of the data sources, randomly selects multiple cases from the data set without replacement and arranges them according to the importance, and assigns weights from high to low according to the ranking until all cases are selected.

5. A method for constructing a multi-source weighted training set for predicting the deformation of a tunneling engineering tunnel according to claim 1, characterized in that: In step S4.2, the Jacquard distance D 1 The calculation formula is: In the above formula, x and y are respectively the sets of discrete feature values of the cases and typical working conditions; In step S4.3, the relative Euclidean distance D 2 The calculation formula is as follows: In the above formula, x i and y i are the values of the case and the typical working condition for the i-th continuous feature respectively, and l is the number of continuous features among the key features.

6. A multi-source weighted training set construction method for predicting the deformation of a cross-engineering tunnel as described in claim 5, characterized in that: The distance conversion function is selected from an inverse proportional function with a positive inverse proportional coefficient, a linear function with a negative slope and a positive intercept, or a monotonically decreasing function with a positive value range.

7. A multi-source weighted training set construction method for predicting the deformation of a cross-engineering tunnel as described in claim 1, characterized in that: The weight combination in step S5 is to obtain a comprehensive weight through combination according to the information source weight and the adaptability weight, and the combination strategy is selected from: (1) The addition method, which stipulates the proportion of the sub-item weight in the combined weight, and takes the result of the proportional addition as the combined weight; (2) The multiplication method, which multiplies all sub-item weights and takes the product result as the combined weight; (3) The minimum method, which selects the minimum value among all sub-item weights as the combined weight; (4) The maximum method, which selects the maximum value among all sub-item weights as the combined weight; and (5) The random method, which takes the minimum value among all sub-item weights as the lower bound of the interval, the maximum value as the upper bound of the interval, and randomly generates a weight within the interval as the combined weight.

8. A method for predicting the deformation of a cross-engineering tunnel based on a machine learning algorithm, including the steps of: Constructing a training set by using the method described in any one of claims 1-7; Training a machine learning algorithm by using the training set; and Predicting the deformation of a cross-engineering tunnel by using the machine learning algorithm.

9. The method as described in claim 8, wherein, The machine learning algorithm is selected from ANN, CNN, XGboost, and lightGBM.

10. A system for predicting the deformation of a cross-engineering tunnel based on a machine learning algorithm, including: One or more processors; And a memory, the memory stores instructions executable by the one or more processors, and the instructions cause the system to execute the method described in any one of claims 8 to 9.

Citation Information

Patent Citations

  • Power missing data filling method based on hybrid strategy

    CN108805193A

  • XGBoost-based underground comprehensive pipe gallery safety condition evaluation method

    CN111950585A