TBM construction core database construction method and system

By quantifying the correlation between continuous and discrete variables of TBM tunneling parameters, a core database was constructed, which solved the problem of large data volume and complex data types of TBM tunneling parameters, and achieved efficient dimensionality reduction and improved modeling efficiency.

CN120763148BActive Publication Date: 2025-11-25CHINA RAILWAY SHISIJU GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511278991.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-25
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

TBM tunneling parameter data is large in volume and complex in type. Existing methods are unable to effectively identify key features, resulting in high data modeling and analysis costs and limited computing resources, making it difficult to apply big data technologies.

Method used

A mixed variable correlation analysis method was adopted to quantify the correlation between continuous and discrete variables. Strongly correlated parameter groups were formed through mapping and grouping. Key parameters were selected for dimensionality compression to construct a core database for TBM construction.

Benefits of technology

It achieves efficient dimensionality reduction and data compression, reduces computational complexity, improves modeling efficiency and engineering interpretability, and is suitable for data modeling and analysis of TBM tunneling parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763148B_ABST
    Figure CN120763148B_ABST
Patent Text Reader

Abstract

The application discloses a TBM construction core database construction method and system, and belongs to the tunneling big data processing technical field of TBM. The method comprises the following steps: grouping TBM tunneling parameter data to obtain a continuous variable set and a discrete variable set; selecting corresponding quantization modes to obtain the correlation between continuous variables, the correlation between discrete variables and the correlation between continuous variables and discrete variables, and obtaining a plurality of groups of strongly correlated variable pairs, and classifying and merging to obtain a plurality of groups of strongly correlated parameter groups; for each group of strongly correlated parameter groups, selecting one variable as a key parameter to perform dimension compression to obtain a simplified parameter set; and taking the simplified parameter set as a core database for subsequent data modeling and analysis processes. The application realizes the compression of TBM tunneling parameter data in the parameter dimension, and reduces the calculation cost of subsequent data modeling and analysis tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of big data processing technology for TBM tunneling, and in particular relates to a method and system for constructing a core database for TBM construction. Background Technology

[0002] TBMs (Hard Rock Tunnel Boring Machines) are crucial pieces of equipment used in hard rock tunnel construction. Due to their high tunneling efficiency, good safety, and low environmental impact, they are widely used in tunnel construction in mountainous and undersea areas. With the rapid development of technologies such as sensors, data transmission, data acquisition, and parallel computing, TBM tunneling parameters such as propulsion speed, cutterhead torque, and cutterhead rotation speed are extensively recorded. These parameters are then aggregated through technologies such as PLCs and the Internet of Things (IoT) and transmitted to intelligent platforms for TBM construction companies, manufacturers, and research institutions, forming a large dataset of TBM tunneling parameters. This dataset provides a data foundation and important reference for guiding TBM design, operation, maintenance, and construction.

[0003] Existing research on TBM tunneling parameter big data mainly focuses on modeling, analysis, and mining of tunneling parameter data, with less attention paid to the TBM tunneling parameter big data itself. The proposed theoretical method has been successfully applied in areas such as TBM tunneling performance prediction, tunneling parameter decision-making, and surrounding rock condition identification, improving TBM tunneling efficiency and providing theoretical and methodological support for intelligent TBM tunneling construction.

[0004] TBMs are "factory-produced" equipment with numerous tunneling parameters, ranging from dozens to hundreds. During data modeling, this abundance of parameters can prevent modeling, analysis, and mining from capturing key features within the data, leading to the "curse of dimensionality." Furthermore, the sheer number of parameters also results in excessively high costs for data modeling, analysis, and mining. Considering the limited computing resources at TBM construction sites, the established modeling, analysis, and mining models face challenges in application. This is one of the key difficulties limiting the application of big data technology to TBM tunneling parameter data.

[0005] Identifying strongly correlated tunneling parameters through TBM tunneling parameter correlation analysis, and then reducing these strongly correlated parameters to form a core parameter set for TBM tunneling, is an effective method for TBM tunneling parameter data compression. TBM tunneling parameters contain both continuous and discrete variables. Existing correlation analysis methods are mostly designed for single continuous or discrete variables, making them ineffective in TBM tunneling parameter correlation analysis. Therefore, there is an urgent need to propose a correlation analysis and data compression method suitable for TBM tunneling parameters with a mixture of continuous and discrete variables. Summary of the Invention

[0006] To address the technical problems existing in the background art, the present invention provides a method and system for constructing a core database for TBM construction.

[0007] This invention adopts the following technical solution: a method for constructing a core database for TBM construction, comprising the following steps:

[0008] Obtain TBM tunneling parameter data, identify continuous and discrete variables, and group them to obtain a set of continuous variables. and discrete variable set ;

[0009] Choose an appropriate quantization method to quantify the correlation between continuous variables and the correlation between discrete variables respectively;

[0010] Construct mapping pairs between continuous and discrete variables, and then map the set of continuous variables based on the values ​​of the discrete variables. Grouping is performed to form multiple subsets of continuous variables; the numerical distributions of the multiple subsets of continuous variables are compared to quantify the correlation between continuous and discrete variables;

[0011] Based on the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous and discrete variables, multiple sets of strongly correlated variable pairs are obtained, and these sets of strongly correlated variable pairs are merged and classified to form several sets of strongly correlated parameter groups.

[0012] For each group of strongly correlated parameters, one variable is selected as the key parameter and its dimensions are compressed to obtain a simplified parameter set. This simplified parameter set is used as the core database for subsequent data modeling and analysis.

[0013] In a further embodiment, a set of continuous variables is defined. : , The number of continuous variables, As variables; the correlation between continuous variables is quantified using the following methods:

[0014] From a set of continuous variables Select all or part of the continuous variables in the selection. Construct data column vectors :

[0015] ,in, For the first One variable, , Representing variables In the The values ​​of each sampling point;

[0016] Based on the data column vector Build OK Column sample data matrix ;

[0017] For any two selected variables and variables Calculate the data column vector and data column vector Correlation coefficient between .

[0018] In a further embodiment, a set of discrete variables is defined. : , q The number of discrete variables, The following quantification methods are used to quantify the correlation between discrete variables:

[0019] In the set of discrete variables Choose any two variables and variables Statistical variables and variables The actual frequency of joint occurrence in the sample And construct a two-dimensional cross-frequency table;

[0020] The theoretical frequencies are calculated using the two-dimensional cross-frequency table. Based on actual frequency and theoretical frequency Determine variables and variables Significance probability ;

[0021] Then, the correlation index of any discrete variable Represented as: .

[0022] In a further embodiment, the following procedure is used to quantify the correlation between continuous and discrete variables:

[0023] Randomly select two subsets of continuous variables from multiple subsets. and Statistical distribution calculations were used to determine the distribution differences. :

[0024] ;

[0025] In the formula, and Two sets of continuous variable subsets respectively and The sample mean, and Each is a subset of two continuous variables. and The sample variance and Each is a subset of two continuous variables. and The number of samples;

[0026] The correlation strength index between continuous and discrete variables Represented as: , The significance probability between continuous and discrete variables is determined by the difference in their distributions. Sure.

[0027] In a further embodiment, the compression method for the simplified parameter set is as follows:

[0028] A key parameter retention mechanism is created, in which one variable is selected as the key parameter to replace the original continuous and discrete variables; the remaining variables are regarded as redundant parameters and retained as alternative parameters.

[0029] The key parameter retention mechanism includes at least one of the following: engineering importance priority mechanism, usage frequency-based priority mechanism, data quality priority mechanism, and relevance-significant priority mechanism.

[0030] In a further embodiment, based on the correlation between continuous variables,

[0031] Multiple pairs of strongly correlated variables are obtained using the following method: If , then represents the variable and variables There is a strong correlation between them, making them a strongly correlated variable pair;

[0032] like , then represents the variable and variables The correlation between them is moderate.

[0033] like , then represents the variable and variables The correlations between them are weak or nonexistent; among them, A positive value indicates a positive correlation. A negative value indicates a negative correlation; among which, This is the threshold for the maximum correlation of continuous variables. This is the minimum correlation threshold for continuous variables.

[0034] In a further embodiment, based on the correlation between discrete variables, multiple pairs of strongly correlated variables are obtained in the following manner:

[0035] When correlation indicators When the value is close to 1, it indicates that the variable... and variables There is a strong correlation between them, making them a strongly correlated variable pair;

[0036] When correlation indicators When it is close to 0, it indicates that the variable and variables The correlation between them is weak or nonexistent.

[0037] In a further embodiment, based on the correlation between continuous and discrete variables, multiple pairs of strongly correlated variables are obtained in the following manner:

[0038] When correlation indicators Greater than or equal to When the value is equal to 0, it indicates that there is a strong correlation between the corresponding continuous variable and discrete variable, which is a strongly correlated variable pair. This is the threshold for the maximum correlation between continuous and discrete variables.

[0039] When correlation indicators Less than When the value is equal to 0, it indicates that the corresponding continuous variable and discrete variable are weakly correlated or uncorrelated. This is the minimum correlation threshold between continuous and discrete variables.

[0040] A TBM construction core database construction system, used to implement the TBM construction core database construction method described above, includes:

[0041] The TBM tunneling parameter processing module is configured to acquire TBM tunneling parameter data, identify continuous and discrete variables, and group them to obtain a set of continuous variables. and discrete variable set ;

[0042] The continuous variable correlation analysis module is set to select an appropriate quantification method to quantify the correlation between continuous variables.

[0043] The discrete variable correlation analysis module is set to select an appropriate quantification method to quantify the correlation between discrete variables respectively;

[0044] The correlation analysis module for continuous and discrete variables is configured to construct mapping pairs between continuous and discrete variables, and to map the set of continuous variables based on the values ​​of the discrete variables. Grouping is performed to form multiple subsets of continuous variables; the numerical distributions of the multiple subsets of continuous variables are compared to quantify the correlation between continuous and discrete variables;

[0045] The key parameter screening and dimensionality compression module is configured to obtain multiple sets of strongly correlated variable pairs based on the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous and discrete variables. These sets of strongly correlated variable pairs are then merged and categorized to form several sets of strongly correlated parameter groups. For each set of strongly correlated parameter groups, one variable is selected as a key parameter for compression to obtain a simplified parameter set, which is then used in the subsequent data modeling and analysis process.

[0046] The beneficial effects of this invention are as follows: Based on a mixed-variable statistical testing strategy, this invention systematically quantifies the correlation relationships among TBM tunneling parameters. It selects appropriate methods to assess the correlations between continuous variables, between discrete variables, and between continuous and discrete variables. This effectively addresses the problems of a large number of TBM tunneling parameters, their diverse types, and complex relationships, providing high-quality variable selection criteria for parameter dimensionality reduction, model optimization, and engineering feature extraction. Compared with traditional empirical variable selection methods, this invention has advantages such as strong quantitative analysis capabilities, standardized operation, and high analytical objectivity. Compared with black-box feature selection algorithms, this invention is transparent, highly interpretable, and more suitable for practical applications in the engineering field. Compared with linear analysis methods that only apply to numerical variables, this invention can simultaneously handle continuous and discrete variables, improving the model's adaptability and generalization ability.

[0047] This invention further establishes a unified variable grouping strategy and key parameter retention mechanism based on variable correlation analysis, achieving orderly compression of the parameter space, effectively eliminating redundant information, and reducing model computational complexity and overfitting risk. Compared with dimensionality reduction methods based on principal component analysis (PCA) or embedded models, this invention can fully preserve the physical meaning of the original variables, improving the engineering feasibility and practicality of subsequent analysis and interpretation.

[0048] This invention utilizes a TBM tunneling parameter processing module, a continuous variable correlation analysis module, a discrete variable correlation analysis module, a variable grouping analysis module, a continuous-discrete variable correlation analysis module, and a key parameter screening and dimensionality compression module to collaboratively achieve TBM tunneling parameter correlation identification and structured dimensionality reduction processing, supporting the efficient construction of the TBM core database and the implementation of data modeling tasks. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating the TBM parameter correlation analysis and data compression method in Example 1.

[0050] Figure 2 This is a heatmap showing the correlation of the Pearson coefficient, a continuous variable in TBM, in Example 1.

[0051] Figure 3This is a heatmap of the correlation between TBM discrete variables and chi-square test results in Example 1.

[0052] Figure 4 This is a heatmap showing the correlation between continuous and discrete variables in the t-test of TBM in Example 1.

[0053] Figure 5 This is a system architecture diagram of a TBM parameter correlation analysis and data compression method in Example 2. Detailed Implementation

[0054] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0055] Example 1

[0056] like Figure 1 As shown in the figure, this embodiment discloses a method for constructing a core database for TBM construction, including the following steps:

[0057] Obtain TBM tunneling parameter data, identify continuous and discrete variables, and group them to obtain a set of continuous variables. and discrete variable set ;

[0058] Choose an appropriate quantization method to quantify the correlation between continuous variables and the correlation between discrete variables respectively;

[0059] Construct mapping pairs between continuous and discrete variables, and then map the set of continuous variables based on the values ​​of the discrete variables. Grouping is performed to form multiple subsets of continuous variables; the numerical distributions of the multiple subsets of continuous variables are compared to quantify the correlation between continuous and discrete variables;

[0060] Based on the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous and discrete variables, multiple sets of strongly correlated variable pairs are obtained, and these sets of strongly correlated variable pairs are merged and classified to form several sets of strongly correlated parameter groups.

[0061] For each group of strongly correlated parameters, one variable is selected as the key parameter and its dimensionality is compressed to obtain a simplified parameter set, which is then used in the subsequent data modeling and analysis process.

[0062] Table 1. Raw data of some TBM parameters

[0063]

[0064] As shown in Table 1, the TBM tunneling parameter data described in this embodiment can be obtained through the data monitoring platform built into the TBM itself, or by retrieving the raw TBM tunneling parameter data through the TBM cloud-edge collaborative monitoring platform. The raw data in this embodiment contains at least 2000 sample data points, including data on 12 parameters such as cutterhead speed, cutterhead thrust, cutterhead torque, propulsion system hydraulic pressure, temperature, and motor current. After data preprocessing (such as missing value imputation and outlier removal), correlation analysis was performed on different types of variables.

[0065] Based on this, the set of continuous variables described in this embodiment and discrete variable set The grouping steps are as follows:

[0066] Step 1.1: Obtain raw parameter data from the TBM tunneling construction site. This data is collected in real time by multiple sensors and covers multiple working condition parameters during the tunneling process.

[0067] Step 1.2: Standardize the TBM tunneling parameter data, including but not limited to: missing value imputation, outlier removal, and time alignment, to ensure data integrity and consistency and improve the reliability of subsequent statistical analysis results.

[0068] Step 1.3: Classify variables according to their physical meaning, data type, value format, and engineering experience:

[0069] Continuous variables include, but are not limited to: feed speed, cutter head speed, feed force, cutter head torque, motor current, motor voltage, motor torque, oil temperature, and other parameters that are expressed in continuous real number form and can be quantified and measured.

[0070] Discrete variables include, but are not limited to, parameters expressed in the form of categories or states, such as fan start / stop, motor switch, gearbox switch, construction status code, tunneling mode selection, and sensor operating status.

[0071] Step 1.4: Based on the grouping in Step 1.3, obtain the set of continuous variables. and discrete variable set and the set of continuous variables Recorded as .in, The number of continuous variables, For variables. Referring to the examples above, a set of continuous variables. The elements in the data are: feed speed, cutter head speed, feed force, cutter head torque, motor current, motor voltage, motor torque, and oil temperature; that is, the current... The value is 8.

[0072] This embodiment employs Pearson correlation coefficient, chi-square test of independence, and t-test to conduct correlation analysis for differences in variable types, aiming to improve the accuracy and suitability of the evaluation results. Pearson correlation coefficient effectively measures the degree of linear correlation between continuous variables and is suitable for data with definable means and variances that follow an approximately normal distribution. Chi-square test of independence is suitable for determining the independence between discrete variables and can reflect the distributional association characteristics of categorical variables. t-test is suitable for combinations of continuous variables and binary discrete variables, measuring the degree of association between them by comparing the significance of differences in the means of continuous variables under different categories. Compared with single correlation quantification methods, these categorical matching statistical tests can fully utilize the statistical characteristics of various variables, thus revealing the correlation between TBM tunneling parameters more scientifically and comprehensively.

[0073] Set of continuous variables This can be further expressed as: {propulsion speed} Cutter head speed Propulsion Cutter head torque Motor current Motor voltage Motor torque Oil temperature The following quantification methods are used to quantify the correlation between continuous variables:

[0074] Select all continuous variables, and for each variable... Construct data column vectors , The construction method is as follows:

[0075] ,in, For the first One variable, , Representing variables In the The values ​​at each sampling point. It is worth noting that the sampling points described in this embodiment are samples of the variable at different time points.

[0076] Based on the data column vector Build OK Column sample data matrix ; to advance speed For example, there are 2000 samples in total, therefore, the corresponding data column vector This is a sample data matrix with 2000 rows and 8 columns.

[0077] For any two selected variables and variables Calculate the data column vector and data column vector Correlation coefficient between This embodiment will use the Pearson correlation coefficient to calculate the correlation coefficient. The calculation formula is as follows:

[0078] ;

[0079] In the formula, and They are data column vectors and data column vector The number of elements, and Variables and variables The sample mean, Data column vector The Middle One element, Data column vector The Middle Each element.

[0080] The data column vector is obtained based on the above calculations. and data column vector Correlation coefficient between Subsequently, multiple pairs of strongly correlated variables were obtained using the following method: If , then represents the variable and variables There is a strong correlation between them, making them a strongly correlated variable pair;

[0081] like , then represents the variable and variables The correlation between them is moderate.

[0082] like , then represents the variable and variables The correlations between them are weak or nonexistent; among them, A positive value indicates a positive correlation. A negative value indicates a negative correlation; among which, This is the threshold for the maximum correlation of continuous variables. This is the minimum correlation threshold for continuous variables. In this embodiment, , .

[0083] Combination Figure 2For example, the correlation coefficients between the propulsion speed and other continuous variables (cutterhead speed, propulsion force, cutterhead torque, motor current, motor voltage, motor torque, and oil temperature) are calculated as follows: , , , , , and Based on the methods used to obtain multiple sets of strongly correlated variable pairs, it can be seen that the propulsion speed forms a strongly correlated variable pair with the following variables: cutterhead rotation speed, propulsion force, cutterhead torque, and motor torque.

[0084] In another embodiment, the set of discrete variables Recorded as , q The number of discrete variables, For variables. Referring to the examples above, the set of discrete variables... The elements in the code are: fan start / stop, motor switch, reducer temperature switch, and gear oil system; at this time, q The value of is 4.

[0085] Based on the examples above, the discrete variable set described in this embodiment... Represented as: {Fan start / stop} Motor switch Gearbox temperature switch The gear oil system is normal. The correlation between discrete variables was quantified using the following methods:

[0086] In the set of discrete variables Choose any two variables and variables Statistical variables and variables The actual frequency of joint occurrence in the sample And construct a two-dimensional cross-frequency table (contingency table);

[0087] The theoretical frequencies are calculated using the two-dimensional cross-frequency table. Based on actual frequency and theoretical frequency Determine variables and variables Significance probability ;

[0088] Then, the correlation index of any discrete variable Represented as: .

[0089] Furthermore, the theoretical frequency was calculated using the following method. Assuming any two variables are selected and variables To ensure that the frequencies are independent of each other, the theoretical frequency of each cell in the two-dimensional cross-frequency table is calculated based on the marginal distribution. The calculation formula is as follows:

[0090] ;in, This represents the sum of the row frequencies of the cells. This represents the sum of column frequencies for each cell. This represents the number of cells.

[0091] Furthermore, the statistic is calculated based on all cells N using the following formula. : .

[0092] Based on the calculated statistics and degrees of freedom The significance probability can be obtained by querying the chi-square distribution table. The significance probability The smaller the value, the higher the actual frequency. The greater the difference between the theoretical frequency of the variables and the independent variable, the stronger the correlation between the variables.

[0093] To quantify the strength of correlation, a correlation index is defined for any discrete variable. Represented as: .

[0094] Furthermore, based on the correlation between discrete variables, multiple pairs of strongly correlated variables are obtained using the following method:

[0095] When correlation indicators When the value is close to 1, it indicates that the variable... and variables There is a strong correlation between them, forming a strongly correlated variable pair; such as ,Right now .

[0096] When correlation indicators When it is close to 0, it indicates that the variable and variables The two are weakly correlated or uncorrelated; for example .

[0097] like Figure 3 As shown, this embodiment selects the start and stop of the fan. and motor switch For example, the corresponding significance probability is calculated. If the value is 0.01, then the correlation index can be determined. The value is 0.99. Therefore, it can be considered that there is a very strong correlation between the start-up and shutdown of the fan and the switching of the motor, that is, a strongly correlated variable pair.

[0098] To facilitate understanding the grouping process of multiple subsets of continuous variables, this embodiment will use a set of continuous variables as an example. :{Propulsion speed Cutter head speed Propulsion Cutter head torque Motor current Motor voltage Motor torque Oil temperature } and discrete variable set Fan start / stop Motor switch Gearbox temperature switch The gear oil system is normal. For example, constructing mapping pairs between continuous and discrete variables to advance speed. Start and stop of the fan For example, according to the start and stop of the wind turbine The various values ​​will correspond to the propulsion speed. Grouping. Furthermore, if the wind turbines start and stop... If the value is {0, 1}, then the start / stop of the fan is extracted. Propulsion speed with a value of 0 All samples form a continuous variable subset Extraction fan start / stop Propulsion speed with a value of 1 All samples form a continuous variable subset .

[0099] Based on this, respectively using and Instead of specifying the exact values, the superordinate represents any two selected subsets of continuous variables. and To further understand this, it means to accelerate the process. According to the start and stop of the fan The variable is divided into two subsets of continuous variables, which are the subsets of continuous variables. and continuous variable subsets Furthermore, continuous variable subsets Start and stop of the wind turbine The set of propulsion velocities during (wind turbine) operation is denoted as... Subset of continuous variables Start and stop of the wind turbine The set of propulsion velocities when the fan is running is denoted as... .

[0100] The following procedure is used to quantify the correlation between continuous and discrete variables, and to calculate the statistical distribution differences. :

[0101] ;

[0102] In the formula, and Two sets of continuous variable subsets respectively and The sample mean, and Each is a subset of two continuous variables. and The sample variance and Each is a subset of two continuous variables. and The number of samples;

[0103] Based on the differences in the calculated distribution and degrees of freedom The significance probability can be obtained by querying the distribution table or by calculation. To facilitate quantitative analysis, the correlation strength index between continuous and discrete variables is used. Represented as: .

[0104] Furthermore, multiple pairs of strongly correlated variables were obtained using the following method:

[0105] When correlation indicators Greater than or equal to When the value is equal to 0, it indicates that there is a strong correlation between the corresponding continuous variable and discrete variable, which is a strongly correlated variable pair. The threshold for the maximum correlation between continuous and discrete variables, such as At that time, that is .

[0106] When correlation indicators Less than When the value is equal to 0, it indicates that the corresponding continuous variable and discrete variable are weakly correlated or uncorrelated. The minimum correlation threshold between continuous and discrete variables is used to... For example, that is .

[0107] Combination Figure 4 In this embodiment, the correlation strength index between propulsion speed (C1) and wind turbine start-up and shutdown (D1) is... The value is 0.98, which indicates a significant correlation between propulsion speed and turbine start-up / shutdown.

[0108] Multiple pairs of strongly correlated variables are obtained based on the correlations between continuous variables, between discrete variables, and between continuous and discrete variables. These pairs are then merged and categorized to form several strong pairs. The methods for obtaining these strongly correlated variable pairs have already been described above, so they will not be elaborated upon further.

[0109] As in the examples above, the continuous variable of feed speed forms a strongly correlated set of parameters with the cutterhead rotation speed, feed force, cutterhead torque, and motor torque, etc.

[0110] Furthermore, the compression method for simplifying the parameter set is as follows:

[0111] A key parameter retention mechanism is created, in which one variable is selected as the key parameter to replace the original continuous and discrete variables; the remaining variables are regarded as redundant parameters and retained as alternative parameters.

[0112] The key parameter retention mechanism includes at least one of the following: engineering importance priority mechanism, usage frequency-based priority mechanism, data quality priority mechanism, and relevance-significant priority mechanism.

[0113] To better understand, the engineering importance optimization mechanism described in this embodiment prioritizes variables that are highly important in engineering experience; the frequency of use-dominated optimization mechanism prioritizes variables that are used frequently; the data quality optimization mechanism prioritizes variables with good data stability, low missing rate, and easy data collection; and the significant correlation optimization mechanism prioritizes variables that have the most significant relationship with the target variable.

[0114] Based on this, the simplified parameter set will be used as the core database for subsequent data modeling and analysis. For example, the final selected simplified parameter set may include: a simplified set of continuous variables. and the simplified set of discrete variables Based on the examples above, , .

[0115] Therefore, by using the above technical solution, strongly correlated parameters are grouped into the same group. For each group of strongly correlated tunneling parameters, one parameter from each group is retained as the key parameter, thus compressing the TBM tunneling parameter data in terms of parameter dimension. Through variable grouping and filtering, the original 12 parameters are compressed into 3 representative parameters, with a parameter dimension compression rate of over 75%, significantly reducing the computational cost and redundancy risk of modeling.

[0116] Example 2

[0117] To implement the TBM construction core database construction method described in Example 1, such as Figure 5This embodiment discloses a TBM construction core database construction system, including:

[0118] The TBM tunneling parameter processing module is configured to acquire TBM tunneling parameter data, identify continuous and discrete variables, and group them to obtain a set of continuous variables. and discrete variable set ;

[0119] The continuous variable correlation analysis module is set to select an appropriate quantification method to quantify the correlation between continuous variables.

[0120] The discrete variable correlation analysis module is set to select an appropriate quantification method to quantify the correlation between discrete variables respectively;

[0121] The correlation analysis module for continuous and discrete variables is configured to construct mapping pairs between continuous and discrete variables, and to map the set of continuous variables based on the values ​​of the discrete variables. Grouping is performed to form multiple subsets of continuous variables; the numerical distributions of the multiple subsets of continuous variables are compared to quantify the correlation between continuous and discrete variables;

[0122] The key parameter screening and dimensionality compression module is configured to obtain multiple sets of strongly correlated variable pairs based on the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous and discrete variables. These sets of strongly correlated variable pairs are then merged and categorized to form several sets of strongly correlated parameter groups. For each set of strongly correlated parameter groups, one variable is selected as a key parameter for compression to obtain a simplified parameter set, which is then used in the subsequent data modeling and analysis process.

Claims

1. A TBM construction core database construction method, characterized in that, The method comprises the following steps: Obtain TBM tunneling parameter data, determine continuous variables and discrete variables therein and group to obtain a continuous variable set and a discrete variable set ; selecting an appropriate quantization method to quantize the correlation between continuous variables and the correlation between discrete variables respectively; constructing pairs of mappings of continuous and discrete variables, based on the values of the discrete variables, to set of continuous variables grouping to form a plurality of groups of continuous variable subsets; comparing the numerical distributions of the plurality of groups of continuous variable subsets to quantify the correlation between the continuous and discrete variables; obtaining a plurality of groups of strongly correlated variable pairs according to the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous variables and discrete variables, and merging and classifying the plurality of groups of strongly correlated variable pairs to form a plurality of groups of strongly correlated parameter groups; for each group of strongly correlated parameter groups, selecting one variable as a key parameter to perform dimension compression to obtain a simplified parameter set, and taking the simplified parameter set as a core database for later data modeling and analysis; defining a set of continuous variables : , is the number of continuous variables, is the variable; the correlation between continuous variables is quantized by using the following quantization method: from a set of continuous variables all or part of the continuous variables selected construct a data column vector : wherein is the th variable, , denotes the value of the variable at the th sampling point; based on the data column vector constructing row a sample data matrix of columns ; for any two variables selected and variable , compute the correlation coefficient between data column vector and data column vector ; define a set of discrete variables : , q is the number of discrete variables, is the variable; the correlation between the discrete variables is quantified using the following quantification: In a discrete variable set , select any two variables and variables , statistical variables and variables , the actual frequency of joint occurrence in the sample , and construct a two-dimensional cross-frequency table; The theoretical frequencies are calculated using the two-dimensional cross-frequency table. Based on actual frequency and theoretical frequency Determine variables and variables Significance probability ; Then, the correlation index of any discrete variable is represented as: ; the correlation between continuous variables and discrete variables is quantized by using the following process: any two sets of continuous variables from among the plurality of sets of continuous variables and statistical distribution calculation for distribution difference ; The correlation strength index between continuous and discrete variables Represented as: , The significance probability between continuous and discrete variables is determined by the difference in their distributions. Sure.

2. The method of claim 1, wherein, The distribution difference The calculation formula is as follows: ; wherein and the sample mean of the subset of continuous variables and the sample variance of the subset of continuous variables and the sample size of the subset of continuous variables and the sample size of the subset of continuous variables and the sample size of the subset of continuous variables and the sample size of the subset of continuous variables 3. The method of claim 1, wherein the method further comprises: The compression method of the simplified parameter set is as follows: A key parameter reservation mechanism is created, and one variable is selected as a key parameter according to the key parameter reservation mechanism to replace the original continuous variable and discrete variable; the remaining variables are regarded as redundant parameters and are reserved as alternative parameters; The key parameter reservation mechanism includes at least one of the following: engineering importance optimization mechanism, usage frequency dominant priority mechanism, data high-quality priority mechanism, and correlation significant priority mechanism.

4. The method of claim 1, wherein, According to the correlation between continuous variables, the following method is used to obtain a plurality of groups of strongly correlated variable pairs: if , it indicates that there is a strong correlation between the variable and the variable , which is a strongly correlated variable pair. If represents a moderate degree of correlation between the variables and the variable ; like , then represents the variable and variables The correlations between them are weak or nonexistent; among them, A positive value indicates a positive correlation. A negative value indicates a negative correlation; among which, This is the threshold for the maximum correlation of continuous variables. This is the minimum correlation threshold for continuous variables.

5. The method of claim 1, wherein, According to the correlation between discrete variables, the following method is used to obtain a plurality of groups of strongly correlated variable pairs: When the correlation index is close to 1, it indicates that there is a strong correlation between the variables and the variables are a pair of strongly correlated variables. When the correlation index is close to 0, then it indicates that the variables are weakly correlated or uncorrelated. between the variables 6. The method of claim 1, wherein, According to the correlation between continuous variables and discrete variables, the following method is used to obtain a plurality of groups of strongly correlated variable pairs: When the correlation index is greater than or equal to , it indicates that there is a strong correlation between the corresponding continuous variable and discrete variable, which is a strong correlation variable pair; is the maximum correlation threshold between the continuous variable and the discrete variable. When the correlation index is less than , it indicates that there is weak correlation or no correlation between the corresponding continuous variable and discrete variable, is the minimum correlation threshold between the continuous variable and the discrete variable.

7. A TBM construction core database construction system for implementing the TBM construction core database construction method according to any one of claims 1 to 6, characterized by, It comprises: The TBM tunneling parameter processing module is configured to acquire TBM tunneling parameter data, determine continuous variables and discrete variables therein and group the continuous variables and the discrete variables to obtain a continuous variable set and a discrete variable set and a discrete variable set ; The continuous variable correlation analysis module is configured to select an appropriate quantization method to quantize the correlation between continuous variables respectively; The discrete variable correlation analysis module is configured to select an appropriate quantization method to quantize the correlation between discrete variables respectively; The continuous variable and discrete variable correlation analysis module is configured to construct mapping pairs of continuous variables and discrete variables, and to group the continuous variables into a plurality of continuous variable subsets based on the values of the discrete variables The continuous variable and discrete variable correlation analysis module is configured to construct mapping pairs of continuous variables and discrete variables, and to group the continuous variables into a plurality of continuous variable subsets based on the values of the discrete variables The continuous variable and discrete variable correlation analysis module is configured to construct mapping pairs of continuous variables and discrete variables, and to group the continuous variables into a plurality of continuous variable subsets based on the values of the discrete variables The key parameter screening and dimension compression module is configured to obtain a plurality of groups of strongly correlated variable pairs according to the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous variables and discrete variables, and merge and classify the plurality of groups of strongly correlated variable pairs to form a plurality of groups of strongly correlated parameter groups; for each group of strongly correlated parameter groups, select one variable as a key parameter to perform dimension compression to obtain a simplified parameter set, and use the simplified parameter set for later data modeling and analysis.

Citation Information

Patent Citations

  • Shield excavation face stratum property real-time sensing and tunneling parameter adjusting method and system

    CN114370284A

  • High-dimensional feature extraction method and device, computer equipment and storage medium

    CN114627965A