TBM construction core database construction method and system
By building a core database of TBM excavation parameters, the problem of large amount of TBM excavation parameter data and mixed types was solved, efficient data compression and modeling optimization were achieved, and the intelligence level of TBM construction was improved.
Patent Information
- Application Number
- CN202511278991.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
The amount of TBM excavation parameter data is large and the types are mixed, which makes it impossible for data modeling, analysis and mining models to obtain internal key features. In addition, due to limited computing resources, existing methods cannot effectively handle the correlation analysis of continuous and discrete variables, resulting in difficulties in data compression and application.
A mixed variable statistical test strategy is adopted to quantify the correlation between TBM excavation parameters. By constructing mapping pairs of continuous and discrete variables, multiple groups of strongly correlated parameters are formed. Key parameters are selected for dimensionality compression to form a core database.
It achieves efficient data compression of TBM excavation parameters, reduces computational complexity and overfitting risks, improves the adaptability of data modeling and engineering interpretability, and is suitable for practical engineering applications.
Smart Images

Figure CN120763148A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of TBM excavation big data processing, and in particular relates to a TBM construction core database construction method and system. Background Art
[0002] TBMs (hard rock tunnel boring machines) are essential equipment for hard rock tunnel construction. Due to their high excavation efficiency, excellent safety, and minimal environmental impact, they are widely used in mountain and submarine tunnel construction. With the rapid development of sensors, data transmission, data acquisition, and parallel computing technologies, TBM excavation parameters such as propulsion speed, cutterhead torque, and cutterhead speed are now widely recorded. These parameters are aggregated through PLC, the Internet of Things, and other technologies onto smart platforms for TBM contractors, manufacturers, and research institutes. This data provides a foundation and valuable reference for guiding TBM design, operation, maintenance, and construction.
[0003] Existing research on TBM excavation parameter big data focuses primarily on modeling, analysis, and mining of excavation parameter data, with relatively little attention paid to TBM excavation parameter big data itself. The proposed theoretical method has been successfully applied in areas such as TBM excavation performance prediction, excavation parameter decision-making, and surrounding rock condition identification, improving TBM excavation efficiency and providing theoretical and methodological support for intelligent TBM excavation construction.
[0004] TBMs are factory-based machines with a vast number of excavation parameters, ranging from dozens to hundreds. During the data modeling process, this multitude of excavation parameters prevents modeling, analysis, and mining models from capturing key features within the data, a phenomenon known as the "curse of dimensionality." Furthermore, this multitude of excavation parameters leads to prohibitively high costs for modeling, analysis, and mining. Furthermore, considering the limited computing resources at TBM construction sites, the established modeling, analysis, and mining models can be difficult to apply. This is one of the challenges limiting the application of big data technology to TBM excavation parameter data.
[0005] An effective method for TBM excavation parameter data compression is to identify strongly correlated excavation parameters through correlation analysis of TBM excavation parameters, and then simplify these strongly correlated excavation parameters to form a core parameter set. TBM excavation parameters have both continuous and discrete variables. Existing correlation analysis methods, which mostly focus on single continuous or discrete variables, are ineffective in TBM excavation parameter correlation analysis. Therefore, a correlation analysis and data compression method suitable for mixed continuous / discrete TBM excavation parameter variables is urgently needed. Summary of the Invention
[0006] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a TBM construction core database construction method and system.
[0007] The present invention adopts the following technical solution: a method for constructing a TBM construction core database, comprising the following steps: Obtain TBM excavation parameter data, determine the continuous variables and discrete variables, and group them to obtain a continuous variable set and a set of discrete variables ; Select an appropriate quantification method to quantify the correlation between continuous variables and the correlation between discrete variables respectively; Construct a mapping pair between continuous variables and discrete variables, and convert the continuous variable set into Grouping to form multiple continuous variable subsets; comparing the numerical distributions of the multiple continuous variable subsets to quantify the correlation between the continuous variable and the discrete variable; According to the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous variables and discrete variables, multiple groups of strongly correlated variable pairs are obtained, and the multiple groups of strongly correlated variable pairs are merged and classified to form several groups of strongly correlated parameter groups; For each group of strongly correlated parameters, one of the variables is selected as the key parameter for dimensional compression to obtain a simplified parameter set, which is used as the core database for subsequent data modeling and analysis processes.
[0008] In a further embodiment, a continuous variable set is defined : , is the number of continuous variables, The following quantitative methods are used to quantify the correlation between continuous variables: From a continuous variable set Select all or part of the continuous variables for the selected variables Constructing a data column vector : ,in, For the variables, , Representing variables In the The value of the sampling points; Based on the data column vector Build OK Sample data matrix of columns ; For any two selected variables and variables , calculate the data column vector and data column vector The correlation coefficient between .
[0009] In a further embodiment, a discrete variable set is defined : , q is the number of discrete variables, is a variable; the following quantitative method is used to quantify the correlation between discrete variables: In the discrete variable set Select any two variables and variables , statistical variables and variables The actual frequency of joint occurrence in the sample , and construct a two-dimensional cross-frequency table; Use the two-dimensional cross-frequency table to calculate the corresponding theoretical frequency , based on actual frequency and theoretical frequency Identify variables and variables The significance probability ; Then, the correlation index of any discrete variable is Expressed as: .
[0010] In a further embodiment, the following process is used to quantify the correlation between continuous and discrete variables: Randomly select two continuous variable subsets from multiple continuous variable subsets and , using statistical distribution calculation to calculate distribution differences : ; Where, and Two groups of continuous variable subsets and The sample mean of and Two groups of continuous variable subsets and The sample variance of and Two groups of continuous variable subsets and The number of samples; The correlation strength index between continuous variables and discrete variables is Expressed as: , The significance probability between continuous variables and discrete variables is calculated by the distribution difference Sure.
[0011] In a further embodiment, the compression method of the reduced parameter set is as follows: A key parameter retention mechanism is created, and one of the variables is selected as a key parameter according to the key parameter retention mechanism to replace the original continuous variable and discrete variable; the remaining variables are regarded as redundant parameters and retained as alternative parameters; Among them, the key parameter retention mechanism includes at least one of the following: an engineering importance priority mechanism, a usage frequency-dominated priority mechanism, a data high quality priority mechanism, and a correlation-significant priority mechanism.
[0012] In a further embodiment, based on the correlation between continuous variables, Use the following method to obtain multiple sets of strongly correlated variable pairs: If , then it means the variable and variables There is a strong correlation between them, which is a strongly correlated variable pair; like , then it means the variable and variables There is a moderate correlation between them; like , then it means the variable and variables There is a weak correlation or no correlation between them; If it is positive, it is positively correlated. If it is a negative value, it is negatively correlated; is the maximum correlation threshold of continuous variables, is the minimum correlation threshold for continuous variables.
[0013] In a further embodiment, based on the correlation between discrete variables, multiple groups of strongly correlated variable pairs are obtained in the following manner: When the correlation index When it is close to 1, it means that the variable and variables There is a strong correlation between them, which is a strongly correlated variable pair; When the correlation index When it is close to 0, it means that the variable and variables There is weak correlation or no correlation between them.
[0014] In a further embodiment, based on the correlation between the continuous variable and the discrete variable, multiple groups of strongly correlated variable pairs are obtained in the following manner: When the correlation index Greater than or equal to When , it means that there is a strong correlation between the corresponding continuous variable and discrete variable, which is a strongly correlated variable pair; is the maximum correlation threshold between continuous and discrete variables; When the correlation index Less than When , it means that the corresponding continuous variable and discrete variable are weakly correlated or unrelated. is the minimum correlation threshold between continuous and discrete variables.
[0015] A TBM construction core database construction system is used to implement the TBM construction core database construction method described above, comprising: The TBM excavation parameter processing module is configured to obtain TBM excavation parameter data, determine the continuous variables and discrete variables, and group them into continuous variable sets. and a set of discrete variables ; The continuous variable correlation analysis module is configured to select an appropriate quantification method to quantify the correlation between continuous variables; The discrete variable correlation analysis module is configured to select an adaptive quantization method to quantify the correlation between discrete variables; The module for correlation analysis of continuous and discrete variables is set to construct mapping pairs of continuous and discrete variables, and to map continuous variables based on the values of discrete variables. Grouping to form multiple continuous variable subsets; comparing the numerical distributions of the multiple continuous variable subsets to quantify the correlation between the continuous variable and the discrete variable; The key parameter screening and dimension compression module is configured to obtain multiple groups of strongly correlated variable pairs based on the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous and discrete variables, and to merge and classify the multiple groups of strongly correlated variable pairs to form several groups of strongly correlated parameter groups; for each group of strongly correlated parameter groups, one of the variables is selected as the key parameter for compression to obtain a streamlined parameter set, which is then used in the subsequent data modeling and analysis process.
[0016] Beneficial effects of the present invention: The present invention systematically quantifies the correlation between TBM excavation parameters based on a mixed variable statistical test strategy. Adaptive methods are selected for correlation evaluation between continuous variables, correlation between discrete variables, and correlation between continuous and discrete variables. It effectively addresses the problems of a large number of TBM excavation parameters, mixed types, and complex relationships, and provides a high-quality variable screening basis for parameter dimensionality reduction, modeling optimization, and engineering feature extraction. Compared with traditional empirical variable selection methods, the present invention has the advantages of strong quantitative analysis capabilities, standardized operations, and high objectivity of analysis; compared with black-box feature selection algorithms, the present invention is transparent and highly interpretable, and is more suitable for practical applications in the engineering field; compared with linear analysis methods that are only applicable to numerical variables, the present invention can process continuous and discrete variables at the same time, thereby improving the adaptability and generalization ability of the model.
[0017] This paper further develops a unified variable grouping strategy and key parameter retention mechanism based on variable correlation analysis, achieving orderly compression of the parameter space, effectively eliminating redundant information, and reducing model computational complexity and the risk of overfitting. Compared to dimensionality reduction methods based on principal component analysis (PCA) or embedded models, this method fully preserves the physical meaning of the original variables, improving the engineering feasibility and practicality of subsequent analysis and interpretation.
[0018] The present invention collaboratively realizes TBM excavation parameter correlation identification and structured dimensionality reduction processing through the TBM excavation parameter processing module, continuous variable correlation analysis module, discrete variable correlation analysis module, variable grouping analysis module, continuous-discrete variable correlation analysis module and key parameter screening and dimensionality compression module, supporting the efficient construction of the TBM core database and the implementation of data modeling tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 4 is a flow chart of the TBM parameter correlation analysis and data compression method of Example 1.
[0020] Figure 2 : This is the Pearson coefficient correlation heat map of TBM continuous variables in Example 1.
[0021] Figure 3 This is the TBM discrete variable chi-square test correlation heat map of Example 1.
[0022] Figure 4 This is a heat map of the t-test correlation between TBM continuous and discrete variables in Example 1.
[0023] Figure 5 This is a system architecture diagram of a TBM parameter correlation analysis and data compression method in Example 2. DETAILED DESCRIPTION
[0024] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.
[0025] Example 1 like Figure 1 As shown, this embodiment discloses a method for constructing a TBM construction core database, including the following steps: Obtain TBM excavation parameter data, determine the continuous variables and discrete variables, and group them to obtain a continuous variable set and a set of discrete variables ; Select an appropriate quantification method to quantify the correlation between continuous variables and the correlation between discrete variables respectively; Construct a mapping pair between continuous variables and discrete variables, and convert the continuous variable set into Grouping to form multiple continuous variable subsets; comparing the numerical distributions of the multiple continuous variable subsets to quantify the correlation between the continuous variable and the discrete variable; According to the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous variables and discrete variables, multiple groups of strongly correlated variable pairs are obtained, and the multiple groups of strongly correlated variable pairs are merged and classified to form several groups of strongly correlated parameter groups; For each group of strongly correlated parameters, one of the variables is selected as the key parameter for dimensionality compression to obtain a simplified parameter set, which is used in the subsequent data modeling and analysis process.
[0026] Table 1 Raw data of some TBM parameters As shown in the partial data in Table 1, the TBM excavation parameter data described in this embodiment can be obtained through the TBM's own onboard data monitoring platform, or by accessing the raw data through the TBM's cloud-edge collaborative monitoring platform. The raw data in this embodiment includes at least 2,000 samples, covering 12 parameters, including cutterhead speed, cutterhead thrust, cutterhead torque, propulsion system oil pressure, temperature, and motor current. After data preprocessing (e.g., missing value filling and outlier removal), correlation analysis was performed on different types of variables.
[0027] Based on this, the continuous variable set described in this embodiment is and a set of discrete variables The grouping steps are as follows: Step 1.1: Obtain raw parameter data from the TBM excavation construction site. This data is collected in real time by multiple sensors and covers multiple working condition parameters during the excavation process. Step 1.2: Perform unified processing on TBM excavation parameter data, including but not limited to: filling in missing values, removing outliers, and time alignment, to ensure data integrity and consistency and improve the reliability of subsequent statistical analysis results.
[0028] Step 1.3: Classify the variables according to their physical meaning, data type, value form, and engineering experience: Continuous variables include but are not limited to: propulsion speed, cutterhead speed, propulsion force, cutterhead torque, motor current, motor voltage, motor torque, oil temperature and other parameters that are expressed in the form of continuous real numbers and can be quantified and measured; Discrete variables include but are not limited to: fan start and stop, motor switch, reducer switch, construction status code, excavation mode selection, sensor operating status, and other parameters expressed in the form of categories or states.
[0029] Step 1.4: Based on the grouping in step 1.3, obtain a set of continuous variables and a set of discrete variables , and the continuous variable set Recorded as .in, is the number of continuous variables, is a variable. Combined with the above example, the continuous variable set The elements in are: propulsion speed, cutter head speed, propulsion force, cutter head torque, motor current, motor voltage, motor torque, oil temperature; that is, the current The value of is 8.
[0030] This example uses the Pearson correlation coefficient, chi-square independence test, and t-test to conduct correlation analysis based on differences in variable types, aiming to improve the accuracy and adaptability of the assessment results. The Pearson correlation coefficient effectively measures the degree of linear correlation between continuous variables and is applicable to data with a well-defined mean and variance that follow an approximately normal distribution. The chi-square independence test is suitable for determining independence between discrete variables and can reflect the distributional association characteristics of categorical variables. The t-test is applicable to combinations of continuous and binary discrete variables, measuring the degree of correlation between the two by comparing the significance of the mean differences between continuous variables in different categories. Compared with a single correlation quantification method, this type of categorical matching statistical test can fully utilize the statistical characteristics of each variable, thereby revealing the correlation between TBM excavation parameters in a more scientific and comprehensive manner.
[0031] Continuous variable set It can be further expressed as: {propulsion speed , Cutter head speed , propulsion , cutter head torque , motor current , motor voltage , motor torque , oil temperature }, the following quantitative methods are used to quantify the correlation between continuous variables: Select all continuous variables and for each variable Constructing a data column vector , , the construction method is as follows: ,in, For the variables, , Representing variables In the It is worth noting that the sampling points described in this embodiment are sampled samples of the variable at different time points.
[0032] Based on the data column vector Build OK Sample data matrix of columns ; to propel speed For example, there are 2000 samples in total, so the corresponding data column vector It is a sample data matrix with 2000 rows and 8 columns.
[0033] For any two selected variables and variables , calculate the data column vector and data column vector The correlation coefficient between In this example, the Pearson correlation coefficient is used to calculate the correlation coefficient , the calculation formula is as follows: ; Where, and are data column vectors and data column vector The number of elements, and Variables and variables The sample mean of is the data column vector Middle elements, is the data column vector Middle elements.
[0034] Based on the above calculation, we get the data column vector and data column vector The correlation coefficient between , then use the following method to obtain multiple sets of strongly correlated variable pairs: Use the following method to obtain multiple sets of strongly correlated variable pairs: If , then it means the variable and variables There is a strong correlation between them, which is a strongly correlated variable pair; like , then it means the variable and variables There is a moderate correlation between them; like , then it means the variable and variables There is a weak correlation or no correlation between them; If it is positive, it is positively correlated. If it is a negative value, it is negatively correlated; is the maximum correlation threshold of continuous variables, is the minimum correlation threshold of the continuous variable. In this embodiment, , .
[0035] Combine Figure 2 For example, the correlation coefficients between propulsion speed and other continuous variables (cutter head speed, propulsion force, cutter head torque, motor current, motor voltage, motor torque, and oil temperature) are calculated as follows: 、 、 、 、 、 and Combining the methods of obtaining multiple sets of strongly correlated variable pairs, it can be seen that the propulsion speed and the following variables constitute a strongly correlated variable pair: cutterhead speed, propulsion force, cutterhead torque, and motor torque.
[0036] In another embodiment, the discrete variable set Recorded as , q is the number of discrete variables, is a variable. Combined with the above example, the discrete variable set The elements in are: fan start and stop, motor switch, reducer temperature switch, gear oil system; at this time, q The value of is 4.
[0037] In combination with the above examples, the discrete variable set described in this embodiment is Expressed as: {fan start and stop , motor switch , reducer temperature switch , gear oil system is normal }, and the following quantitative methods are used to quantify the correlation between discrete variables: In the discrete variable set Select any two variables and variables , statistical variables and variables The actual frequency of joint occurrence in the sample , and construct a two-dimensional cross-frequency table (contingency table); Use the two-dimensional cross-frequency table to calculate the corresponding theoretical frequency , based on actual frequency and theoretical frequency Identify variables and variables The significance probability ; Then, the correlation index of any discrete variable is Expressed as: .
[0038] Furthermore, the theoretical frequency is calculated using the following method: :Assume that any two variables are selected and variables To be independent of each other, calculate the theoretical frequency of each cell in the two-dimensional cross-frequency table based on the marginal distribution , the calculation formula is as follows: ;in, is the sum of the row frequencies of the cells, is the sum of the column frequencies of the cells, is the number of cells.
[0039] Furthermore, the statistics are calculated based on all cells N using the following formula: : .
[0040] According to the calculated statistics and degrees of freedom Query the chi-square distribution table to obtain the corresponding significance probability , where the significance probability The smaller the value, the more actual frequency The greater the difference from the theoretical frequency of the independent variable, the stronger the correlation between the variables.
[0041] In order to quantify the strength of correlation, we define the correlation index of any discrete variable Expressed as: .
[0042] Furthermore, based on the correlation between discrete variables, multiple sets of strongly correlated variable pairs are obtained in the following way: When the correlation index When it is close to 1, it means that the variable and variables There is a strong correlation between them, which is a strongly correlated variable pair; ,Right now .
[0043] When the correlation index When it is close to 0, it means that the variable and variables There is a weak correlation or no correlation between them; .
[0044] like Figure 3 As shown, this embodiment selects the fan start and stop and motor switch As an example, the corresponding significance probability is calculated is 0.01, the correlation index can be determined Therefore, it can be considered that there is a strong correlation between the start and stop of the fan and the motor switch, that is, a strongly correlated variable pair.
[0045] In order to facilitate the understanding of the grouping process of multiple continuous variable subsets, this embodiment will use the continuous variable set : {Advance speed , Cutter head speed , propulsion , cutter head torque , motor current , motor voltage , motor torque , oil temperature } and discrete variable sets :{Fan start and stop , motor switch , reducer temperature switch , gear oil system is normal For example, construct a mapping pair of continuous variables and discrete variables to promote speed and fan start and stop For example, according to the fan start and stop The various values of will correspond to the advancement speed Furthermore, if the fan starts and stops The value of is {0, 1}, then the fan start and stop The propulsion speed is 0 The entire sample forms a continuous variable subset , extract fan start and stop The propulsion speed is 1 The entire sample forms a continuous variable subset .
[0046] Based on this, we use and Instead of a specific value, the upper position represents two arbitrarily selected subsets of continuous variables and , further understood as, will promote the speed According to the fan start and stop The continuous variable subsets are divided into two groups, namely and continuous variable subsets . Further, the continuous variable subset Start and stop the fan The propulsion speed set when (wind is off) is recorded as ; Continuous variable subset Start and stop the fan The propulsion speed set when (fan is on) is recorded as .
[0047] The following process is used to quantify the correlation between continuous and discrete variables and to calculate the distribution difference : ; Where, and Two groups of continuous variable subsets and The sample mean of and Two groups of continuous variable subsets and The sample variance of and Two groups of continuous variable subsets and The number of samples; According to the calculated distribution difference and degrees of freedom Query the distribution table or calculate the corresponding significance probability In order to facilitate quantitative analysis, the correlation strength index between continuous variables and discrete variables is Expressed as: .
[0048] Furthermore, multiple sets of strongly correlated variable pairs are obtained using the following method: When the correlation index Greater than or equal to When , it means that there is a strong correlation between the corresponding continuous variable and discrete variable, which is a strongly correlated variable pair; is the maximum correlation threshold between continuous variables and discrete variables, such as When .
[0049] When the correlation index Less than When , it means that the corresponding continuous variable and discrete variable are weakly correlated or unrelated. is the minimum correlation threshold between continuous variables and discrete variables, For example, .
[0050] Combine Figure 4 In this embodiment, the correlation strength index between the propulsion speed (C1) and the fan start and stop (D1) is The value is 0.98, which indicates that there is a significant correlation between propulsion speed and fan start and stop.
[0051] Based on the correlations between continuous variables, between discrete variables, and between continuous and discrete variables, we obtain multiple sets of strongly correlated variable pairs. These pairs are then merged and classified into several groups of strong phases. The corresponding methods for obtaining strongly correlated variable pairs have been described above, so we will not elaborate on them in detail.
[0052] As in the above example, the continuous variable propulsion speed and the cutter head rotation speed, propulsion force, cutter head torque and motor torque form a group of strongly correlated parameters, and so on.
[0053] Furthermore, the compression method of the simplified parameter set is as follows: A key parameter retention mechanism is created, and one of the variables is selected as a key parameter according to the key parameter retention mechanism to replace the original continuous variable and discrete variable; the remaining variables are regarded as redundant parameters and retained as alternative parameters; Among them, the key parameter retention mechanism includes at least one of the following: an engineering importance priority mechanism, a usage frequency-dominated priority mechanism, a data high quality priority mechanism, and a correlation-significant priority mechanism.
[0054] For a better understanding, the engineering importance priority mechanism described in this embodiment is to give priority to corresponding variables with high importance in engineering experience; the frequency-dominant priority mechanism is to give priority to variables with high frequency of use; the data high quality priority mechanism is to give priority to variables with good data stability, low missing rate and easy collection; the correlation significance priority mechanism is to give priority to variables with the most significant relationship with the target variable.
[0055] Based on this, the simplified parameter set is used as the core database for the subsequent data modeling and analysis process. The simplified parameter set finally selected includes: and a reduced set of discrete variables , combined with the above examples, , .
[0056] Therefore, the above technical solution grouped strongly correlated parameters into the same group. For each group of strongly correlated tunneling parameters, one parameter from each group was retained as the key parameter, achieving parameter-dimensional compression of TBM tunneling parameter data. Through variable grouping and screening, the original 12 parameters were compressed to three representative parameters, achieving a parameter-dimensional compression rate exceeding 75%, significantly reducing modeling computational costs and redundancy risks.
[0057] Example 2 In order to implement the TBM construction core database construction method described in Example 1, Figure 5 This embodiment discloses a TBM construction core database construction system, including: The TBM excavation parameter processing module is configured to obtain TBM excavation parameter data, determine the continuous variables and discrete variables, and group them into continuous variable sets. and a set of discrete variables ; The continuous variable correlation analysis module is configured to select an appropriate quantification method to quantify the correlation between continuous variables; The discrete variable correlation analysis module is configured to select an adaptive quantization method to quantify the correlation between discrete variables; The module for correlation analysis of continuous and discrete variables is set to construct mapping pairs of continuous and discrete variables, and to map continuous variables based on the values of discrete variables. Grouping to form multiple continuous variable subsets; comparing the numerical distributions of the multiple continuous variable subsets to quantify the correlation between the continuous variable and the discrete variable; The key parameter screening and dimension compression module is configured to obtain multiple groups of strongly correlated variable pairs based on the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous and discrete variables, and to merge and classify the multiple groups of strongly correlated variable pairs to form several groups of strongly correlated parameter groups; for each group of strongly correlated parameter groups, one of the variables is selected as the key parameter for compression to obtain a streamlined parameter set, which is then used in the subsequent data modeling and analysis process.
Claims
1. A method for constructing a core database for TBM construction, characterized in that: The following steps are involved: Obtain TBM excavation parameter data, determine the continuous variables and discrete variables, and group them to obtain a continuous variable set and a set of discrete variables ; Select an appropriate quantification method to quantify the correlation between continuous variables and the correlation between discrete variables respectively; Construct a mapping pair between continuous variables and discrete variables, and convert the continuous variable set into Grouping to form multiple continuous variable subsets; comparing the numerical distributions of the multiple continuous variable subsets to quantify the correlation between the continuous variable and the discrete variable; According to the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous variables and discrete variables, multiple groups of strongly correlated variable pairs are obtained, and the multiple groups of strongly correlated variable pairs are merged and classified to form several groups of strongly correlated parameter groups; For each group of strongly correlated parameters, one of the variables is selected as the key parameter for dimensional compression to obtain a simplified parameter set, which is used as the core database for subsequent data modeling and analysis processes.
2. A TBM construction core database construction method according to claim 1, characterized in that: Define a set of continuous variables : , is the number of continuous variables, is a variable; The following quantitative methods were used to quantify the correlation between continuous variables: From a continuous variable set Select all or part of the continuous variables for the selected variables Constructing a data column vector : ,in, For the variables, , Representing variables In the The value of the sampling points; Based on the data column vector Build OK Sample data matrix of columns ; For any two selected variables and variables , calculate the data column vector and data column vector The correlation coefficient between .
3. A TBM construction core database construction method according to claim 1, characterized in that: Discrete variable set : , q is the number of discrete variables, is a variable; The correlation between discrete variables is quantified using the following quantitative methods: In the discrete variable set Select any two variables and variables , statistical variables and variables The actual frequency of joint occurrence in the sample , and construct a two-dimensional cross-frequency table; Use the two-dimensional cross-frequency table to calculate the corresponding theoretical frequency , based on actual frequency and theoretical frequency Identify variables and variables The significance probability ; Then, the correlation index of any discrete variable is Expressed as: .
4. A TBM construction core database construction method according to claim 1, characterized in that: The following procedure was used to quantify the correlation between continuous and discrete variables: Randomly select two continuous variable subsets from multiple continuous variable subsets and , using statistical distribution calculation to calculate distribution differences : ; Where, and Two groups of continuous variable subsets and The sample mean of and Two groups of continuous variable subsets and The sample variance of and Two groups of continuous variable subsets and The number of samples; The correlation strength index between continuous variables and discrete variables is Expressed as: , The significance probability between continuous variables and discrete variables is calculated by the distribution difference Sure.
5. A TBM construction core database construction method according to claim 1, characterized in that: The compression method for the simplified parameter set is as follows: Creating a key parameter retention mechanism, and selecting one of the variables as a key parameter according to the key parameter retention mechanism to replace the original continuous variable and discrete variable; The remaining variables were considered redundant parameters and retained as alternative parameters; Among them, the key parameter retention mechanism includes at least one of the following: an engineering importance priority mechanism, a usage frequency-dominated priority mechanism, a data high quality priority mechanism, and a correlation-significant priority mechanism.
6. A TBM construction core database construction method according to claim 2, characterized in that: According to the correlation between continuous variables, multiple groups of strongly correlated variable pairs are obtained in the following way: , then it means the variable and variables There is a strong correlation between them, which is a strongly correlated variable pair; like , then it means the variable and variables There is a moderate correlation between them; like , then it means the variable and variables There is a weak correlation or no correlation between them; If it is positive, it is positively correlated. If it is a negative value, it is negatively correlated; is the maximum correlation threshold of continuous variables, is the minimum correlation threshold for continuous variables.
7. A TBM construction core database construction method according to claim 3, characterized in that: According to the correlation between discrete variables, multiple groups of strongly correlated variable pairs are obtained in the following way: When the correlation index When it is close to 1, it means that the variable and variables There is a strong correlation between them, which is a strongly correlated variable pair; When the correlation index When it is close to 0, it means that the variable and variables There is weak correlation or no correlation between them.
8. A TBM construction core database construction method according to claim 4, characterized in that: Based on the correlation between continuous variables and discrete variables, multiple sets of strongly correlated variable pairs are obtained in the following way: When the correlation index Greater than or equal to When , it means that there is a strong correlation between the corresponding continuous variable and discrete variable, which is a strongly correlated variable pair; is the maximum correlation threshold between continuous and discrete variables; When the correlation index Less than When , it means that the corresponding continuous variable and discrete variable are weakly correlated or unrelated. is the minimum correlation threshold between continuous and discrete variables.
9. A TBM construction core database construction system, used to implement the TBM construction core database construction method according to any one of claims 1 to 8, characterized in that: include: The TBM excavation parameter processing module is configured to obtain TBM excavation parameter data, determine the continuous variables and discrete variables, and group them into continuous variable sets. and a set of discrete variables ; The continuous variable correlation analysis module is configured to select an appropriate quantification method to quantify the correlation between continuous variables; The discrete variable correlation analysis module is configured to select an adaptive quantization method to quantify the correlation between discrete variables; The module for correlation analysis of continuous and discrete variables is set to construct mapping pairs of continuous and discrete variables, and to map continuous variables based on the values of discrete variables. Grouping to form multiple continuous variable subsets; comparing the numerical distributions of the multiple continuous variable subsets to quantify the correlation between the continuous variable and the discrete variable; The key parameter screening and dimension compression module is configured to obtain multiple groups of strongly correlated variable pairs based on the correlation between continuous variables, the correlation between discrete variables, and the correlation between continuous and discrete variables, and to merge and classify the multiple groups of strongly correlated variable pairs to form several groups of strongly correlated parameter groups; for each group of strongly correlated parameter groups, one of the variables is selected as the key parameter for compression to obtain a streamlined parameter set, which is then used in the subsequent data modeling and analysis process.
Citation Information
Patent Citations
Risk rapid assessment method for large power grid
CN111160772A
Shield excavation face stratum property real-time sensing and tunneling parameter adjusting method and system
CN114370284A
High-dimensional feature extraction method and device, computer equipment and storage medium
CN114627965A
Real-time multi-step prediction method for tunneling key parameters of shield tunneling machine based on ST-GCN-LSTM
CN120277367A
Traffic accident severity influence factor analysis method based on local cascade integration
CN120277552A