Boulder stratum identification method and device based on tunneling data, equipment and medium
By extracting and clustering the excavation data in shield construction, and identifying the lonely stone strata in unsupervised and supervised models, the problem of identifying the lonely stone strata in shield construction is solved, and accurate identification and early warning is achieved to reduce construction risks.
Patent Information
- Application Number
- CN202510116027.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology is difficult to accurately identify the lonely stone formation in shield construction, resulting in construction hazards and safety hazards, and lacks real-time monitoring measures.
By obtaining feature items of excavation data, a pre-configured unsupervised clustering model is used for clustering processing, combined with a supervised strata identification model, a solitary strata is identified and an early warning is issued.
Accurate identification of lonely stone formations has been achieved, reducing construction hazards and improving construction safety.
Smart Images

Figure CN120336883A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underground detection, and particularly to a method, device, equipment and medium for identifying boulder formations based on tunneling data. Background Art
[0002] Boulder formations have unique geological characteristics in shield tunneling. Their irregular shapes and uneven distributions make shield tunneling face more complex and difficult challenges. It is difficult to accurately judge the type and nature of boulder formations and even more difficult to predict their impact on shield tunneling based solely on experience and intuition. In actual projects, if the boulder formation cannot be identified in a timely and accurate manner, it will cause great harm to shield tunneling, such as damaging the shield machine, delaying the construction period, etc., and even leading to safety accidents. Therefore, it is crucial to study the development trend and law of monitoring data of boulder formations, but there is currently a lack of monitoring measures for real-time identification of boulder formations during tunneling. Summary of the Invention
[0003] In view of the problems existing in the prior art, the present invention provides a method, device, equipment and medium for identifying boulder formations based on tunneling data.
[0004] The present invention provides a method for identifying boulder formations based on tunneling data, including:
[0005] Obtaining the data characteristics of the tunneling data of the rock mass; the data characteristics are composed of data corresponding to multiple feature items;
[0006] Performing unsupervised clustering processing on the data characteristics respectively based on a variety of pre-configured unsupervised clustering models to obtain an unsupervised clustering label result corresponding to each feature item and matching the preset number of formation types;
[0007] Constructing new data characteristics based on the data characteristics and the unsupervised clustering label result;
[0008] Inputting the new data characteristics into a supervised formation identification model to obtain the formation type of the rock mass output by the supervised formation identification model;
[0009] When it is determined that the formation type is a boulder formation, an early warning message is sent;
[0010] Wherein, the supervised formation identification model is a model obtained by machine learning training with the data characteristics determined based on sample tunneling data and the clustering label result determined by performing unsupervised clustering on the data characteristics determined based on sample tunneling data as inputs, and the formation type label corresponding to the sample tunneling data as an output, and is used for identifying the formation of tunneling data.
[0011] A method for identifying boulder formations based on tunneling data provided by the present invention, the data characteristics of obtaining the tunneling data of the rock mass include:
[0012] Obtain a plurality of pre-configured feature items, and construct the data characteristics of the tunneling data of the rock mass based on the data corresponding to the plurality of feature items;
[0013] The obtaining steps of the plurality of pre-configured feature items include:
[0014] Obtain the sample tunneling data of the rock mass, and the sample tunneling data includes a plurality of data items;
[0015] Perform feature extraction on the tunneling data respectively based on a plurality of pre-configured feature extraction strategies to obtain a plurality of feature extraction results; the feature extraction results are the sorting results of each of the data items;
[0016] Based on the plurality of feature extraction results, screen out a plurality of feature items from the plurality of data items.
[0017] A method for identifying boulder formations based on tunneling data provided by the present invention, the screening out a plurality of feature items from the plurality of data items based on the plurality of feature extraction results includes:
[0018] Obtain the first weight value corresponding to each feature extraction strategy, and obtain the second weight value of each data item in the feature extraction result corresponding to each feature extraction strategy;
[0019] Based on the plurality of feature extraction results, and the first weight value and the second weight value corresponding to the feature extraction results, determine the comprehensive feature extraction result; the comprehensive feature extraction result is the sorting result of each of the data items;
[0020] Based on the comprehensive feature extraction result and the screening number of the feature items, screen out a plurality of feature items from the plurality of data items.
[0021] A method for identifying boulder formations based on tunneling data provided by the present invention, the obtaining steps of the supervised formation identification model include:
[0022] Obtain the sample tunneling data of the rock mass, and the sample tunneling data includes a plurality of data items;
[0023] Perform feature extraction on the tunneling data respectively based on a plurality of pre-configured feature extraction strategies to obtain a plurality of feature extraction results; the feature extraction results are the sorting results of each of the data items;
[0024] Based on the plurality of feature extraction results, screen out a plurality of feature items from the plurality of data items;
[0025] Obtain the data characteristics of the sample tunneling data of the rock mass based on screening out multiple feature items;
[0026] Based on a variety of pre-configured unsupervised clustering models, perform unsupervised clustering processing on the data characteristics respectively, and obtain the unsupervised clustering label results corresponding to each feature item and matching the preset number of formation types;
[0027] Construct new data characteristics of the sample tunneling data based on the data characteristics and the unsupervised clustering label results;
[0028] Divide the new data characteristics of the sample tunneling data into a training set and a test set;
[0029] Train a model based on the new data characteristics in the training set and the formation type labels of the sample tunneling data, and evaluate the model based on the new data characteristics in the test set and the formation type labels of the sample tunneling data to obtain a supervised formation identification model.
[0030] According to a method for identifying boulder formations based on tunneling data provided by the present invention, during the process of evaluating the model based on the new data characteristics in the test set and the formation type labels of the sample tunneling data, when it is determined that the classification evaluation index, confusion matrix, and performance index all meet the conditions, determine the supervised formation identification model.
[0031] According to a method for identifying boulder formations based on tunneling data provided by the present invention, before obtaining the data characteristics of the tunneling data of the rock mass, the method further includes obtaining the tunneling data of the rock mass, including:
[0032] Obtain the tunneling data of multiple tunneling sections based on a preset extraction strategy for tunneling sections;
[0033] Perform anomaly analysis on the tunneling data of each tunneling section based on a preset anomaly identification strategy, determine the anomaly type of each tunneling section, and configure an anomaly processing strategy for the tunneling data of the corresponding tunneling section based on the anomaly type;
[0034] When it is determined that there is an anomaly processing strategy without an interpolation method in the configured anomaly processing strategy, configure a first processing flow, and execute the first processing flow on the tunneling data based on the anomaly processing strategy without an interpolation method;
[0035] After the first processing flow is executed, perform noise reduction processing on all the tunneling data;
[0036] Configure a second processing flow, and perform interpolation processing on the noise-reduced tunneling data based on the anomaly processing strategy with an interpolation method to obtain the tunneling data of the rock mass.
[0037] A method for identifying boulder formations based on tunneling data provided by the present invention performs interpolation processing on the tunneling data after noise reduction to obtain the tunneling data of the rock mass, including:
[0038] Determine the data missing type corresponding to the tunneling section and the missing rate of each missing area corresponding to the tunneling section;
[0039] Based on the data missing type and the missing rate, determine the interpolation measure corresponding to the missing area.
[0040] The present invention also provides an apparatus for identifying boulder formations based on tunneling data, including:
[0041] An acquisition module for acquiring the data characteristics of the tunneling data of the rock mass; the data characteristics are composed of data corresponding to multiple feature items;
[0042] A first identification module for performing unsupervised clustering processing on the data characteristics respectively based on a plurality of pre-configured unsupervised clustering models to obtain an unsupervised clustering label result corresponding to each feature item and matching the preset number of formation types;
[0043] A construction module for constructing new data characteristics based on the data characteristics and the unsupervised clustering label result;
[0044] A second identification module for inputting the new data characteristics into a supervised formation identification model to obtain the formation type of the rock mass output by the supervised formation identification model;
[0045] A processing module for sending out a warning message when the formation type is determined to be a boulder formation;
[0046] Wherein, the supervised formation identification model is a model trained by machine learning with the data characteristics determined based on the sample tunneling data and the clustering label result determined by performing unsupervised clustering on the data characteristics determined based on the sample tunneling data as inputs, and the formation type label corresponding to the sample tunneling data as an output, and is used for identifying the formation of the tunneling data.
[0047] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements any one of the above-mentioned methods for identifying boulder formations based on tunneling data.
[0048] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the above-mentioned methods for identifying boulder formations based on tunneling data.
[0049] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements any one of the above-described boulder stratum identification methods based on tunneling data.
[0050] A boulder stratum identification method, device, equipment and medium based on tunneling data provided by the present invention perform unsupervised clustering processing on the data features of tunneling data respectively through a variety of pre-configured unsupervised clustering models, obtain an unsupervised clustering label result corresponding to the matching of the data features and the number of stratum types, then construct new data features based on the data features and the unsupervised clustering label result, and finally input the new data features into a supervised stratum identification model to obtain the stratum type of the rock mass output by the supervised stratum identification model, realizing a richer and more comprehensive feature representation of the tunneling data and accurately identifying the stratum type of the rock mass. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 is a schematic flowchart of the boulder stratum identification method based on tunneling data provided by the present invention.
[0053] Figure 2 is a schematic flowchart of the comprehensive extraction process of boulder stratum features provided by the present invention.
[0054] Figure 3 is an overall process diagram of the identification of the boulder stratum provided by the present invention.
[0055] Figure 4 is a schematic structural diagram of the boulder stratum identification device based on tunneling data provided by the present invention.
[0056] Figure 5 is a schematic structural diagram of the electronic equipment provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0058] The following will be combined withFigures 1 - 5 Describe a boulder stratum identification method, device, equipment and medium based on tunneling data of the present invention.
[0059] Figure 1 The flowchart of a boulder stratum identification method based on tunneling data provided by the present invention is shown. Refer to Figure 1 This method includes the following steps:
[0060] Step 11: Obtain the data characteristics of the tunneling data of the rock mass, and the data characteristics are composed of data corresponding to multiple feature items.
[0061] Step 12: Based on a variety of pre-configured unsupervised clustering models, perform unsupervised clustering processing on the data characteristics respectively, and obtain unsupervised clustering label results corresponding to each feature item and matching the preset number of stratum types.
[0062] Step 13: Construct new data characteristics based on the data characteristics and the unsupervised clustering label results.
[0063] Step 14: Input the new data characteristics into a supervised stratum identification model to obtain the stratum type of the rock mass output by the supervised stratum identification model.
[0064] Step 15: When the stratum type is determined to be a boulder stratum, send out a warning message.
[0065] Among them, the supervised stratum identification model is a model obtained through machine learning training by taking the data characteristics determined based on the sample tunneling data, the clustering label results determined by performing unsupervised clustering on the data characteristics determined based on the sample tunneling data as inputs, and the stratum type label corresponding to the sample tunneling data as outputs, and is used for stratum identification of tunneling data.
[0066] Regarding steps 11 to 15, it should be noted that the boulder stratum has unique geological characteristics in shield tunneling construction. Its irregular shape and uneven distribution make shield tunneling construction face more complex and difficult challenges. It is difficult to accurately judge the type and nature of the boulder stratum only by experience and intuition, and it is even more difficult to predict its impact on shield tunneling construction. In actual projects, if the boulder stratum cannot be identified in a timely and accurate manner, it will cause great harm to shield tunneling construction, such as damaging the shield machine, delaying the construction period, etc., and even leading to the occurrence of safety accidents. Therefore, it is crucial to study the development trend and law of the monitoring data of the boulder stratum.
[0067] During the process of shield tunneling data acquisition, measured data such as cutterhead rotation speed, tunneling speed, and soil chamber pressure not only contain the characteristics of geotechnical engineering but also have typical big data characteristics. Generally speaking, it has the following characteristics: multi-source heterogeneity, spatio-temporality, real-time nature, continuity, relevance, massiveness, and low value density. That is to say, the tunneling data includes various types of data. However, the types of data that can identify the boulder stratum may be fewer. Therefore, several data types (feature items as data features) that can identify the boulder stratum can be obtained in advance through analysis, and the data corresponding to these data types are constructed into the data features of the tunneling data. For example, the oil temperature of the gear oil, the pressure of the right middle soil chamber, the oil temperature of the main fuel tank, the total cumulative amount of EP2 lubricating grease, the cumulative amount of the current ring of the shield tail seal, the cumulative amount of the current ring of the HBW seal grease, the current cumulative working time of the cutterhead, the VMT guiding pitch angle, the oil temperature of the screw conveyor, and the cutterhead rotation speed, a total of 10 types of data are helpful for the identification of the boulder stratum. Therefore, the required data types can be obtained in advance, and then for the collected tunneling data, the corresponding data content is collected according to the data type, and then the data content is constructed to obtain the data features of the tunneling data.
[0068] In the present invention, due to the lack of a process for real-time identifying the boulder stratum based on the tunneling data, it is impossible to fully fit and capture the characteristics of the tunneling data, and it also leads to different and deviated identification results when different processes for identifying the boulder stratum are adopted, which is not conducive to the accuracy of boulder stratum identification. Therefore, in order to eliminate this deviation, it is proposed to combine the clustering results of different strata, and fuse the stratum type clustering results with the overall data characteristics, which can increase the richness of data and the accuracy of stratum identification during the identification process. Therefore, in the present invention, multiple unsupervised clustering models need to be pre-configured first. These models can be conventional classification models (such as random forest RF, support vector classifier, K-nearest neighbor algorithm, multi-layer perceptron, gradient boosting classifier), or newly designed clustering models. Based on the pre-configured multiple unsupervised clustering models, unsupervised clustering processing is respectively performed on the data features to obtain unsupervised clustering label results corresponding to each feature item and matching the preset number of stratum types. That is to say, the collected tunneling data is clustered, and the data is clustered and divided. If the configured stratum types are 3 types, then the tunneling data can be clustered and divided into three clusters, and each cluster is configured with a corresponding type label.
[0069] In the present invention, by combining the clustering label results of multiple unsupervised clustering models and then integrating these clustering label results into the data features, the adaptability to tunneling data is improved, and then the comprehensive processing ability of the supervised formation recognition model for the overall data is facilitated, so as to obtain a more accurate recognition result of the formation type. That is to say, the supervised formation recognition model is a model trained by machine learning for identifying the formation of tunneling data, which takes the data features determined based on the sample tunneling data, the clustering label results determined by unsupervised clustering of the data features determined based on the sample tunneling data as inputs, and the formation type labels corresponding to the sample tunneling data as outputs.
[0070] The boulder formation has unique geological characteristics in shield tunneling. Its irregular shape and uneven distribution make shield tunneling face more complex and difficult challenges. In the boulder formation, the cutterhead thrust fluctuates greatly, indicating that the geological conditions are more complex and the rock formation is harder, and a greater thrust is required to advance the shield machine. Therefore, when it is determined that the formation type of the current tunneling area is the boulder formation, a warning message is issued.
[0071] The boulder formation recognition method based on tunneling data provided by the present invention performs unsupervised clustering processing on the data features of tunneling data respectively based on a variety of pre-configured unsupervised clustering models, obtains unsupervised clustering label results corresponding to the matching of the data features and the number of formation types, then constructs new data features based on the data features and the unsupervised clustering label results, and finally inputs the new data features into the supervised formation recognition model to obtain the formation type of the rock mass output by the supervised formation recognition model, realizing a richer and more comprehensive feature representation of the tunneling data and accurately identifying the formation type of the rock mass.
[0072] In a further method of the above method, it is mainly an explanatory description of the processing process of obtaining the data features of the tunneling data of the rock mass, which is specifically as follows:
[0073] Obtain a plurality of pre-configured feature items, and construct the data features of the tunneling data of the rock mass based on the data corresponding to the plurality of feature items.
[0074] Among them, the obtaining steps of the plurality of pre-configured feature items include:
[0075] Obtain the sample tunneling data of the rock mass, and the sample tunneling data includes a plurality of data items.
[0076] Perform feature extraction on the tunneling data respectively based on a variety of pre-configured feature extraction strategies to obtain a plurality of feature extraction results; the feature extraction results are the sorting results of each data item.
[0077] Based on the plurality of feature extraction results, screen out a plurality of feature items from the plurality of data items.
[0078] In this regard, it should be noted that the multiple pre-configured feature extraction strategies can be conventional feature selection algorithms (such as Pearson correlation, Spearman correlation, analysis of variance method, random forest feature importance, mutual information method), or newly designed feature selection algorithms.
[0079] For each feature extraction strategy, for the feature extraction of tunneling data, the sorting results of each data item can be obtained. This sorting result characterizes the contribution degree of each data item to the subsequent formation classification in different strategies.
[0080] After obtaining multiple feature extraction results, based on the multiple feature extraction results, multiple feature items are screened out from multiple data items. The data items screened out at this time are used as the required feature items. These feature items are more comprehensive and adaptable in terms of the contribution degree to the subsequent formation classification than a single strategy.
[0081] Furthermore, the processing process of screening multiple feature items from multiple data items based on multiple feature extraction results is explained as follows:
[0082] Obtain the first weight value corresponding to each feature extraction strategy, and obtain the second weight value of each data item in the feature extraction result corresponding to each feature extraction strategy under the feature extraction strategy.
[0083] Based on multiple feature extraction results, as well as the first weight value and the second weight value corresponding to the feature extraction results, determine the comprehensive feature extraction result; the comprehensive feature extraction result is the sorting result of each data item.
[0084] Based on the comprehensive feature extraction result and the screening number of feature items, screen out multiple feature items from multiple data items.
[0085] In this regard, it should be noted that in the present invention, by ranking or scoring the feature extraction results of each feature extraction strategy, and then different feature extraction strategies have different precision degrees in the feature selection process, different weight values can be configured, which are regarded as the first weight value here. And different data items have different contribution degrees to the formation type classification, and different weight values can be configured, which are regarded as the second weight value here. Then, each feature extraction result is weighted and summed with the corresponding weight value, and finally a comprehensive feature ranking or score is obtained, which is regarded as the comprehensive feature extraction result.
[0086] Finally, according to the comprehensive ranking or score, the features are finally sorted. The higher the ranking of the feature, the more important the feature is considered, because they are recognized as having higher value in multiple feature selection algorithms.
[0087] See Figure 2The flow diagram of the comprehensive method for extracting the characteristics of boulder formations is shown. The characteristic factors in the figure are the data items x1... x of the tunneling data n , and the target parameter is the boulder formation type. After sorting the data items by the feature extraction methods m1... mn, different characteristic results Rank m1-y{x1... x n} are obtained,..., Rank mn-y{x1... x n}, and then weighted summation is performed to obtain the final ranking Rank m-y{x1... x n .
[0088] In a further method of the above method, the acquisition steps of the supervised formation recognition model are mainly explained as follows:
[0089] Obtain the sample tunneling data of the rock mass, and the sample tunneling data includes multiple data items.
[0090] Based on multiple pre-configured feature extraction strategies, feature extraction is respectively performed on the tunneling data to obtain multiple feature extraction results; the feature extraction results are the sorting results of each data item.
[0091] Based on multiple feature extraction results, multiple feature items are screened out from multiple data items.
[0092] Based on the screened multiple feature items, the data characteristics of the sample tunneling data of the rock mass are obtained.
[0093] Based on multiple pre-configured unsupervised clustering models, unsupervised clustering processing is respectively performed on the data characteristics to obtain the unsupervised clustering label results corresponding to each feature item and matching the preset number of formation types.
[0094] Based on the data characteristics and the unsupervised clustering label results, new data characteristics of the sample tunneling data are constructed.
[0095] Based on the new data characteristics of the sample tunneling data, it is divided into a training set and a test set.
[0096] Based on the new data characteristics in the training set and the formation type labels of the sample tunneling data, the model is trained, and based on the new data characteristics in the test set and the formation type labels of the sample tunneling data, the model is evaluated to obtain the supervised formation recognition model.
[0097] In this regard, it should be noted that refer to Figure 3, due to various defects of Gradient Boosting Tree (GBC), K-Nearest Neighbor (KNN), Multi-Layer Perceptron (MLP), Random Forest (RF), and Support Vector Machine (SVC) in boulder stratum identification, in order to make full use of shield data and obtain the best boulder stratum identification results, GBC, KNN, MLP, RF, and SVC are used as base learners, and their prediction results are pieced together into a new feature matrix (for example, if there are 10 feature items and they are classified and predicted by 5 algorithms, the number of items of the new features is fused to 15 items). Then, the new feature matrix and the true labels are input into the supervised RF algorithm for training to obtain the fused model, and then the data is processed through the model to predict the stratum type (different stratum numbers correspond to different stratum types).
[0098] It should also be noted that during the process of training the model with new features, when evaluating the model based on the new data features in the test set and the stratum type labels of the sample tunneling data, when it is determined that the classification evaluation indicators, confusion matrix, and performance indicators all meet the conditions, the supervised stratum identification model is determined.
[0099] The classification evaluation indicators include accuracy (ACC), precision (PRE) i , recall (REC) i , and the harmonic mean F of precision and recall 1i ;
[0100]
[0101] Among them, i is the number of stratum types, TP i is the number of samples determined to be i and actually being i, N i is the number of samples actually being i, and P i is the number of samples determined to be i;
[0102] The confusion matrix is an n×n matrix used to compare the differences between the prediction results and the actual labels. Among them, n is the number of stratum types. In the matrix, the element in the i-th row and j-th column represents the number of samples with the actual type being i and the predicted type being j;
[0103] The performance indicator is characterized by a curve, which represents the trade-off relationship between the true positive rate and the false positive rate at different thresholds.
[0104] In the further method of the above method, it mainly explains the processing process of obtaining the tunneling data of the rock mass, which is specifically as follows:
[0105] Obtain the tunneling data of multiple tunneling sections based on the preset extraction strategy of the tunneling sections.
[0106] Perform anomaly analysis on the tunneling data of each tunneling section based on a preset anomaly recognition strategy to determine the anomaly type of each tunneling section, and configure an anomaly handling strategy for the tunneling data of the corresponding tunneling section based on the anomaly type.
[0107] When it is determined that there is an anomaly handling strategy without an interpolation method in the configured anomaly handling strategy, configure a first processing flow, and execute the first processing flow on the tunneling data based on the anomaly handling strategy without an interpolation method.
[0108] After the first processing flow is executed, perform noise reduction processing on all the tunneling data.
[0109] Configure a second processing flow, and perform interpolation processing on the noise-reduced tunneling data based on the anomaly handling strategy with an interpolation method to obtain the tunneling data of the rock mass.
[0110] In this regard, it should be noted that in the present invention, during the tunneling process of the shield machine, due to reasons such as the removal of muck trucks, worker breaks, replacement of worn cutters, and equipment failures, there are many useless and abnormal non-tunneling state data in the original data. To simplify the original data to the greatest extent, it is necessary to extract the tunneling data in the tunneling state, that is, the collected tunneling data includes the data of multiple tunneling sections. Here, a tunneling section refers to a complete working section from the start to the end of the shield machine without any pause in the middle. Most of the parameter values of the shield machine in the non-tunneling section are weak or even zero, so the mapping relationship established with each parameter is meaningless.
[0111] The tunneling section can reflect the working mode of the shield machine and can generally be divided into four stages: the starting section, the rising section, the stable section, and the falling section. The driver first slowly increases the cutter head speed from zero, and at the same time observes the numerical values of the cutter head torque and the cutter head thrust; when the equipment reaches a relatively ideal tunneling state, keep the parameters of this state unchanged until the tunneling ends, complete the excavation work, and enter the shutdown support or other processes.
[0112] Therefore, when collecting tunneling data, it must be strictly able to distinguish the tunneling section, that is, it is necessary to strictly control the tunneling section and only collect the data in the tunneling state.
[0113] Therefore, obtain the data of multiple tunneling sections based on a preset extraction strategy for tunneling sections.
[0114] The extraction strategy for tunneling sections includes:
[0115]
[0116] Among them, is the first point where the cutter head speed in tunneling section i is greater than zero, is the first point where the cutter head speed in tunneling section i is equal to zero, is the end displacement of the tunneling section i, is the starting displacement of the tunneling section i, l 油缸 is the theoretical elongation of the propulsion cylinder, Time p-min is the minimum time interval between adjacent tunneling sections.
[0117] During tunneling, different factors (such as too hard formation geology, inaccurate acquisition devices, poor transmission signals, etc.) may cause abnormal tunneling data. Therefore, after the required tunneling data is collected, abnormal determination is performed on each tunneling section. At this time, an abnormal recognition strategy is configured, and based on the configured abnormal recognition strategy, feature analysis is performed on the tunneling data of each tunneling section to identify the abnormal type of each tunneling section. Different abnormal types result in different processing methods for the tunneling data of that tunneling section. Therefore, an abnormal processing strategy needs to be configured for the tunneling data of the corresponding tunneling section based on the abnormal type. Table 1 shows the correspondence table among the tunneling section, recognition strategy, and processing strategy.
[0118] Table 1 shows the correspondence among the tunneling section, recognition strategy, and processing strategy
[0119]
[0120] As can be seen from Table 1 above, the processing strategies include deleting this type of tunneling section, deleting and interpolating relevant data, and interpolating relevant data. For these strategies, when the tunneling data of a certain tunneling section is required, the strategy should at least include the interpolation method of the data. When the tunneling data of a certain tunneling section is not required, the interpolation method of the data is not included in the strategy. Therefore, the present invention needs to consider from two processing processes. First, it is necessary to determine whether there is an abnormal processing strategy without an interpolation method in the allocated abnormal processing strategy. If so, the first processing process is configured. In this processing process, the first processing process is executed on the tunneling data based on the abnormal processing strategy without an interpolation method. Correspondingly, when it is determined that there is an abnormal processing strategy with an interpolation method in the allocated abnormal processing strategy, the second processing process is configured. In this processing process, interpolation processing is performed on the tunneling data. Configuring two processing processes is to distinguish and execute sequentially in two application scenarios with and without an interpolation method.
[0121] In the present invention, data denoising refers to reducing or removing noise in data through various methods and technologies during data processing to improve the quality and reliability of the data, so that the data can more truly reflect its inherent laws and characteristics. Therefore, to ensure the quality and reliability of the tunneling data, after the first processing process is completed, all tunneling data is denoised. In fact, after deleting the worthless data, data denoising should be performed first so that in the second processing process, interpolation processing can be performed on the data.
[0122] In the present invention, the tunneling data after noise reduction processing is interpolated to obtain the processing process of the tunneling data of the rock mass. It should be noted that:
[0123] Determine the data missing type corresponding to the tunneling section and the missing rate of each missing area corresponding to the tunneling section;
[0124] Based on the data missing type and the missing rate, determine the interpolation measure corresponding to the missing area.
[0125] Data missing in actual engineering can be divided into three types: interval missing, scattered missing, and continuous time missing. The specific introduction is as follows:
[0126] Interval missing: An entire block of data in the data set is missing. In this case, the data within a certain time period of the data set is completely lost.
[0127] Scattered missing: There are scattered missing at multiple positions in the data set. In this case, the data of multiple data points or multiple time periods in the data set is lost.
[0128] Continuous time missing: The data set is missing multiple times, but each missing only lasts for 1 s.
[0129] For different data missing types, different interpolation measures are configured. The interpolation measures can include linear fitting interpolation, cubic function fitting interpolation, RF fitting interpolation, support vector machine interpolation, K-nearest neighbor interpolation, moving average interpolation, mean fitting interpolation, MLP interpolation, etc.
[0130] In the present invention, when selecting different interpolation measures, it is also necessary to limit based on the missing rate. For this reason, it is necessary to obtain the missing rate of each missing area corresponding to the tunneling section, and then based on the data missing type and the missing rate, determine the interpolation measure corresponding to the missing area.
[0131] The calculation method of the missing rate is as follows:
[0132]
[0133] In the present invention, Tables 2 to 4 are the selection rules for the interpolation measures for interval missing, scattered missing, and continuous time missing.
[0134] Table 2 is the selection rule for the interpolation measure for interval missing
[0135]
[0136] Table 3 is the selection rule for the interpolation measure for scattered missing
[0137]
[0138] Table 4 is the selection rule for the interpolation measure for continuous time missing
[0139]
[0140] The parameters in Tables 2 to 4 are the original penetration degree (p), cutter head torque (T), and total thrust (F).
[0141] The above rules are used for missing value imputation of the actual monitoring data:
[0142] (1) If the actual missing rate < 2%, no repair is performed;
[0143] (2) If 2% ≤ actual missing rate < 3.5%, repair is performed according to a missing rate of 2%;
[0144] (3) If 3.5% ≤ actual missing rate < 7.5%, repair is performed according to a missing rate of 5%;
[0145] (4) If 7.5% ≤ actual missing rate < 15%, repair is performed according to a missing rate of 10%;
[0146] (5) If 15% ≤ actual missing rate, repair is performed according to a missing rate of 20%.
[0147] The following describes the boulder stratum identification device provided by the present invention based on tunneling data. The boulder stratum identification device based on tunneling data described below can be correspondingly referred to the boulder stratum identification method based on tunneling data described above.
[0148] Figure 4 The flow schematic diagram of a boulder stratum identification device provided by the present invention is shown. Refer to Figure 4 , the device includes an acquisition module 41, a first identification module 42, a construction module 43, a second identification module 44, and a processing module 45, where:
[0149] The acquisition module is used to acquire the data characteristics of the tunneling data of the rock mass;
[0150] The first identification module is used to respectively identify the stratum types of the data characteristics based on a variety of pre-configured unsupervised stratum classification models, and obtain the stratum type results corresponding to each stratum classification model;
[0151] The construction module is used to construct new data characteristics based on the data characteristics and the stratum type results;
[0152] The second identification module is used to input the new data characteristics into a supervised stratum classification model to obtain the stratum type of the rock mass output by the supervised stratum classification model;
[0153] The processing module is used to issue a warning message when the stratum type is determined to be a boulder stratum;
[0154] Among them, the supervised formation classification model is obtained by machine learning training with the data features determined based on the sample tunneling data, the formation type results determined based on the data features, and the formation type labels of the sample tunneling data as inputs, and is used to classify the formation of the tunneling data.
[0155] Since the device in the embodiment of the present invention has the same principle as the method in the above embodiment, the more detailed explanation content will not be repeated here.
[0156] It should be noted that in the embodiment of the present invention, the relevant functional modules can be implemented by a hardware processor.
[0157] The boulder formation identification device based on tunneling data provided by the present invention respectively identifies the formation types of the data features of the tunneling data through a variety of pre-configured unsupervised formation classification models, obtains the formation type results corresponding to each formation classification model, then constructs new data features based on the data features and the formation type results, and finally inputs the new data features into the supervised formation classification model to obtain the formation type of the rock mass output by the supervised formation classification model, realizing a richer and more comprehensive feature representation of the tunneling data and accurately identifying the formation type of the rock mass.
[0158] Figure 5 An example of the physical structure diagram of an electronic device is shown as Figure 5 shown. The electronic device may include: a processor 51 (processor), a communication interface 52 (Communications Interface), a memory 53 (memory), and a communication bus 54. Among them, the processor 51, the communication interface 52, and the memory 53 communicate with each other through the communication bus 54. The processor 51 can call the logical instructions in the memory 53 to execute the boulder formation identification method based on tunneling data, and the method includes: obtaining the data features of the tunneling data of the rock mass; respectively identifying the formation types of the data features through a variety of pre-configured unsupervised formation classification models, and obtaining the formation type results corresponding to each formation classification model; constructing new data features based on the data features and the formation type results; inputting the new data features into the supervised formation classification model to obtain the formation type of the rock mass output by the supervised formation classification model; and sending out a warning message when the formation type is determined to be a boulder formation. Among them, the supervised formation classification model is obtained by machine learning training with the data features determined based on the sample tunneling data, the formation type results determined based on the data features, and the formation type labels of the sample tunneling data as inputs, and is used to classify the formation of the tunneling data.
[0159] In addition, when the logical instructions in the above-mentioned memory 53 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0160] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the boulder stratum identification method based on tunneling data provided by the above-mentioned various methods. The method includes: obtaining the data characteristics of the tunneling data of the rock mass; respectively performing stratum type identification on the data characteristics based on a variety of pre-configured unsupervised stratum classification models to obtain stratum type results corresponding to each stratum classification model; constructing new data characteristics based on the data characteristics and the stratum type results; inputting the new data characteristics into a supervised stratum classification model to obtain the stratum type of the rock mass output by the supervised stratum classification model; and when the stratum type is determined to be a boulder stratum, sending out a warning message. Among them, the supervised stratum classification model is a model that is trained through machine learning by using the data characteristics determined based on the sample tunneling data, the stratum type results determined based on the data characteristics, and the stratum type labels of the sample tunneling data as inputs and is used for stratum classification of tunneling data.
[0161] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for identifying boulder formations based on tunneling data provided by the above-mentioned various methods. The method includes: obtaining the data characteristics of the tunneling data of the rock mass; respectively identifying the formation types of the data characteristics based on a variety of pre-configured unsupervised formation classification models to obtain formation type results corresponding to each formation classification model; constructing new data characteristics based on the data characteristics and the formation type results; inputting the new data characteristics into a supervised formation classification model to obtain the formation type of the rock mass output by the supervised formation classification model; and when the formation type is determined to be a boulder formation, sending out a warning message. Among them, the supervised formation classification model is a model obtained by machine learning training with the data characteristics determined based on the sample tunneling data, the formation type results determined based on the data characteristics, and the formation type labels of the sample tunneling data as inputs, and is used for classifying the formation of the tunneling data.
[0162] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0164] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying boulder formations based on tunneling data, characterized in that, Including: Obtaining data characteristics of the tunneling data of the rock mass; the data characteristics are composed of data corresponding to multiple feature items; Performing unsupervised clustering processing on the data characteristics respectively based on multiple pre-configured unsupervised clustering models, and obtaining unsupervised clustering label results corresponding to each feature item and matching the preset number of formation types; Constructing new data characteristics based on the data characteristics and the unsupervised clustering label results; Inputting the new data characteristics into a supervised formation identification model to obtain the formation type of the rock mass output by the supervised formation identification model; When it is determined that the formation type is a boulder formation, sending out a warning message; Wherein, the supervised formation identification model is a model obtained through machine learning training by taking the data characteristics determined based on the sample tunneling data, the clustering label results determined by performing unsupervised clustering on the data characteristics determined based on the sample tunneling data as inputs, and taking the formation type labels corresponding to the sample tunneling data as outputs, and is used for identifying the formation of the tunneling data.
2. The method for identifying boulder stratum based on tunneling data according to claim 1, wherein The obtaining of the data characteristics of the tunneling data of the rock mass includes: Obtaining multiple pre-configured feature items, and constructing the data characteristics of the tunneling data of the rock mass based on the data corresponding to the multiple feature items; The obtaining steps of the multiple pre-configured feature items include: Obtaining sample tunneling data of the rock mass, and the sample tunneling data includes multiple data items; Performing feature extraction on the tunneling data respectively based on multiple pre-configured feature extraction strategies to obtain multiple feature extraction results; the feature extraction results are the sorting results of each data item; Based on the multiple feature extraction results, screening out multiple feature items from multiple data items.
3. The method for identifying boulder stratum based on tunneling data according to claim 2, characterized in that, The screening out of multiple feature items from multiple data items based on the multiple feature extraction results includes: Obtaining the first weight value corresponding to each feature extraction strategy, and obtaining the second weight value of each data item in the feature extraction result corresponding to each feature extraction strategy under the feature extraction strategy; Based on the multiple feature extraction results, and the first weight value and the second weight value corresponding to the feature extraction results, determining a comprehensive feature extraction result; the comprehensive feature extraction result is the sorting result of each data item; Based on the comprehensive feature extraction result and the screening number of feature items, screening out multiple feature items from multiple data items.
4. The method for identifying boulder stratum based on tunneling data according to claim 1, characterized in that The obtaining steps of the supervised formation identification model include: Obtaining sample tunneling data of the rock mass, and the sample tunneling data includes multiple data items; Performing feature extraction on the tunneling data respectively based on multiple pre-configured feature extraction strategies to obtain multiple feature extraction results; the feature extraction results are the sorting results of each data item; Based on the multiple feature extraction results, screening out multiple feature items from multiple data items; Obtaining the data characteristics of the sample tunneling data of the rock mass based on the screened multiple feature items; Performing unsupervised clustering processing on the data characteristics respectively based on multiple pre-configured unsupervised clustering models, and obtaining unsupervised clustering label results corresponding to each feature item and matching the preset number of formation types; Construct new data features of the sample tunneling data based on the data features and the unsupervised clustering label results; Divide into a training set and a test set based on the new data features of the sample tunneling data; Train a model based on the new data features in the training set and the formation type labels of the sample tunneling data, and evaluate the model based on the new data features in the test set and the formation type labels of the sample tunneling data to obtain a supervised formation identification model.
5. The method for identifying boulder stratum based on tunneling data according to claim 1, wherein, During the process of evaluating the model based on the new data features in the test set and the formation type labels of the sample tunneling data, when it is determined that the classification evaluation index, confusion matrix, and performance index all meet the conditions, determine the supervised formation identification model.
6. The method for identifying boulder stratum based on tunneling data according to claim 1, characterized in that, Before obtaining the data features of the tunneling data of the rock mass, the method further includes obtaining the tunneling data of the rock mass, including: Obtain the tunneling data of multiple tunneling sections based on a preset extraction strategy for the tunneling sections; Perform anomaly analysis on the tunneling data of each tunneling section based on a preset anomaly identification strategy, determine the anomaly type of each tunneling section, and configure an anomaly handling strategy for the tunneling data of the corresponding tunneling section based on the anomaly type; When it is determined that there is an anomaly handling strategy without an interpolation method in the configured anomaly handling strategy, configure a first processing flow, and execute the first processing flow on the tunneling data based on the anomaly handling strategy without an interpolation method; After the first processing flow is executed, perform noise reduction processing on all the tunneling data; Configure a second processing flow, and perform interpolation processing on the noise-reduced tunneling data based on the anomaly handling strategy with an interpolation method to obtain the tunneling data of the rock mass.
7. The method for identifying boulder stratum based on tunneling data according to claim 6, wherein Perform interpolation processing on the noise-reduced tunneling data to obtain the tunneling data of the rock mass, including: Determine the data missing type corresponding to the tunneling section and the missing rate of each missing area corresponding to the tunneling section; Based on the data missing type and the missing rate, determine the interpolation measure corresponding to the missing area.
8. An apparatus for identifying boulder formations based on tunneling data, characterized in that, Include: An acquisition module for acquiring the data features of the tunneling data of the rock mass; the data features are composed of data corresponding to multiple feature items; A first identification module for performing unsupervised clustering processing on the data features respectively based on a plurality of pre-configured unsupervised clustering models to obtain an unsupervised clustering label result corresponding to each feature item and matching the preset number of formation types; A construction module for constructing new data features based on the data features and the unsupervised clustering label results; A second identification module for inputting the new data features into a supervised formation identification model to obtain the formation type of the rock mass output by the supervised formation identification model; A processing module for sending out a warning message when it is determined that the formation type is a boulder formation; Wherein, the supervised formation identification model is a model obtained by machine learning training with the data features determined based on the sample tunneling data and the clustering label results determined by unsupervised clustering of the data features determined for the sample tunneling data as inputs, and the formation type labels corresponding to the sample tunneling data as outputs, and is used for identifying the formation of the tunneling data.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the boulder stratum identification method based on tunneling data as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the boulder stratum identification method based on tunneling data as described in any one of claims 1-7.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN111666275A
Geological type identification method and device, storage medium and computer equipment
CN112364917A
Anomaly detection and attack initiator analysis system based on network flow message
CN112422513A
Model-based risk identification method and device, computer equipment and storage medium
CN113327037A
Semi-supervised shield tunnel face geological type prediction method and system
CN113431635A