Shield muck resource utilization classification method based on XGBoost

By constructing a slag classification method through the XGBoost algorithm, the problem of insufficient capture of slag data changes during shield construction was solved, and the accurate classification and efficient reuse of slag resources were achieved.

CN120724221AInactive Publication Date: 2025-09-30SICHUAN HIGHWAY PLANNING SURVEY DESIGN AND RESEARCH INSTITUTE LTD +3
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511212674.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-09-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies fail to effectively capture the high-frequency changes in slag data during shield construction, resulting in delayed classification labels, affecting reuse effects, increasing the difficulty of soil remediation and wasting resources.

Method used

The XGBoost algorithm is used to construct a combination of partition nodes and attribute offsets based on the difference number of slag particle size, particle friction curvature, dry bulk density and density change of ring cutter samples, combined with the discrete ratio of propulsion speed and cutter head torque. The disturbance intensity mapping and classification reference table are generated, the amplification points and label conflicts are identified, the label segment constraint parameter matrix is ​​established, and the use path matching and process block sequence are generated.

Benefits of technology

Accurately capture the physical property variation points during the process of slag advancement, enhance the anti-interference ability of the classification mechanism, improve the classification robustness and scene adaptability, reduce the decline in classification accuracy, and improve the accuracy of slag resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724221A_ABST
    Figure CN120724221A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field, in particular to an XGBoost-based shield muck resource utilization classification method, which comprises the following steps of: constructing an attribute offset combination by taking a muck particle size difference number, a particle friction curvature, a dry volume weight and a compactness change value as basic indexes and a linkage propulsion speed and a cutterhead torque discrete ratio; according to the method, key physical property variation points in the propulsion process are accurately captured, so that the disturbance trend is clearly presented in the time sequence dimension, and the cooperative capability of the particle size extreme value and the propulsion accumulated value in disturbance recognition is effectively enhanced. The method is advantaged in that improvement of perception capability of boundary unstable areas is facilitated, a process block sequence is constructed through combination of parameter fluctuation sections and path numbers, anti-interference capability of a classification mechanism is enhanced, and stronger scene adaptability and classification robustness are shown in a muck complex fluctuation environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of resource utilization classification, and in particular to an XGBoost-based shield slag resource utilization classification method. Background Art

[0002] Resource utilization classification is a technology that classifies, identifies and determines the utilization of solid waste generated during engineering construction, industrial production and urban development, and determines its reuse value in building material recycling, earth backfill and road base. The large amount of slag generated during shield construction has complex physical properties and diverse composition. If it is not classified and treated, it will be difficult to achieve efficient utilization, which can easily cause resource waste and environmental burden. The resource utilization classification technology is committed to building a set of identification mechanisms with strong operability, clear evaluation standards and high intelligence level to support the implementation of the solid waste resource recycling system.

[0003] Resource utilization classification refers to the systematic classification and processing of engineering waste based on multiple parameters such as physical properties, mechanical indicators, chemical composition and moisture content. The purpose is to identify available attributes, determine the most suitable reuse scenario, and guide subsequent processing and direct utilization based on the classification results. The goal of classification is to improve the reuse rate of engineering waste, reduce raw material consumption, reduce landfill space, and achieve green transformation of solid waste disposal and closed-loop management of resources throughout the entire process.

[0004] Traditional methods rely on static indicators such as physical properties, mechanical indicators, moisture content and chemical composition as input, ignoring the dynamic changes of data during the construction process. Traditional methods do not introduce disturbance trends, label conflicts and path matching factors, resulting in the recognition mechanism remaining at feature points and static numerical dimensions. It is difficult to capture the high-frequency change intervals during construction disturbances, including data drift and periodic anomalies generated during the project advancement. It is impossible to respond and adjust in time, resulting in classification label lags, which is manifested as a decrease in the accuracy of the reuse plan, affecting the reuse effect, increasing the difficulty of subsequent soil remediation treatment, causing project delays and resource waste, and forming a substantial obstacle to the goal of green construction. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings of the prior art and propose a shield slag resource utilization classification method based on XGBoost.

[0006] In order to achieve the above objectives, the present invention adopts the following technical solution: a shield slag resource utilization classification method based on XGBoost, comprising the following steps: Step 1: Based on the four original data items of soil particle size difference, particle friction curvature, dry bulk density of ring cutter specimens, and density change, the propulsion speed sequence and cutterhead torque discrete ratio during construction are matched to establish a partition node and attribute offset combination, and a list of partition node and attribute offset combinations is marked; Step 2: Based on the partition node and attribute offset combination list, extract the mud and sand offset direction, generate a corresponding symbol sequence, count the intersection frequency areas, extract the particle size threshold points and the cumulative value of the advance to form a ratio comparison sequence, and obtain the disturbance intensity map and classification reference table; Step 3: Based on the disturbance intensity map and classification reference table, extract the offset values ​​of the tunneling parameters and physical property combination items, identify the amplification points and set the node intervals, sort and filter the label conflict numbers in the split nodes, and generate the split judgment path and boundary stability classification labels; Step 4: Based on the splitting judgment path and the boundary stability classification label, extract the number segment label, read the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset value, identify the extreme values ​​of the three items, establish the index number and the change point mapping, and obtain the label segment constraint parameter matrix; Step 5: Based on the label segment constraint parameter matrix, read the earth pressure boundary coefficient and cyclic offset value by block, filter the floating parameter segment, select the intersection that meets the preset path number to bind the path number and process segment combination, and generate the use path matching and process block sequence.

[0007] As a further solution of the present invention, the specific steps of generating the partition node and attribute offset combination list are: Based on the four original data items of soil particle size difference, particle friction curvature, dry bulk density of ring cutter specimens, and density change, the advancement speed sequence and cutterhead torque discrete ratio during construction were matched, aligned along the time axis, and a sample number sequence was established. The data were then paired in chronological order, and the order of the data was arranged to obtain the sample number sequence. Based on the aligned sample number sequence, the difference combination of dry density and mud content parameters is extracted, and the interval division is performed based on the particle size difference number and density change value. The boundary of the sample section is determined according to the extreme value of the data, and the interval range is marked to obtain the sample section boundary data; Based on the sample paragraph boundary data, a partition node and attribute offset combination list is generated, the nodes are divided according to the segment boundaries, the attribute value intervals are marked, the parameter difference direction is calculated, and the node information matching the sample paragraph boundary data is generated to obtain a partition node and attribute offset combination list.

[0008] As a further solution of the present invention, the specific steps of generating the disturbance intensity map and classification reference table are: Based on the partition node and attribute offset combination list, extract the offset direction of each section of mud and sand, calculate the offset direction and attribute change in sequence, mark the offset direction, generate a corresponding symbol sequence, and correspond the symbols to the data segments to obtain a symbol sequence; Based on the intersection frequency area of ​​the symbol sequence, traverse the symbols and record the number of occurrences of the intersection points, calculate the distribution of the frequency area, extract the particle size extreme value and the propulsion cumulative value to form a ratio comparison sequence, sort by the particle size extreme value and the propulsion value ratio, screen out the disturbance reference points, organize the numbered sequence structure, and obtain the disturbance reference point sequence; Based on the disturbance reference point sequence, a disturbance intensity mapping and a classification reference table are constructed, the disturbance reference ranking indexes of the disturbance points are summarized, the correlation value between the sample number and the disturbance intensity is calculated, the disturbance intensity information corresponding to the sample number is generated, and the disturbance intensity mapping and the classification reference table are obtained.

[0009] As a further solution of the present invention, the specific steps of generating the split determination path and the boundary stability classification label are: Based on the disturbance intensity mapping and classification reference table, a judgment path is constructed and offset values ​​of tunneling parameters and physical property combination items are extracted. After traversing the nodes according to the sample number, the change values ​​related to the tunneling parameters are extracted, the offset of the physical property items is calculated, the amplification points are identified, and the node interval is determined according to the amplification amount to obtain the node interval and amplification point data; Based on the node interval and amplification point data, the leaf node category count density is sorted, the count density of the category in each node is counted in turn, and sorted by density. The numbers of the split nodes with label conflicts are screened, the conflict ratio is extracted, and the difference values ​​between the nodes are calculated to obtain the conflict labels and difference ratio data; Based on the conflicting labels and difference ratio data, principal component analysis is used to match stable path nodes, sort them according to the matching of labels and paths, filter out nodes that conflict with labels one by one, generate split judgment paths and category labels corresponding to nodes, and obtain split judgment paths and boundary stable classification labels.

[0010] As a further aspect of the present invention, the principal component analysis is generated according to the formula: ; in: For the The sample in The weighted projection values ​​in the direction of the principal components, For the summation symbol, is the feature dimension index, is the total number of features, is the sample index, is the principal component index, For the The sample in The value of the original feature dimension, For the The mean of the original features in all samples, For the The original features are The load values ​​in the principal component directions are For the The difference weight coefficient corresponding to the feature dimension, For the The path stability influence coefficient corresponding to each characteristic dimension.

[0011] The difference value between nodes is used to measure the degree of difference in label distribution in different nodes. It is defined based on a clear statistical distance and distribution similarity measurement method. To enhance the clarity of expression, the difference value should be specifically set as the distance measurement result of the label distribution vector between nodes. The node can be represented as a label frequency vector, the dimension corresponds to different category labels, and the element value is the frequency of occurrence in the node. The difference value between two nodes can be calculated by the inverse function of cosine similarity and Euclidean distance. This method can quantitatively describe the degree of deviation of the category structure between adjacent nodes, thereby assisting in identifying label mutation trends, boundary volatility and potential misclassification areas in actual operations, and enhancing the reliability of the basis for splitting paths and boundary judgments.

[0012] As a further solution of the present invention, the specific steps of generating the label segment constraint parameter matrix are: Based on the splitting judgment path and the boundary stability classification label, the number segment label is extracted, the number segment is traversed and the corresponding earth pressure release coefficient and propulsion fluctuation ratio, and the slag discharge synchronization offset value are read, and then the row data is cleaned and standardized to obtain the number segment label data; Based on the numbered segment label data, the change points in the three parameters are identified, the change amplitudes of the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset value are calculated, the extreme value points of the data are determined, and a mapping relationship between the index number and the change point is established to obtain the change point mapping data; Based on the change point mapping data, a segment output matrix is ​​constructed, the numbered segments are associated with the change points, and the variation ranges of the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset value are compared. The constraint parameter range of each numbered segment is calculated to obtain the label segment constraint parameter matrix; As a further solution of the present invention, the specific steps of generating the usage path matching and process block sequence are: Based on the label segment constraint parameter matrix, the algorithm obtains the earth pressure boundary coefficient and cyclic offset values ​​according to the block, sequentially records and calculates the value changes, and screens out the sections where the parameter fluctuation is greater than the target range to obtain the floating parameter section data; Based on the floating parameter segment data, counting the items that overlap with the path condition set, traversing the floating parameter segment, comparing the intersection items with the path condition set, calculating the number of overlapping items, selecting the path numbers whose intersection number meets the target number, and obtaining the extreme value of the overlapping path number data; Based on the extreme values ​​of the overlapping path number data, a dynamic time warping algorithm is used to bind the path number and the process segment combination, and matching is performed according to the order of the path number and the process segment, so that the path and the process correspond to each other, and the purpose path matching and process block sequence are obtained; As a further solution of the present invention, the dynamic time warping algorithm is according to the formula: ; in: For the The path number and The cumulative alignment distance value of each process section, is the index position in the path number sequence, is the index position in the process segment sequence, Number the path Item and process section Matching strength correction factor between items, The first The multidimensional feature vector corresponding to the number, The first The multidimensional feature vector corresponding to each process, is the path feature vector Item and process characteristic vector The Euclidean distance between items, is the power exponential correction factor of the Euclidean distance, From the path number Item is aligned to the process section The cumulative distance value of the item, From the path number Item is aligned to the process section The cumulative distance value of the item, From the path number Item is aligned to the process section The cumulative distance value of the item, Computes the sign for the Euclidean distance between vectors, It is the coordinate position in the two-dimensional path matching table.

[0013] The establishment of the combination of path number and process segment should be based on the overlapping relationship of time series and the process continuity of construction events. The path number represents the data sampling sequence during the shield advancement process. The number corresponds to the advancement characteristic record within a specific period of time, including advancement speed, cutter head torque and slag discharge volume. The process segment corresponds to the operation stage division defined in the construction log, which includes the specific process type, start and end time, and the corresponding operation parameter range. For effective binding, the advancement timestamp should be used as the alignment benchmark to determine whether the sample time in the path number sequence completely falls within the time range of the process segment. When the time inclusion relationship is met and the path number sequence maintains the continuity and stability of the advancement parameters in the process segment, it is determined that the path number and process segment form a combination relationship. The coherent adaptation combination relationship provides an executable basis for the subsequent use of path and process logic matching.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are: 1. This method uses the soil particle size difference, particle friction curvature, dry bulk density, and density variation as basic indicators, and links propulsion speed with the cutterhead torque discrete ratio to construct an attribute offset combination. This accurately captures key physical property variation points during the propulsion process, forming a list of partition nodes and offset relationships, thus avoiding misjudging complex soil conditions with a single indicator. 2. In this invention, the deviation direction is expressed through a symbol sequence and the intersection frequency region is extracted, so that the disturbance trend is clearly presented in the time series dimension, and the synergy between the particle size extreme value and the propulsion cumulative value in disturbance identification is effectively enhanced; 3. In this invention, by constructing a mapping table for disturbance intensity and extracting offset amplification, label conflict, and physical property combination items, the splitting path and classification label are accurately located, which helps improve the ability to perceive unstable boundary areas. Further, by reading the earth pressure release coefficient, propulsion fluctuation ratio, and slag discharge offset value; 4. In the present invention, a process block sequence is constructed by combining parameter fluctuation segments with path numbers, so that the output results have the ability to support practical uses. This processing logic makes the recognition process no longer rely on single-point numerical judgment, but is instead controlled by the three-level joint control of time series, structure, and category. This enhances the anti-interference ability of the classification mechanism, reduces the decline in classification accuracy caused by parameter disturbances, and demonstrates stronger scene adaptability and classification robustness in the complex and fluctuating environment of slag. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, constitute a part of this application, and do not constitute a limitation of the embodiments of the present invention. In the drawings: Figure 1 It is a structural schematic diagram of the present invention; DETAILED DESCRIPTION

[0016] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the examples and accompanying drawings. The exemplary embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention. It should be noted that the present invention is already in the actual development and use stage.

[0017] Example 1 See also Figure 1 The present invention provides a technical solution: a shield slag resource utilization classification method based on XGBoost, comprising the following steps: Step 1: Based on the four original data items of soil particle size difference, particle friction curvature, dry bulk density of ring cutter specimens, and density change, the propulsion speed sequence and cutterhead torque discrete ratio during construction are matched to establish a partition node and attribute offset combination, and a list of partition node and attribute offset combinations is marked; Step 2: Based on the partition node and attribute offset combination list, extract the direction of mud and sand offset, generate the corresponding symbol sequence, count the intersection frequency areas, extract the particle size extreme points and the cumulative value of the advance to form a ratio comparison sequence, and obtain the disturbance intensity map and classification reference table; Step 3: Based on the disturbance intensity map and classification reference table, extract the offset values ​​of the tunneling parameters and physical property combination items, identify the amplification points and set the node intervals, sort and filter the label conflict numbers in the split nodes, and generate the split judgment path and boundary stability classification labels; Step 4: Based on the splitting judgment path and boundary stability classification labels, extract the numbered segment labels, read the earth pressure release coefficient, propulsion fluctuation ratio, and slag discharge synchronization offset value, identify the extreme values ​​of the three items, establish a mapping between index numbers and change points, and obtain the label segment constraint parameter matrix; Step 5: Based on the label segment constraint parameter matrix, read the earth pressure boundary coefficient and cyclic offset value by block, filter the floating parameter segment, select the intersection that meets the preset path number to bind the path number and process segment combination, and generate the use path matching and process block sequence.

[0018] The specific steps to generate a list of partition nodes and attribute offset combinations are: Based on the four original data items of soil particle size difference, particle friction curvature, dry bulk density of ring cutter specimens, and density change, the advancement speed sequence and cutterhead torque discrete ratio during construction were matched, aligned along the time axis, and a sample number sequence was established. The data were then paired according to chronological order, and the order of the sample data was arranged to obtain a sample number sequence. Based on the aligned sample number sequence, the difference combination of dry density and mud content parameters is extracted, and the interval division is performed based on the particle size difference number and density change value. The boundary of the sample section is determined according to the extreme value of the data, and the interval range is marked to obtain the sample section boundary data; Based on the sample paragraph boundary data, a partition node and attribute offset combination list is generated, nodes are divided according to the segment boundary, attribute value intervals are marked, the parameter difference direction is calculated, and node information matching the sample paragraph boundary data is generated to obtain a partition node and attribute offset combination list; Based on the four original data items of particle size difference, particle friction curvature, dry bulk density of ring cutter specimens, and density change, a time series alignment algorithm was used to align the propulsion speed sequence and the discrete ratio of the cutter head torque during construction along the time axis. The data points were matched with the timestamps to align the order of the construction data. The merge function in the pandas library was used to merge the four original data items with the propulsion speed and cutter head torque sequences according to the timestamps. The order of the data columns was arranged to generate a sample number sequence. The samples were assigned numbers in the merged data frame and saved as a sample number sequence. Based on the aligned sample number sequence, a differential calculation algorithm is used to calculate the difference between the dry density and mud content parameters. The diff() function in numpy is used to perform differential calculation on the dry density and mud content data in the sample to obtain the sample difference sequence. Combined with the particle size difference number and the density change value, an interval partitioning algorithm is used. The find_peaks() function in the scipy library is used to extract the extreme values ​​in the data. The extreme values ​​of the data are used to determine the interval range, and the boundary threshold is set to mark the boundary of the sample paragraph. The boundary data of the sample paragraph is sorted into a sequence according to the interval range to generate the sample paragraph boundary data. Based on the sample paragraph boundary data, a partitioning algorithm is adopted to divide the nodes according to the boundary of the segment. The linspace() function in numpy is used to generate equally spaced nodes in the segment, mark the attribute value interval, calculate the attribute value difference of the node, and use the shift() function in pandas to calculate the attribute difference direction of adjacent nodes, generate the node offset value, and combine the interval information extracted from the sample paragraph boundary data to generate node information that matches the segment boundary. Then, the nodes and attribute offsets are combined in the order of the segment and organized into a data table to generate a list of partition node and attribute offset combinations.

[0019] The specific steps to generate the disturbance intensity map and classification reference table are: Based on the partition node and attribute offset combination list, the offset direction of each segment of mud and sand is extracted, the offset direction and attribute change are calculated in sequence, the offset direction is marked, and the corresponding symbol sequence is generated. The corresponding symbols and data segments are obtained to obtain the symbol sequence; Based on the intersection frequency area of ​​the symbol sequence, traverse the symbols and record the number of occurrences of the intersection points, calculate the distribution of the frequency area, extract the particle size extreme value and the cumulative value of the propulsion ratio comparison sequence, sort by the particle size extreme value and the propulsion ratio, screen out the disturbance reference points, organize the numbered sequence structure, and obtain the disturbance reference point sequence; Based on the disturbance reference point sequence, a disturbance intensity mapping and classification reference table are constructed, the disturbance reference ranking index of the disturbance point is summarized, the correlation value between the sample number and the disturbance intensity is calculated, the disturbance intensity information corresponding to the sample number is generated, and the disturbance intensity mapping and classification reference table are obtained; Based on the list of partition nodes and attribute offset combinations, a direction calculation algorithm is used. The gradient() function in the numpy library is used to calculate the offset direction of the mud and sand to obtain the direction change value. The shift() function in pandas is then used to calculate the attribute change between each data point and the previous data segment, generating a sequence of attribute changes. The if-else statement is used to mark each offset direction as forward or reverse, generating a corresponding symbol sequence. Finally, the concat() function in pandas is used to sequentially merge the offset direction and the attribute change sequence, corresponding symbols to data segments, and generating a symbol sequence. Based on the symbol sequence, a frequency calculation algorithm is adopted. The unique() function in the numpy library is used to traverse the symbols and record the number of times the symbols appear in the sequence. The value_counts() function in pandas is used to calculate the frequency of the intersection points, and a frequency data table of the intersection points is generated. Then, according to the distribution of the frequency areas, the particle size extreme values ​​and the cumulative values ​​of advancement are extracted. The max() function in numpy is used to extract the particle size extreme values, and the cumsum() function in pandas is used to calculate the cumulative values ​​of advancement to form a ratio comparison sequence. Then, the argsort() function in numpy is used to sort the comparison values, and the perturbation reference points with a ratio of the set value are selected. Finally, the sort_values() function in pandas is used to organize the numbering structure of the perturbation reference points to generate a perturbation reference point sequence. Based on the disturbance reference point sequence, a disturbance intensity mapping algorithm is adopted. The std() function in numpy is used to calculate the standard deviation of the disturbance reference point to generate a preliminary mapping of the disturbance intensity. The disturbance points are then grouped according to the category of the disturbance reference point through the groupby() function in pandas. The disturbance reference sorting index of the disturbance point is summarized, and the mean() function in numpy is used to calculate the average disturbance intensity value of the group. The merge() function in pandas is used to associate the disturbance intensity of the disturbance point with the sample number to generate the disturbance intensity information corresponding to the sample number, which is organized into a disturbance intensity mapping and classification reference table.

[0020] The specific steps for generating split decision paths and boundary stability classification labels are: Based on the disturbance intensity mapping and classification reference table, a judgment path is constructed and the offset values ​​of the tunneling parameters and physical property combination items are extracted. After traversing the nodes according to the sample number, the change values ​​related to the tunneling parameters are extracted, the physical property item offset is calculated, the amplification point is identified, and the node interval is determined based on the amplification value, thus obtaining the node interval and amplification point data; Based on the node interval and increase point data, the leaf node category count density is sorted, the category count density of each node is counted in turn, and the count density is sorted by density. The numbers of the split nodes with label conflicts are screened, the conflict ratio is extracted, and the difference values ​​between the nodes are calculated to obtain the conflict label and difference ratio data; Based on the conflicting labels and difference ratio data, principal component analysis is used to match stable path nodes, sort them according to the matching of labels and paths, filter out nodes that conflict with labels one by one, generate split decision paths and category labels corresponding to the nodes, and obtain split decision paths and boundary stable classification labels; Based on the disturbance intensity mapping and classification reference table, a path determination algorithm is adopted. The iterrows() function in pandas is used to traverse the nodes by sample number, extract the change values ​​related to the tunneling parameters, and use the diff() function in numpy to calculate the change in the tunneling parameters. Based on the calculation results, the offset value of the property combination item is extracted. The offset of the physical property item is calculated using the gradient() function in numpy. Then, the amplification point is identified using the where() function in numpy. The extreme value position of the amplification point is determined based on the amplification amount using the argmax() function in numpy. The node interval is determined. Finally, the amplification point and the node interval are combined and organized into a data table to generate the node interval and amplification point data. Based on the node interval and increase point data, a density calculation and conflict analysis algorithm is used. The groupby() function in Pandas is used to group the leaf node categories, and the size() function is used to count the count density of the categories within the node. The argsort() function in NumPy is then used to sort the nodes by density, screening out the numbers with label conflicts in the split nodes. The apply() function in Pandas is used to calculate the conflict ratio and the number of label conflicts in the nodes. The mean() function in NumPy is used to calculate the difference between the nodes. The label conflicts and difference ratios are organized into a data table to generate conflict label and difference ratio data. Based on the conflicting labels and difference ratio data, principal component analysis (PCA) was used. The fit_transform() function in sklearn.decomposition.PCA was used to reduce the dimensionality of the label and path matching, extract the principal components, calculate the principal component contribution of the nodes, and use the argsort() function in numpy to sort the nodes by contribution. The nodes that conflict with the labels were screened out, and the merge() function in pandas was used to merge the label and path matching to generate a split decision path. Then, the category label corresponding to the node was generated according to the category label of the node, and the split decision path and the boundary stable classification label were organized.

[0021] Generate principal component analysis according to the formula: ; in: For the The sample in The weighted projection values ​​in the direction of the principal components, For the summation symbol, is the feature dimension index, is the total number of features, is the sample index, is the principal component index, For the The sample in The value of the original feature dimension is The mean of the original features in all samples, For the The original features are The load values ​​in the principal component directions are For the The difference weight coefficient corresponding to the feature dimension, For the The path stability influence coefficient corresponding to each characteristic dimension; Execution process: First, the samples in the input sample set The corresponding original feature vector is analyzed to extract multidimensional indicators including particle size, moisture, density and mud content, which are recorded as , calculate the feature dimension The sample mean of , perform data centering to eliminate the dimension effect, and obtain the eigenvector based on the covariance matrix of the entire sample , define the principal component direction, and introduce the difference weight coefficient when performing principal component analysis , calculate the ratio of the variance of the current feature in the conflict label subset to the variance of all samples to measure the sensitivity of the feature to label splitting, combined with the path stability influence coefficient , determined based on the label information gain contribution ratio generated by the path splitting feature, the intensity of the effect of the measurement feature on the path stability boundary, and the calculation of the projection result of the sample in the direction of the principal component ,Depend on Subtract the mean Then multiply by the load value , multiplied by the corresponding and Then add up and get all Arrange in descending order, filter the principal component direction most relevant to the conflicting label, generate split path information for node judgment, mark the samples with corresponding resource category labels, and then perform high-precision training and classification prediction on the XGBoost model.

[0022] The difference value between nodes is used to measure the degree of difference in label distribution in different nodes. It is defined based on a clear statistical distance and distribution similarity measurement method. To enhance the clarity of expression, the difference value should be specifically set as the distance measurement result of the label distribution vector between nodes. The node can be represented as a label frequency vector, the dimension corresponds to different category labels, and the element value is the frequency of occurrence in the node. The difference value between two nodes can be calculated by the inverse function of cosine similarity and Euclidean distance. This method can quantitatively describe the degree of deviation of the category structure between adjacent nodes, thereby assisting in identifying label mutation trends, boundary volatility and potential misclassification areas in actual operations, and enhancing the reliability of the basis for splitting paths and boundary judgments.

[0023] The specific steps to generate the label segment constraint parameter matrix are: Based on the splitting judgment path and boundary stability classification label, the number segment label is extracted, the number segment is traversed and the corresponding earth pressure release coefficient and propulsion fluctuation ratio, slag discharge synchronization offset value are read, and the row data is cleaned and standardized to obtain the number segment label data; Based on the numbered segment label data, the change points in the three parameters are identified, the change amplitudes of the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset value are calculated, the extreme value points of the data are determined, and the mapping relationship between the index number and the change point is established to obtain the change point mapping data; Based on the change point mapping data, a segment output matrix is ​​constructed, and the numbered segments are associated with the change points. The variation ranges of the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset are compared. The constraint parameter range of each numbered segment is calculated to obtain the label segment constraint parameter matrix. Based on the splitting judgment path and boundary stability classification labels, a data cleaning and standardization algorithm was used. The apply() function in pandas was used to traverse the numbered segments and extract the corresponding earth pressure release coefficient, propulsion fluctuation ratio, and slag discharge synchronization offset value. The fit_transform() function in sklearn.preprocessing.StandardScaler was used for standardization, adjusting the data mean to 0 and standard deviation to 1. The dropna() function in pandas was used to clean up the missing values ​​in the data, organize it into a standardized data table, and generate numbered segment label data. Based on the numbered segment label data, a change point detection algorithm was used. The diff() function in numpy was used to perform differential calculations on the earth pressure release coefficient, propulsion fluctuation ratio, and slag discharge synchronization offset value to obtain the parameter change amplitude. The scipy.signal.find_peaks() function was then used to identify the extreme point locations from the differential results. A threshold parameter was set to filter out important extreme points. The merge() function in pandas was used to map the extreme point locations to the index numbers and organize them into change point mapping data. Based on the change point mapping data, a segment output construction algorithm is adopted. The vstack() function in numpy is used to connect the numbered segments with the change point mapping, and the two data sets are merged into a new matrix. The variation range of the earth pressure release coefficient, propulsion fluctuation ratio, and slag discharge synchronization offset value of the numbered segments is calculated using the min() and max() functions in numpy to obtain the constraint parameter range of the numbered segments. The constraint range is associated with the label, and the concat() function in pandas is used to combine the constraint parameter range of the numbered segments and the corresponding labels to generate the constraint parameter matrix of the label segment.

[0024] The specific steps for generating usage path matching and process block sequence are as follows: Based on the label segment constraint parameter matrix, the algorithm obtains the earth pressure boundary coefficient and cyclic offset values ​​according to the block, sequentially records and calculates the value changes, and screens out the sections where the parameter fluctuation is greater than the target range to obtain the floating parameter section data; Based on the floating parameter segment data, count the items that overlap with the path condition set, traverse the floating parameter segment, compare the intersection items with the path condition set, calculate the number of overlapping items, select the path number whose intersection number meets the target number, and obtain the extreme value of the overlapping path number data; Based on the extreme values ​​of the overlapping path number data, the dynamic time warping algorithm is used to bind the path number and process segment combination, and the path and process are matched according to the order of the path number and process segment to make the path and process correspond to each other, thus obtaining the purpose path matching and process block sequence; Based on the label segment constraint parameter matrix, a block acquisition algorithm is adopted. The iloc[] function in Pandas is used to split the matrix into blocks. The earth pressure boundary coefficient and cyclic offset values ​​in the blocks are extracted. The diff() function in NumPy is used to calculate the numerical changes of the parameters in the blocks. The where() function in NumPy is used to filter out the segments with fluctuations greater than the target range. A fluctuation range threshold is set. The screening conditions meet the set target range. The filtered floating parameter segment data is sorted and saved to generate floating parameter segment data. Based on floating parameter segment data, an intersection calculation and path selection algorithm is adopted. The merge() function in Pandas is used to traverse the floating parameter segment and the path condition set, and the intersection items of the two are calculated. The intersect1d() function in NumPy is used to perform the intersection calculation to obtain the overlapping items of the segment and the path condition set. The count() function in Pandas is then used to calculate the number of intersection items of each path. The path numbers whose intersection number meets the requirement are selected according to the target number. The max() function in NumPy is used to calculate the extreme value of the selected path number data, and the extreme value of the overlapping path number data is generated. Based on the extreme values ​​of overlapping path number data, the dynamic time warping (DTW) algorithm is adopted. The fastdtw() function in the fastdtw library is used to perform dynamic time warping of path numbers and process segments. The sequential data of path numbers and process segments are obtained through the iloc[] function in pandas. The fastdtw() function is used to match the time series of path numbers and process segments. The matching is performed according to the order of path numbers and process segments. The matching results are sorted by the argsort() function in numpy to make the paths and processes correspond to each other, and generate usage path matching and process block sequences.

[0025] Dynamic Time Warping algorithm, according to the formula: ; in: For the The path number and The cumulative alignment distance value of each process section, is the index position in the path number sequence, is the index position in the process segment sequence, Number the path Item and process section Matching strength correction factor between items, The first The multidimensional feature vector corresponding to the number, The first The multidimensional feature vector corresponding to each process, is the path feature vector Item and process characteristic vector The Euclidean distance between items, is the power exponential correction factor of the Euclidean distance, From the path number Item is aligned to the process section The cumulative distance value of the item, From the path number Item is aligned to the process section The cumulative distance value of the item, From the path number Item is aligned to the process section The cumulative distance value of the item, Computes the sign for the Euclidean distance between vectors, It is the coordinate position in the two-dimensional path matching table.

[0026] Execution process: Number the path sequences with the coincident path number extreme value characteristics and construct the path number feature vector , the vector contains the disturbance frequency of the path, the propulsion stability coefficient, the relative position of the path and the matching entropy index, and encodes the process segment into the corresponding process feature vector , including construction stage number, process duration, pressure fluctuation degree and stage label index, defining distance index parameters Used to control the penalty degree of the distance between the path and the process, and introduce the matching strength correction factor Adjust the structural relationship between the path number and the process segment. The overlap of the path number, the label disturbance rate and the stability parameter of the process segment jointly determine the adjustment relationship. In the process of aligning the path and the process, recursively calculate each group The value of The path number and The cumulative matching distance of each process section is compared with the previous position when updating. 、 、 The accumulated value of the path cost decision is used to obtain the alignment path, and the constructed usage path and process block sequence are composed of all The backtracking path is determined to achieve accurate binding and classification-assisted input between the path number and the resource utilization category process.

[0027] The establishment of the combination of path number and process segment should be based on the overlapping relationship of time series and the process continuity of construction events. The path number represents the data sampling sequence during the shield advancement process. The number corresponds to the advancement characteristic record within a specific period of time, including advancement speed, cutter head torque and slag discharge volume. The process segment corresponds to the operation stage division defined in the construction log, which includes the specific process type, start and end time, and the corresponding operation parameter range. For effective binding, the advancement timestamp should be used as the alignment benchmark to determine whether the sample time in the path number sequence completely falls within the time range of the process segment. When the time inclusion relationship is met and the path number sequence maintains the continuity and stability of the advancement parameters in the process segment, it is determined that the path number and process segment form a combination relationship. The coherent adaptation combination relationship provides an executable basis for the subsequent use of path and process logic matching.

[0028] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A shield slag resource utilization classification method based on XGBoost, characterized by: The following steps are involved: Step 1: Based on the four original data items of soil particle size difference, particle friction curvature, dry bulk density of ring cutter specimens, and density change, the propulsion speed sequence and cutterhead torque discrete ratio during construction are matched to establish a partition node and attribute offset combination, and a list of partition node and attribute offset combinations is marked; Step 2: Based on the partition node and attribute offset combination list, extract the direction of mud and sand offset, generate a corresponding symbol sequence, count the intersection frequency areas, extract the particle size extreme points and the cumulative value of the advance to form a ratio comparison sequence, and obtain the disturbance intensity map and classification reference table; Step 3: Based on the disturbance intensity map and classification reference table, extract the offset values ​​of the tunneling parameters and physical property combination items, identify the amplification points and set the node intervals, sort and filter the label conflict numbers in the split nodes, and generate the split judgment path and boundary stability classification labels; Step 4: Based on the splitting judgment path and the boundary stability classification label, extract the number segment label, read the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset value, identify the extreme values ​​of the three items, establish the index number and the change point mapping, and obtain the label segment constraint parameter matrix; Step 5: Based on the label segment constraint parameter matrix, read the earth pressure boundary coefficient and cyclic offset value by block, filter the floating parameter segment, select the intersection that meets the preset path number to bind the path number and process segment combination, and generate the use path matching and process block sequence.

2. The shield slag resource utilization classification method based on XGBoost according to claim 1 is characterized in that: The specific steps for generating the partition node and attribute offset combination list are: Based on the four original data items of soil particle size difference, particle friction curvature, dry bulk density of ring cutter specimens, and density change, the advancement speed sequence and cutterhead torque discrete ratio during construction were matched, aligned along the time axis, and a sample number sequence was established. The data were then paired in chronological order, the sample data were listed, and the sample number sequence was obtained. Based on the aligned sample number sequence, the difference combination of dry density and mud content parameters is extracted, and the interval division is performed based on the particle size difference number and density change value. The boundary of the sample section is determined according to the extreme value of the data, and the interval range is marked to obtain the sample section boundary data; Based on the sample paragraph boundary data, a partition node and attribute offset combination list is generated, the nodes are divided according to the segment boundaries, the attribute value intervals are marked, the parameter difference direction is calculated, and the node information matching the sample paragraph boundary data is generated to obtain a partition node and attribute offset combination list.

3. The shield slag resource utilization classification method based on XGBoost according to claim 1 is characterized in that: The specific steps of generating the disturbance intensity map and classification reference table are as follows: Based on the partition node and attribute offset combination list, extract the offset direction of each section of mud and sand, calculate the offset direction and attribute change in sequence, mark the offset direction, generate a corresponding symbol sequence, and correspond the symbols to the data segments to obtain a symbol sequence; Based on the intersection frequency area of ​​the symbol sequence, traverse the symbols and record the number of occurrences of the intersection points, calculate the distribution of the frequency area, extract the particle size extreme value and the propulsion cumulative value to form a ratio comparison sequence, sort by the particle size extreme value and the propulsion value ratio, screen out the disturbance reference points, organize the numbered sequence structure, and obtain the disturbance reference point sequence; Based on the disturbance reference point sequence, a disturbance intensity mapping and a classification reference table are constructed, the disturbance reference ranking indexes of the disturbance points are summarized, the correlation value between the sample number and the disturbance intensity is calculated, the disturbance intensity information corresponding to the sample number is generated, and the disturbance intensity mapping and the classification reference table are obtained.

4. The shield slag resource utilization classification method based on XGBoost according to claim 1 is characterized in that: The specific steps of generating the split determination path and boundary stability classification label are as follows: Based on the disturbance intensity mapping and classification reference table, a judgment path is constructed and offset values ​​of tunneling parameters and physical property combination items are extracted. After traversing the nodes according to the sample number, the change values ​​related to the tunneling parameters are extracted, the offset of the physical property items is calculated, the amplification points are identified, and the node interval is determined according to the amplification amount to obtain the node interval and amplification point data; Based on the node interval and amplification point data, the leaf node category count density is sorted, the count density of the category in each node is counted in turn, and sorted by density. The numbers of the split nodes with label conflicts are screened, the conflict ratio is extracted, and the difference values ​​between the nodes are calculated to obtain the conflict labels and difference ratio data; Based on the conflicting labels and difference ratio data, principal component analysis is used to match stable path nodes, sort them according to the matching of labels and paths, filter out nodes that conflict with labels one by one, generate split judgment paths and category labels corresponding to nodes, and obtain split judgment paths and boundary stable classification labels.

5. The shield slag resource utilization classification method based on XGBoost according to claim 1 is characterized in that: The principal component analysis is based on the formula: ; in: For the The sample in The weighted projection values ​​in the direction of the principal components, For the summation symbol, is the feature dimension index, is the total number of features, is the sample index, is the principal component index, For the The sample in The value of the original feature dimension, For the The mean of the original features in all samples, For the The original features are The load values ​​in the principal component directions are For the The difference weight coefficient corresponding to the feature dimension, For the The path stability influence coefficient corresponding to each characteristic dimension.

6. The shield slag resource utilization classification method based on XGBoost according to claim 4 is characterized in that: The difference value between the nodes is used to measure the degree of difference in the label distribution in different nodes, and is defined based on a clear statistical distance and distribution similarity measurement method. To enhance the clarity of expression, the difference value should be specifically set as the distance measurement result of the label distribution vector between the nodes. The node can be represented as a label frequency vector, the dimension corresponds to different category labels, and the element value is the frequency of occurrence in the node. The difference value between two nodes can be calculated by the inverse function of the cosine similarity and the Euclidean distance. This method can quantitatively describe the degree of deviation of the category structure between adjacent nodes, thereby assisting in identifying label mutation trends, boundary volatility and potential misclassification areas in actual operations, and enhancing the reliability of the basis for splitting paths and boundary judgments.

7. The shield slag resource utilization classification method based on XGBoost according to claim 1 is characterized in that: The specific steps of generating the label segment constraint parameter matrix are: Based on the splitting judgment path and the boundary stability classification label, the number segment label is extracted, the number segment is traversed and the corresponding earth pressure release coefficient and propulsion fluctuation ratio, and the slag discharge synchronization offset value are read, and then the row data is cleaned and standardized to obtain the number segment label data; Based on the numbered segment label data, the change points in the three parameters are identified, the change amplitudes of the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset value are calculated, the extreme value points of the data are determined, and a mapping relationship between the index number and the change point is established to obtain the change point mapping data; Based on the change point mapping data, a segment output matrix is ​​constructed, the numbered segments are associated with the change points, and the variation ranges of the earth pressure release coefficient, the propulsion fluctuation ratio, and the slag discharge synchronization offset value are compared. The constraint parameter range of each numbered segment is calculated to obtain the label segment constraint parameter matrix.

8. The shield slag resource utilization classification method based on XGBoost according to claim 1 is characterized in that: The specific steps for generating the usage path matching and process block sequence are as follows: Based on the label segment constraint parameter matrix, the algorithm obtains the earth pressure boundary coefficient and cyclic offset values ​​according to the block, sequentially records and calculates the value changes, and screens out the sections where the parameter fluctuation is greater than the target range to obtain the floating parameter section data; Based on the floating parameter segment data, counting the items that overlap with the path condition set, traversing the floating parameter segment, comparing the intersection items with the path condition set, calculating the number of overlapping items, selecting the path numbers whose intersection number meets the target number, and obtaining the extreme value of the overlapping path number data; Based on the extreme values ​​of the overlapping path number data, the dynamic time warping algorithm is used to bind the path number and process segment combination, and matching is performed according to the order of the path number and process segment, so that the path and the process correspond to each other, and the purpose path matching and process block sequence are obtained.

9. The shield slag resource utilization classification method based on XGBoost according to claim 1 is characterized in that: The dynamic time warping algorithm is based on the formula: ; in: For the The path number and The cumulative alignment distance value of each process section, is the index position in the path number sequence, is the index position in the process segment sequence, Number the path Item and process section Matching strength correction factor between items, The first The multidimensional feature vector corresponding to the number, The first The multidimensional feature vector corresponding to each process, is the path feature vector Item and process characteristic vector The Euclidean distance between items, is the power exponential correction factor of the Euclidean distance, From the path number Item is aligned to the process section The cumulative distance value of the item, From the path number Item is aligned to the process section The cumulative distance value of the item, From the path number Item is aligned to the process section The cumulative distance value of the item, Computes the sign for the Euclidean distance between vectors, It is the coordinate position in the two-dimensional path matching table.

10. The shield slag resource utilization classification method based on XGBoost according to claim 8 is characterized in that: The establishment of the combination of path number and process segment needs to be based on the overlapping relationship of time series and the process continuity of construction events. The path number represents the data sampling sequence during the shield advancement process. The number corresponds to the advancement characteristic record within a specific period of time, including advancement speed, cutter head torque and slag discharge volume. The process segment corresponds to the operation stage division defined in the construction log, including the specific process type, start and end time, and the corresponding operation parameter range. For effective binding, the advancement timestamp should be used as the alignment benchmark to determine whether the sample time in the path number sequence completely falls within the time range of the process segment. When the time inclusion relationship is met and the path number sequence maintains the continuity and stability of the advancement parameters in the process segment, it is determined that the path number and the process segment form a combination relationship, and the combination relationship is coherently adapted to provide an executable basis for the subsequent use of path and process logic matching.

Citation Information

Cited By

  • Method and system for predicting resource utilization of shield muck

    CN120971183A

  • A method and system for predicting the resource utilization of shield muck

    CN120971183B