High-temperature alloy GH4169 milling surface roughness modeling method based on XGBoost-RFR algorithm

By combining the XGBoost-RFR algorithm with static milling process parameters and dynamic signals, the complexity of surface roughness modeling during the milling of the high-temperature alloy GH4169 was solved, and accurate milling surface roughness prediction was achieved, which is suitable for the processing of key components in fields such as aerospace.

CN120809006APending Publication Date: 2025-10-17FUZHOU UNIV ZHICHENG COLLEGE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510918554.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional physical modeling methods are difficult to accurately describe the surface roughness of the high-temperature alloy GH4169 during milling, especially under multi-factor coupling and nonlinear characteristics, which makes the modeling complex and inaccurate.

Method used

The XGBoost-RFR algorithm is used to combine the static milling process parameters and dynamic signals of the high-temperature alloy GH4169, and an accurate milling surface roughness model is established through feature selection and model training.

Benefits of technology

It achieves accurate modeling of the milling surface roughness of the high-temperature alloy GH4169, improves the reliability and accuracy of the model, and is suitable for the processing of key components in fields such as aerospace.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120809006A_ABST
    Figure CN120809006A_ABST
Patent Text Reader

Abstract

The invention provides a high-temperature alloy GH4169 milling surface roughness modeling method based on an XGBoost-RFR algorithm. The high-temperature alloy GH4169 milling surface roughness modeling method comprises the following steps that a model is established; the method comprises the steps that S1, static milling process parameters serve as experimental variables, an experimental scheme is designed, dynamic signals in the milling process are collected in the experimental process, and after the experiment is finished, the milling surface roughness of the high-temperature alloy GH4169 is measured; s2, time domain features of the dynamic signals are extracted, and static milling process parameters are combined to be used for model training of a surface roughness model XGBoost-RFR; s3, performing feature selection by using an XGBoost algorithm, and realizing feature optimization by calculating feature importance scores of all features; s4, constructing a milling surface roughness model of the high-temperature alloy GH4169 by taking the features obtained by optimization in the step S3 as input features of an RFR algorithm and the milling surface roughness of the high-temperature alloy GH4169 as output response; according to the invention, a more accurate and reliable model can be established.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of production manufacturing technology, and particularly relates to a high-temperature alloy GH4169 milling surface roughness modeling method based on an XGBoost-RFR algorithm. BACKGROUND

[0002] High-temperature alloys have high strength and hardness, excellent thermal stability, corrosion resistance and fatigue resistance, and are widely used in high-end equipment fields such as aerospace, gas turbines and petroleum chemical industry, and play an important role in the core parts of aircraft engines and gas turbines. According to the difference of the base composition, high-temperature alloys can be divided into three types of iron-based, nickel-based and cobalt-based, among which nickel-based high-temperature alloys are the most widely used. The nickel-based high-temperature alloy takes nickel as the base and dissolves other metal elements, so that it has excellent corrosion resistance and high-temperature oxidation resistance. According to statistics, more than 50% of aircraft engine materials are composed of nickel-based high-temperature alloys. Among the nickel-based high-temperature alloys, GH4169 is the most widely used. Because the high-temperature alloy GH4169 has low thermal conductivity, high strength and high hardness, the adhesion between the material and the tool is strong during the machining process, resulting in poor machining performance, which is only 5% to 20% of that of 45 steel. Therefore, the high-temperature alloy GH4169 is considered as a typical difficult-to-machine material. Milling is one of the main machining methods of GH4169, and the high-temperature alloy GH4169 is often used in key parts in important fields such as aerospace, which work under extreme conditions such as high temperature and high pressure for a long time, so the machining surface quality has higher requirements. Milling is a complex process with multiple factors coupling, and the surface roughness is not only affected by the milling process parameters, but also affected by the vibration, tool wear, process system stability and other dynamic factors. These dynamic data and milling process parameters can provide data support for surface roughness modeling research. The traditional physical modeling method usually depends on specific physical assumptions, and establishes accurate theoretical formula on this basis, but the coupling effect and nonlinear characteristics of milling are difficult to accurately describe by mathematical expressions, so it is more complex to use this method for modeling. The machine learning modeling method is a data-driven modeling method, which can avoid the complex underlying cutting mechanism, directly analyze the processing data, learn the potential relationship between variables, and thus establish a more accurate and reliable model. SUMMARY

[0003] The application provides a high-temperature alloy GH4169 milling surface roughness modeling method based on an XGBoost-RFR algorithm.

[0004] The application adopts the following technical solutions.

[0005] A high-temperature alloy GH4169 milling surface roughness modeling method based on an XGBoost-RFR algorithm comprises the following steps.

[0006] Step S1: taking static milling process parameters as experimental variables, designing an experimental scheme, collecting dynamic signals in the milling process during the experiment, and measuring the high-temperature alloy GH4169 milling surface roughness after the experiment.

[0007] Step S2: extracting time domain features of the dynamic signals, and combining the static milling process parameters (milling speed, feed per tooth, and axial cutting depth) to train the surface roughness model XGBoost-RFR.

[0008] Step S3: performing feature selection using the XGBoost algorithm, and realizing feature optimization by calculating the feature importance scores of all features.

[0009] Step S4: using the features selected in step S3 as the input features of the RFR algorithm, taking the high-temperature alloy GH4169 milling surface roughness as the output response, and constructing the high-temperature alloy GH4169 milling surface roughness model.

[0010] The static milling process parameters include the milling speed, the feed per tooth, and the axial cutting depth, and the dynamic signals in the milling process include the electric signal and the vibration signal.

[0011] Step S1 comprises the following steps, i.e., step S1.1, taking the milling speed, the feed per tooth, and the axial cutting depth as experimental factors, setting the levels of each factor, and designing an orthogonal experiment. min ≤v c ≤v max ; the feed per tooth f min ≤f z ≤f max ; and the axial cutting depth a pmin ≤a p ≤a pmax ;

[0012] wherein v min , v max are minimum and maximum values of the milling speed v c , respectively; f min , f max are minimum and maximum values of the feed per tooth f z , respectively; a pmin , a pmax are minimum and maximum values of the axial depth of cut a p , respectively;

[0013] Step S1.2, according to the experimental scheme, carry out the high-temperature alloy milling experiment, collect the dynamic signal (electric signal, vibration signal) data in the milling process during the experiment, specifically including 18 groups of electric signal parameters (as shown in Table 1) and 3 groups of vibration signals (X axis, Y axis, Z axis) of the machine tool spindle, a total of 21 groups of signals; after the experiment is completed, the surface roughness value after milling is measured.

[0014] Table 1 Electric signal parameter details

[0015]

[0016] Step S2 includes the following steps.

[0017] Step S2.1, extract 11 time domain features (as shown in Table 2) of each group of dynamic signals (electric signals, vibration signals), get 231 (21x11) dynamic characteristics, combine the static milling process parameters (milling speed, feed per tooth, axial depth of cut) to get 234 features, which are used to construct the training data set of XGBoost-RFR model;

[0018] Table 2 Time domain feature calculation formula

[0019]

[0020]

[0021] Step S2.2, the dynamic characteristics obtained by feature extraction and the static milling process parameters are used as input features; the corresponding high-temperature alloy GH4169 milling surface roughness value is used as the output response, and the training data set G = {(x i , y i ) | i = 1, 2,..., n} of the model is constructed, wherein x i represents the input feature vector of the training data, y i represents the output response of the training data, that is, the high-temperature alloy GH4169 milling surface roughness value, and n represents the sample number of the training data set. The training data set G is divided into training set G t and test set Gv ;

[0022] Step S2.3, normalize the input feature vector of the training data set G = {(x i ,y i )|i = 1, 2,..., n}, the normalization method is as follows:

[0023]

[0024] Wherein, x si represents the normalized input feature vector; x i represents the original input feature vector of the training sample; Min(x) represents the minimum value in the input feature vector; Max(x) represents the maximum value in the input feature vector.

[0025] Step S3 includes the following steps:

[0026] Step S3.1, XGBoost is an ensemble learning algorithm, its basic idea is to gradually build a strong model by iteration. In the specific implementation process, XGBoost first uses a weak learner (CART tree is used in the present application) to establish an initial model; then calculate the residual between the model output value and the true value of the initial model, and use it as the new round of training data to train a new weak learner to correct the model error of the previous round, which is repeated continuously. After completing the training of all weak learners, XGBoost outputs the final output result of the model by weightedly integrating the prediction results of each weak learner as the model output result, the model output result is as follows:

[0027]

[0028] Wherein, represents the model output value; K represents the number of trees; f k (x i ) represents the output of the kth CART tree; x i represents the input feature vector of the training sample; F represents the model structure space composed of all CART trees, that is, the strong model obtained by integration;

[0029] Step S3.2, XGBoost trains with the objective function minimization, the objective function L includes the loss function and the penalty term, which is as follows:

[0030]

[0031]

[0032] Wherein, n represents the number of samples; y i represents the true value of the sample output response; is the residual error between the model output value and the true value; Ω(f k ) represents a penalty term for controlling the complexity of the model; γ represents a penalty coefficient for controlling the number of branches of the tree; T represents the number of leaf nodes; λ represents a regularization coefficient of the leaf node weight; w j represents the weight of the leaf node j.

[0033] Step S3.3, when training the t-th CART tree, the objective function is represented as:

[0034]

[0035] wherein, represents the cumulative output value of the first t-1 CART trees; f t (x i ) represents the output of the t-th CART tree.

[0036] Step S3.4, for optimizing the objective function, the loss function is Taylor expanded to be a quadratic function, and the objective function is further simplified by removing the constant term, and the expanded objective function L (t) is as follows:

[0037]

[0038] wherein, g i represents the first-order partial derivative of the loss function with respect to the prediction result ; h i represents the second-order partial derivative of the loss function with respect to the prediction result .

[0039] Step S3.5, if the structure of the t-th CART tree has been determined, i.e., all samples in the training data have been assigned to the T leaf nodes, the weight w j of each leaf node j represents the predicted value of all samples in the node, i.e., f t (x i ) = w j , x i ∈j. Therefore, the objective function can be decomposed according to the contribution of each leaf node, and the decomposed objective function L (t) is as follows:

[0040]

[0041] wherein, I j represents the sample set belonging to the leaf node j; w j represents the weight of the leaf node j.

[0042] Step S3.6, substitute the optimal weight obtained in step S3.5 into the objective function L in S3.5, to obtain the minimized objective function as shown below: (t) Take the derivative of L j to determine the optimal weight , so that L (t) is minimized, and the optimal weight expression is as shown below:

[0043]

[0044] Step S3.7, substitute the optimal weight obtained in step S3.6 into the objective function L in S3.5, to obtain the minimized objective function as shown below: (t)

[0045]

[0046] Step S3.8, when performing node division, traverse all candidate features and division points, respectively calculate the sum of gradients and second-order derivatives of left and right child nodes to determine the node split gain Gain; compare the split gains Gain of all division nodes, select the node with the largest split gain Gain>0 as the best split scheme of the current node for node division. Repeat the above division process, and stop the division when the split gain Gain is less than a threshold value or the maximum number of iterations (i.e., the number of trees) is reached, and finally obtain K trees. Integrate all trees to obtain the final XGBoost model. The specific calculation method of the split gain Gain is as shown below:

[0047]

[0048] wherein G left and G right represent the sum of g i of all samples in the left and right child nodes of the current candidate division node; H left and H right represent the sum of h i of all samples in the left and right child nodes of the current candidate division node;

[0049] Step S3.9, XGBoost performs feature selection based on feature importance, first calculates the feature importance and feature contribution rate of all features in the input feature vector, and selects the features that have a significant impact on the model, and then trains the model using the selected features. The calculation methods of the feature importance I q and the feature contribution rate C q are as shown below:

[0050]

[0051]

[0052] ​​Among them, I q Indicates the feature importance of the qth feature; N q Indicates the total number of times feature q is used for node partitioning in the training of all trees; S q (t) represents the node set in the t-th tree that is partitioned using feature q; Gain s Indicates the splitting gain value of node partitioning using feature q on node s;

[0053] Step S3.10: After calculating the feature importance and contribution rate of all features, the feature importance vector I=[I1,I2,...,I q ,...,I m ], where m represents the number of features and the feature importance I q The larger the value, the more important the feature q is, and vice versa. Features with larger feature importance values ​​are preferred as input features in the training data of the RFR model.

[0054] Step S4 includes the following steps:

[0055] Step S4.1: Use the features obtained in step S3 and the milling surface roughness to construct the training data set G of the RFR model. selected , and divided into training sets in proportion and test set There is a sampling with replacement, forming a training set with the same amount of data as the training set Repeat n times and finally get n sub-training sets;

[0056] Step S4.2: Use grid search combined with O-fold cross validation to determine the number of trees (n_estimators) and the maximum depth (max_depth) of the trees in the RFR algorithm. Assume that there are r hyperparameters that need to be adjusted, and the set of candidate values ​​for the i-th hyperparameter is Then the overall space of hyperparameters Ψ and the total number of combinations β can be expressed as:

[0057]

[0058]

[0059] Step S4.3: Calculate the loss function of the model In the o-fold cross validation, the loss function is calculated as follows:

[0060]

[0061] Among them, |D (o) | represents the number of samples in the o-th fold dataset; yi represents the true value of the i-th sample; represents the predicted value of the i-th sample under the hyperparameters

[0062] The average loss after 10-fold cross-validation is as follows:

[0063]

[0064] Step S4.4, find the optimal parameter combination Optimize the model performance; the optimal hyperparameter combination is the hyperparameter combination with the minimum average error in the validation process:

[0065]

[0066] Step S4.5, after determining the optimal hyperparameter combination, use each sub-training set Train a decision tree independently, and in the construction process of each decision tree, divide each region into two sub-regions in the input space by recursion, and determine the output value of the corresponding region model according to the samples in the region. Finally, a decision tree is generated;

[0067] Step S4.6, train multiple decision trees by repeating the training process of the decision tree, and finally obtain p decision trees M1, M2,..., M i p p . After obtaining the p decision trees, use each decision tree for prediction. The obtained prediction results are represented as c1, c2,..., c i p p ;

[0068] Step S4.7, the prediction results of the p decision trees are comprehensively averaged as the final prediction result of the RFR model; the final prediction result P of the RFR algorithm is the arithmetic mean of the prediction results of the p tree models, and the specific expression is as follows:

[0069]

[0070] The modeling method is used for process simulation of a GH4169 milling process of a hard alloy cutter and under dry milling conditions.

[0071] The cutting speed range of the GH4169 milling process is 90-180 m / min.

[0072] ​In step S1, the surface roughness of each milling experiment needs to be obtained after the experiment is completed, and the surface roughness is measured by a surface roughness measuring instrument, and the specific method is: after each milling experiment is completed, the surface roughness of the workpiece is measured three times, and the average value is taken as the surface roughness value of GH4169 after milling.

[0073] The reliability verification method of the high-temperature alloy GH4169 milling surface roughness model is to randomly select three groups of process parameter combinations from the orthogonal experiment; and the electric signal and the vibration signal collected during the processing of each group of process parameters are respectively substituted into the model for instance verification.

[0074] The method disclosed by the application takes the high-temperature alloy GH4169 static milling process parameters (milling speed, feed per tooth, axial depth of cut) and dynamic signals (electric signal, vibration signal) as the training data set of the high-temperature alloy GH4169 surface roughness model, and the application also proposes an XGBoost-RFR algorithm, which can realize feature selection and high-temperature alloy GH4169 milling surface roughness modeling.

[0075] The machine learning modeling method adopted by the application is a data-driven modeling method, which can avoid complex underlying cutting mechanism, directly analyze the potential relationship between variables through processing data, and thus establish a more accurate and reliable model. BRIEF DESCRIPTION OF DRAWINGS

[0076] The application will be further described in detail below in combination with the drawings and specific embodiments:

[0077] FIG. 1 is a flowchart of the application; Figure 1

[0078] FIG. 2 is a schematic diagram of the feature importance column and the feature cumulative contribution rate curve when the XGBoost algorithm is used to evaluate the feature importance and the feature cumulative contribution rate of all features in the training data in the embodiment; Figure 2

[0079] FIG. 3 is a schematic diagram of the top 10 features in terms of feature importance when the XGBoost algorithm is used to evaluate the feature importance and the feature cumulative contribution rate of all features in the training data in the embodiment. Figure 3 DETAILED DESCRIPTION As shown in the figure, a high-temperature alloy GH4169 milling surface roughness modeling method based on an XGBoost-RFR algorithm includes the following steps:

[0080]

[0081] ​​​Step S1: Take static milling process parameters as experimental variables, design an experimental scheme, collect dynamic signals during the milling process during the experiment, and measure the surface roughness of the high-temperature alloy GH4169 after the experiment;

[0082] Step S2: Extract the time domain features of the dynamic signals, and combine the static milling process parameters (milling speed, feed per tooth, axial depth of cut) for model training of the surface roughness model XGBoost-RFR;

[0083] Step S3: Feature selection using the XGBoost algorithm, and feature optimization by calculating the feature importance scores of all features;

[0084] Step S4: Use the features selected in step S3 as the input features of the RFR algorithm, and the surface roughness of the high-temperature alloy GH4169 as the output response to build a high-temperature alloy GH4169 milling surface roughness model.

[0085] The static milling process parameters include milling speed, feed per tooth, and axial depth of cut, and the dynamic signals during the milling process include electrical signals and vibration signals.

[0086] Step S1 includes the following steps, step S1.1, taking the milling speed, feed per tooth, and axial depth of cut as experimental factors, setting the levels of each factor and designing an orthogonal experiment. The constraints of each factor are as follows: the milling speed v min ≤v c ≤v max ; the feed per tooth f min ≤f z ≤f max ; the axial depth of cut a pmin ≤a p ≤a pmax ;

[0087] wherein v min , v max are the minimum and maximum values of the milling speed v c ; f min , f max are the minimum and maximum values of the feed per tooth f z ; a pmin , a pmax are the minimum and maximum values of the axial depth of cut a p ;

[0088] Step S1.2, according to the experimental scheme, carry out high-temperature alloy milling experiment, collect dynamic signal (electric signal, vibration signal) data in the milling process during the experiment, specifically including 18 groups of electric signal parameters (as shown in Table 1) and 3 groups of vibration signals (X axis, Y axis, Z axis) of the main shaft of the machine tool, a total of 21 groups of signals; after the experiment is completed, the surface roughness value after milling is measured.

[0089] Table 3 Electric signal parameter details

[0090]

[0091]

[0092] Step S2 includes the following steps:

[0093] Step S2.1, extract 11 time domain features (as shown in Table 2) of each group of dynamic signals (electric signals, vibration signals), get 231 (21x11) dynamic characteristics, combine the static milling process parameters (milling speed, feed per tooth, axial depth of cut) to get 234 characteristics, which are used to construct the training data set of XGBoost-RFR model;

[0094] Table 4 Time domain feature calculation formula

[0095]

[0096] Step S2.2, the dynamic characteristics obtained by feature extraction and static milling process parameters are used as input features; the corresponding high-temperature alloy GH4169 milling surface roughness value is used as the output response to form the training data set G = {(x i ,y i )|i=1,2,...,n} of the model, wherein x i represents the input feature vector of the training data, y i represents the output response of the training data, that is, the high-temperature alloy GH4169 milling surface roughness value, and n represents the sample number of the training data set. The training data set G is divided into training set G t and test set G v in proportion;

[0097] Step S2.3, the input feature vector of the training data set G = {(x i ,y i )|i=1,2,...,n} is normalized, and the normalization method is as follows:

[0098]

[0099] Wherein, x si represents the normalized input feature vector; xi represents the minimum value in the input feature vector; Max(x) represents the maximum value in the input feature vector.

[0100] Step S3 includes the following steps:

[0101] Step S3.1, XGBoost is an ensemble learning algorithm, and its basic idea is to gradually build a strong model through iteration. In the specific implementation process, XGBoost first uses a weak learner (CART tree is used in the present application) to establish an initial model; then calculates the residual between the model output value and the true value of the initial model, which is used as new training data for training a new weak learner to correct the model error of the previous round, and the process is repeated continuously. After completing the training of all weak learners, XGBoost outputs the final output result of the model by weighting the prediction results of each weak learner and outputting the final output result of the model as follows:

[0102]

[0103] wherein, represents the model output value; K represents the number of trees; f k (x i ) represents the output of the kth CART tree; x i represents the input feature vector of the training sample; F represents the model structure space composed of all CART trees, i.e. the strong model obtained by integration;

[0104] Step S3.2, XGBoost trains with the objective of minimizing the objective function, and the objective function L includes a loss function and a penalty term, which is as follows:

[0105]

[0106]

[0107] wherein, n represents the number of samples; y i represents the true value of the sample output response; is the loss function, i.e. the residual between the model output value and the true value; Ω(f k ) represents the penalty term, which is used to control the complexity of the model; γ represents the penalty coefficient, which is used to control the number of branches of the tree; T represents the number of leaf nodes; λ represents the regularization coefficient of the leaf node weight; w j represents the weight of the leaf node j.

[0108] Step S3.3, when training the tth CART tree, the objective function is represented as:

[0109]

[0110] wherein, denotes the accumulated output value of the first t-1 CART trees; f t (x i ) denotes the output of the t-th CART tree;

[0111] Step S3.4, for the optimization of the objective function, the loss function is Taylor expanded, approximated as a quadratic function, and further simplified by removing the constant term, the expanded objective function L (t) is as follows:

[0112]

[0113] wherein, g i denotes the first-order partial derivative of the loss function with respect to the prediction result ; h i denotes the second-order partial derivative of the loss function with respect to the prediction result .

[0114] Step S3.5, if the structure of the t-th CART tree has been determined, i.e., all samples in the training data have been assigned to the T leaf nodes, the weight w j of each leaf node j represents the prediction value of all samples in the node, i.e., f t (x i ) = w j , x i ∈j. Therefore, the objective function can be decomposed according to the contribution of each leaf node, and the decomposed objective function L (t) is as follows:

[0115]

[0116] wherein, I j denotes the sample set belonging to the leaf node j; w j denotes the weight of the leaf node j;

[0117] Step S3.6, the objective function L (t) in S3.5 is derived with respect to the weight w j to determine the optimal weight that minimizes L (t) , and the optimal weight expression is as follows:

[0118]

[0119] Step S3.7, the optimal weight obtained in step S3.6 is substituted into the objective function L (t)In this case, the following minimum objective function is obtained:

[0120]

[0121] Step S3.8, when performing node partitioning, the gradient and second derivative sums of the left and right child nodes are calculated respectively, and the node split gain Gain is determined. The split gains Gain of all partition nodes are compared, and the node with the largest split gain Gain > 0 is selected as the best split scheme of the current node for node partitioning. The above partitioning process is repeated, and when the split gain Gain is less than a threshold value, or the maximum number of iterations (i.e. the number of trees) is reached, the partitioning is stopped, and finally K trees are obtained. All trees are integrated to obtain the final XGBoost model. The specific calculation method of the split gain Gain is as follows:

[0122]

[0123] where G left and G right represent the sum of g i of all samples in the left and right child nodes of the current candidate partition node; H left and H right represent the sum of h i of all samples in the left and right child nodes of the current candidate partition node;

[0124] Step S3.9, XGBoost performs feature selection based on feature importance. First, the feature importance and feature contribution rate of all features in the input feature vector are calculated, and the features that have a significant impact on the model are selected according to the feature importance and feature contribution rate, and then the selected features are used for model training. The calculation methods of the feature importance I q and the feature contribution rate C q are as follows:

[0125]

[0126]

[0127] where I q represents the feature importance of the qth feature; N q represents the total number of times that the feature q is used for node partitioning in all tree training; S q (t) represents the node set in which the feature q is used for node partitioning in the tth tree; Gain s represents the split gain value of using the feature q for node partitioning on node s;

[0128] Step S3.10, after calculating the feature importance and contribution rate of all features, a feature importance vector I = [I1, I2,..., Im] containing all features is obtained, where m represents the number of features, and the feature importance Iqrepresents the importance of the feature q. The greater the value of Iq, the more important the feature q is, and vice versa. Preferably, the features with larger feature importance values are used as input features in the training data of the RFR model. q ,...,I m ], wherein m represents the number of features, and the feature importance I q The greater the value of Iq, the more important the feature q is, and vice versa. Preferably, the features with larger feature importance values are used as input features in the training data of the RFR model.

[0129] Step S4 includes the following steps:

[0130] Step S4.1, the features selected in step S3 are used together with the milling surface roughness to construct a training data set G of the RFR model selected , and divided into a training set and a test set with replacement sampling, forming a training set with the same amount of data as the training set, repeated n times, and finally obtaining n sub-training sets;

[0131] Step S4.2, the grid search method is used to determine the number of trees (n_estimators) and the maximum depth (max_depth) in the RFR algorithm combined with O-fold cross-validation. Assuming that there are r hyperparameters to be adjusted, the set of candidate values of the i-th hyperparameter is The overall space of hyperparameters Ψ and the total number of combinations β can be represented as:

[0132]

[0133]

[0134] Step S4.3, calculate the loss function of the model In the o-fold cross-validation, the loss function is calculated as follows:

[0135]

[0136] where |D (o) | represents the number of samples in the o-fold data set; y i represents the true value of the i-th sample; represents the predicted value of the i-th sample under the hyperparameter

[0137] The average loss after O-fold cross-validation is as follows:

[0138]

[0139] ​Step S4.4, finding the optimal parameter combination Optimizing the model performance; the optimal hyperparameter combination is the one with the minimum average error in the validation process:

[0140]

[0141] Step S4.5, after determining the optimal hyperparameter combination, use each sub-training set Independently train a decision tree, and in the construction process of each decision tree, divide each region into two sub-regions in the input space by recursion, and determine the output value of the corresponding region model according to the samples in the region. Finally, a decision tree is generated;

[0142] Step S4.6, train multiple decision trees by repeating the training process of the decision tree, and finally obtain p decision trees M1, M2,..., M i ,...,M p . After obtaining the p decision trees, use each decision tree for prediction. The obtained prediction results are represented as c1, c2,..., c i ,...,c p ;

[0143] Step S4.7, the prediction results of the p decision trees are comprehensively averaged as the final prediction result of the RFR model; the final prediction result P of the RFR algorithm is the arithmetic average of the prediction results of the p tree models, and the specific expression is as follows:

[0144]

[0145] The modeling method is used for the process simulation of the GH4169 milling process of the hard alloy cutter and under the dry milling condition.

[0146] The cutting speed range of the GH4169 milling process is 90-180 m / min.

[0147] In step S1, the surface roughness of each milling experiment needs to be obtained after the experiment is completed. The surface roughness is measured by a surface roughness measuring instrument, and the specific method is: after each milling experiment is completed, the surface roughness of the workpiece is measured three times and the average value is taken as the surface roughness value after the GH4169 milling process.

[0148] The reliability verification method of the high-temperature alloy GH4169 milling surface roughness model is to randomly select three groups of process parameter combinations (the 2nd, 8th and 16th groups) from the orthogonal experiment; the electrical signals and vibration signals collected during the processing of each group of process parameters are combined and substituted into the model for instance verification, and the verification results are shown in Table 13. As can be seen from the table, the maximum absolute error between the model output value of XGBoost-RFR and the actual measured value is 0.041 μm, the minimum absolute error is 0.015 μm, the average absolute error is 0.027 μm, and the root mean square error is 0.029 μm. The comprehensive results show that the XGBoost-RFR algorithm can effectively model the high-temperature alloy GH4169 milling surface roughness.

[0149] Table 5 Model verification results

[0150]

[0151] Embodiment:

[0152] This example is carried out on a vertical machining center, and a high-temperature alloy GH4169 plate with lxd x h = 100 x 80 x 5 mm is side milled. The radial depth of cut is set to a fixed value of 5 mm, and the milling speed, feed per tooth and axial depth of cut are used as factors. A three-factor four-level orthogonal experiment is designed, and the orthogonal experiment table is shown in Table 3.

[0153] Table 6 GH4169 orthogonal experiment design factors and level table

[0154]

[0155] Table 7 Orthogonal experiment design table

[0156]

[0157]

[0158] According to the experimental design in Table 4, a high-temperature alloy GH4169 plate with lxd x h = 100 x 80 x 5 mm is side milled. After the experiment is completed, the surface roughness of each milling experiment needs to be obtained. The surface roughness is measured by a surface roughness measuring instrument. The specific method is as follows: after each milling experiment is completed, the surface roughness of the workpiece is measured three times and the average value is taken as the surface roughness value after GH4169 milling. Table 5 shows the GH4169 milling experiment results, and Tables 6 and 7 show the time domain features extracted from part of the electrical signals and part of the vibration signals.

[0159] Table 8 High-temperature alloy GH4169 milling experiment results

[0160]

[0161] Table 9A Partial time domain features extracted from phase current Ua

[0162]

[0163]

[0164] Table 10 Partial time domain features extracted from spindle X direction vibration signal

[0165]

[0166] XGBoost feature selection. Before feature selection using the XGBoost algorithm, the hyperparameters of the algorithm need to be set, including the number of trees (n_estimators), learning rate (learning_rate) and maximum depth of the tree (max_depth). n_estimators represents the number of weak learners in the model, and an increase in the number helps to improve the feature selection ability, but too much can lead to an increase in the complexity of the XGBoost model and overfitting; learning_rate affects the convergence speed of the model, and a lower value will make the algorithm converge more slowly, but generally better generalization performance will be obtained; a larger max_depth can capture more complex information, but there is also the risk of overfitting. To better improve the performance of the algorithm, a grid search method combined with 5-fold cross-validation is used to optimize the hyperparameters (number of trees, learning rate, maximum depth) in the XGBoost algorithm. The parameter candidate space is: the number of trees candidate interval range is 1-300, step length is 5; learning rate candidate values: 0.01, 0.05, 0.1 and 0.2; the maximum depth of the tree candidate values: 5, 6, 10 and 20; the final determined hyperparameter combination is shown in Table 11.

[0167] Table 11 Hyperparameters of XGBoost algorithm

[0168]

[0169] After the parameter setting is completed, the XGBoost algorithm is used to evaluate the feature importance and feature cumulative contribution rate of all features in the training data. After the evaluation is completed, the feature importance is sorted, and the feature importance bar chart and feature cumulative contribution rate curve chart (such as Figure 2 ) are drawn. As can be seen from the figure, the cumulative contribution rate of the top 30 features in terms of feature importance is 99.04%, therefore, the top 30 features are selected as the input features of the subsequent RFR model. Figure 3The top 10 features in terms of feature importance are plotted. The numbers in the figure correspond to the features as shown in Table 11. From the table, it can be seen that only the feed per tooth (ranked eighth) is a milling process parameter, and the rest are time-domain features of dynamic signals. This shows that dynamic signals (electrical signals, vibration signals) have an important influence on the surface roughness of high-temperature alloy GH4169 milling.

[0170] Table 12 Feature Number and Name Comparison Table

[0171]

[0172] RFR surface roughness modeling. The 30 features selected by feature selection are used as input features, and the milling surface roughness is used as the output response to construct the training data set G of the RFR model selected , and the data set is divided into training set and test set according to the ratio of 8:2. Then, the grid search method is used in combination with 5-fold cross-validation to determine the number of decision trees (n_estimators) and the maximum depth of the tree (max_depth) in the RFR algorithm. The number of trees ranges from 1 to 300, with a step size of 5, and the maximum depth parameters include None (default value), 5, 10, and 20. After hyperparameter optimization, the number of trees is 126 and the maximum depth is the default value (None). Using this parameter combination, the model is retrained to obtain the high-temperature alloy GH4169 milling surface roughness model XGBoost-RFR. The mean absolute error MAE, root mean square error RMSE, and determination coefficient R 2 are used to evaluate the performance of the model, as shown in Table 12.

[0173] Table 13 Model Evaluation Index Output Results

[0174]

Claims

1. A surface roughness modeling method for milling high-temperature alloy GH4169 based on the XGBoost-RFR algorithm, characterized by: The following steps are included: Step S1: Using static milling process parameters as experimental variables, an experimental plan is designed. During the experiment, dynamic signals during the milling process are collected. After the experiment, the milling surface roughness of the high-temperature alloy GH4169 is measured. Step S2: extracting the time domain features of the dynamic signal and combining them with the static milling process parameters for model training of the surface roughness model XGBoost-RFR; Step S3: Use the XGBoost algorithm to perform feature selection and achieve feature optimization by calculating the feature importance scores of all features; Step S4: Using the features obtained in step S3 as input features of the RFR algorithm and the milling surface roughness of the high-temperature alloy GH4169 as the output response, a milling surface roughness model of the high-temperature alloy GH4169 is constructed.

2. The surface roughness modeling method for milling of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 1 is characterized in that: Static milling process parameters include milling speed, feed per tooth, and axial cutting depth. Dynamic signals during the milling process include electrical signals and vibration signals.

3. The surface roughness modeling method for milling of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 1 is characterized in that: Step S1 includes the following steps, Step S1.1: Take milling speed, feed per tooth, and axial depth of cut as experimental factors, set the levels of each factor, and design an orthogonal experiment. The constraints of each factor are as follows: milling speed v min ≤v c ≤v max ; Feed per tooth f min ≤f z ≤f max ;Axial cutting depth a pmin ≤a p ≤a pmax ; Among them, v min 、v max The milling speed v c The minimum and maximum values ​​of f min 、f max The feed per tooth is f z The minimum and maximum values ​​of a pmin 、a pmax Axial cutting depth a p the minimum and maximum values ​​of Step S1.2: Perform a high-temperature alloy milling experiment according to the experimental plan, collect dynamic signal data during the milling process, and measure the surface roughness value after milling after the experiment is completed.

4. The milling surface roughness modeling method of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 1 is characterized by: Step S2 includes the following steps: Step S2.1: Extract multiple time-domain features of each group of dynamic signals to obtain corresponding dynamic features, and combine them with static milling process parameters to obtain 234 features, which are used to construct the training dataset of the XGBoost-RFR model; Step S2.2: The dynamic features and static milling process parameters obtained through feature extraction are used as input features; the corresponding milling surface roughness values ​​of the high-temperature alloy GH4169 are used as output responses to form the training data set G = {(x i ,y i )|i=1,2,...,n}, where x i Represents the input feature vector of the training data, y i represents the output response of the training data, i.e., the milling surface roughness value of the high-temperature alloy GH4169, and n represents the number of samples in the training data set; the training data set G is divided into training sets G t and the test set G v ; Step S2.3: Set the training data set G = {(x i ,y i )|i=1,2,...,n} is normalized. The normalization method is as follows: Among them, x si represents the normalized input feature vector; x i Represents the original input feature vector of the training sample; Min(x) represents the minimum value in the input feature vector; Max(x) represents the maximum value in the input feature vector.

5. The milling surface roughness modeling method of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 1 is characterized in that: Step S3 includes the following steps: Step S3.1: First, use a weak learner to build an initial model; then calculate the residual between the model output value of the initial model and the true value, and use it as the new round of training data to train a new weak learner to correct the model error of the previous round. This process is repeated. After completing the training of all weak learners, XGBoost outputs the prediction results of each weak learner as the final output of the model through weighted integration. The model output result is as follows: in, Represents the model output value; K represents the number of trees; f k (x i ) represents the output of the k-th CART tree; x i Represents the input feature vector of the training sample; F represents the model structure space composed of all CART trees, that is, the strong model obtained by integration; Step S3.2: XGBoost is trained with the goal of minimizing the objective function L. The objective function L includes a loss function and a penalty term, as shown below: Where n represents the number of samples; y i Represents the true value of the sample output response; is the loss function, i.e. the residual between the model output value and the true value; Ω(f k ) represents the penalty term, which is used to control the complexity of the model; γ represents the penalty coefficient, which is used to control the number of branches in the tree; T represents the number of leaf nodes; λ represents the regularization coefficient of the leaf node weight; w j represents the weight of leaf node j; Step S3.3: When training the t-th CART tree, the objective function is expressed as: in, represents the cumulative output value of the first t-1 CART trees; f t (x i ) represents the output of the t-th CART tree; Step S3.4: Taylor expand the loss function to approximate it as a quadratic function, and further simplify the objective function by removing the constant term. The expanded objective function L (t) As shown below: Among them, g i Represents the loss function Forecast results The first partial derivative of h i Represents the loss function Forecast results The second-order partial derivative of ; Step S3.5: If the structure of the t-th CART tree has been determined, that is, all samples in the training data have been assigned to a total of T leaf nodes, then the weight w of each leaf node j is j Represents the predicted value of all samples in the node, that is, f t (x i )=w j ,x i ∈j, decompose the objective function according to the contribution of each leaf node, and the decomposed objective function L (t) As shown below: Among them, I j represents the sample set belonging to leaf node j; w j represents the weight of leaf node j; Step S3.6: Set the objective function L in S3.5 to (t) For weight w j Derivative, determine the optimal weight Make L (t) Minimize, the optimal weight expression is as follows: Step S3.7: The optimal weight obtained in step S3.6 Substitute the objective function L of S3.5 into (t) The minimization objective function is obtained as follows: Step S3.8: When performing node splitting, traverse all candidate features and splitting points, calculate the sum of the gradient and second-order derivative of the left and right child nodes respectively, and determine the node splitting gain Gain; compare the splitting gains of all split nodes, and select the node with the largest splitting gain Gain>0 as the optimal splitting scheme for the current node for node splitting; repeat the above splitting process. When the splitting gain Gain is less than the threshold or reaches the preset maximum number of iterations, stop the splitting and finally obtain K trees; integrate all trees to obtain the final XGBoost model; the specific calculation method of the splitting gain Gain is as follows: Among them, G left and G right Represents the g of all samples in the left and right child nodes of the current candidate partition node i The sum of H left and H right Represents the h of all samples in the left and right child nodes of the current candidate partition node i sum; Step S3.9, XGBoost performs feature selection based on feature importance. First, the feature importance and feature contribution rate of all features in the input feature vector are calculated, and the features that have a significant impact on the model are selected based on the feature importance and contribution rate. The selected features are then used for model training. q and characteristic contribution rate C q The calculation method is as follows: Among them, I q Indicates the feature importance of the qth feature; N q Indicates the total number of times feature q is used for node partitioning in the training of all trees; S q (t) represents the node set in the t-th tree that is partitioned using feature q; Gain s Indicates the splitting gain value of node partitioning using feature q on node s; Step S3.10: After calculating the feature importance and contribution rate of all features, the feature importance vector I=[I1,I2,...,I q ,...,I m ], where m represents the number of features and the feature importance I q The larger the value, the more important the feature q is, and vice versa. Features with larger feature importance values ​​are used as input features in the training data of the RFR model.

6. The milling surface roughness modeling method of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 1 is characterized in that: Step S4 includes the following steps: Step S4.1: Use the features obtained in step S3 and the milling surface roughness to construct the training data set G of the RFR model. selected , and divided into training sets in proportion and test set There is a sampling with replacement, forming a training set with the same amount of data as the training set Repeat n times to obtain n sub-training sets; Step S4.2: Use grid search combined with O-fold cross validation to determine the number of trees n_estimators and the maximum depth max_depth in the RFR algorithm. Assume that there are r hyperparameters that need to be adjusted, and the set of candidate values ​​for the i-th hyperparameter is Then the overall space of hyperparameters Ψ and the total number of combinations β are expressed as: Step S4.3: Calculate the loss function of the model In the o-fold cross validation, the loss function is calculated as follows: Among them, |D (o) | represents the number of samples in the o-th fold dataset; y i Represents the true value of the i-th sample; Indicates the hyperparameters The predicted value of the next i-th sample; The average loss after O-fold cross validation is as follows: Step S4.4: Find the optimal parameter combination Optimize model performance; the optimal hyperparameter combination is the one with the smallest average error during validation: Step S4.5: After determining the optimal hyperparameter combination, use each sub-training set Train a decision tree independently. During the construction of each decision tree, recursively divide each region in the input space into two sub-regions, and determine the output value of the corresponding regional model based on the samples in the region; finally, a decision tree is generated; Step S4.6: Train multiple decision trees by repeating the training process of the decision tree, and finally obtain p decision trees M1, M2, ..., M i ,...,M p .; After obtaining p decision trees, use each decision tree to make predictions; the prediction results are expressed as c1, c2, ..., c i ,...,c p ; Step S4.7: The prediction results of the p decision trees are comprehensively averaged as the final prediction result of the RFR model. The final prediction result P of the RFR algorithm is the arithmetic mean of the prediction results of the p tree models. The specific expression is as follows:

7. The milling surface roughness modeling method of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 1 is characterized by: The modeling method is used for process simulation of the milling process of GH4169 using carbide tools and under dry milling conditions.

8. The milling surface roughness modeling method of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 7 is characterized in that: The cutting speed range of the GH4169 milling process is 90 to 180 m / min.

9. The milling surface roughness modeling method of high-temperature alloy GH4169 based on the XGBoost-RFR algorithm according to claim 7 is characterized in that: In step S1, after the experiment is completed, the surface roughness of each milling experiment needs to be obtained. The surface roughness is measured using a surface roughness measuring instrument. The specific method is as follows: after each milling experiment is completed, the surface roughness of the workpiece is measured multiple times and the average value is taken as the surface roughness value after GH4169 milling.

10. The milling surface roughness modeling method of high-temperature alloy GH4169 based on XGBoost-RFR algorithm according to claim 7, characterized in that: The reliability verification method of the milling surface roughness model of the high-temperature alloy GH4169 is to randomly select process parameter combinations from orthogonal experiments; combine the electrical signals and vibration signals collected during the machining process of each set of process parameters, and substitute them into the model for example verification.

Citation Information

Cited By

  • Online prediction method for milling surface roughness of self-supervised comparative learning thin-walled workpiece

    CN122310473A