A wear prediction method and system based on optimized random forest
Patent Information
- Application Number
- CN202610784145.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-28
AI Technical Summary
[0003]然而现有方法难以兼顾参数优化与预测精度:基于物理模型的预测方法对复杂工况的适应性差,且需大量经验校准参数,泛化能力有限
[0030] This invention constructs a DE-RF radial wear prediction model by organically combining the improved differential evolution algorithm (DE) with random forest (RF). It solves the technical problems of insufficient parameter optimization of random forest, low optimization accuracy of ordinary differential evolution algorithm, and easy premature convergence in the prior art, and realizes the synergistic optimization of parameter optimization and feature extraction.
Smart Images

Figure CN122654864A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mechanical device failure prediction technology, and in particular to a wear prediction method and system based on optimized random forest. Background Technology
[0002] As a core load-bearing component in deep-sea oil and gas extraction and supercritical medium pressurization and transportation equipment, the spiral-structured pressure-bearing pipe component continuously endures thermal erosion and chemical corrosion from the internal high-temperature and high-pressure medium, as well as mechanical coupling wear between the moving mating pairs and the spiral working surface of the inner wall, during continuous equipment operation. These three types of damage superimpose and promote each other, ultimately forming irreparable cumulative damage, manifested as a continuous increase in radial wear on the inner wall of the component. When the radial wear reaches a preset critical failure threshold, core operating indicators such as equipment conveying accuracy, working pressure stability, and operational efficiency decline significantly, and the service life of the core component ends. Accurately predicting the radial wear on the inner wall of the component has significant engineering application value and practical guiding significance for realizing predictive maintenance of industrial equipment, reducing equipment failure and safety risks, and ensuring the continuous and stable operation of high-end equipment.
[0003] However, existing methods struggle to balance parameter optimization and prediction accuracy: prediction methods based on physical models are poorly adaptable to complex working conditions and require extensive empirical parameter calibration, resulting in limited generalization ability.
[0004] Limited model practicality and scalability: Existing methods lack a complete technical solution from data acquisition, preprocessing, model building to online prediction, limiting their engineering application value; moreover, most methods are only applicable to specific types of mechanical parts and are difficult to extend to wear prediction in other scenarios.
[0005] Existing methods have limited prediction accuracy under complex working conditions and the prediction results have high uncertainty, making it difficult to meet the actual needs of industrial sites for high-precision and high-reliability predictions. In particular, in the prediction of radial wear of key components such as pressure-bearing pipe components with spiral structures, even small prediction deviations may lead to serious safety hazards. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a wear prediction method and system that can balance parameter optimization and prediction accuracy, and is highly practical and reliable.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0008] A wear prediction method based on optimized random forest, the key of which includes the following steps:
[0009] S1. Data acquisition and sample construction: The radial wear of hollow pressure-bearing circular tube components with continuous spiral structure on the inner wall is measured and the full-life time series data are obtained. The radial wear is used as the label to divide the training set and the validation set.
[0010] S2. Data preprocessing: Missing values are removed, outliers are eliminated, and data types are converted for the training and validation sets respectively. MinMax normalization is used to eliminate dimensional differences.
[0011] S3. Construct a DE-RF model, which includes an improved differential evolution algorithm module and a random forest module. The improved differential evolution algorithm module is used to optimize the key parameters of the random forest module, and the random forest module is used to predict radial wear.
[0012] S4. Model training and optimization: The sum of the MAE weighting and the average relative error weighting is used as the fitness function to train and improve the DE-RF model, and the optimal model is saved; at the same time, an RF baseline model is constructed.
[0013] S5. Online wear prediction: Input real-time online data into the trained model, output radial wear prediction values, and evaluate the accuracy of the prediction results.
[0014] Preferably, in step S1, data acquisition and sample construction, the radial wear amount full-lifetime time-series data includes input feature parameters and output target parameters, wherein the input feature parameter is the total service time and the output target parameter is the radial wear amount.
[0015] Preferably, in step S2, data preprocessing, the normalization process calculates the mean and standard deviation based on the training set, and the validation set is transformed using the same normalization parameters.
[0016] Preferably, in step S3. Constructing the DE-RF model, the parameters of the improved differential evolution algorithm module are: population size 30, maximum number of iterations 80, initial mutation factor 0.5, crossover probability 0.9, number of consecutive iterations without improvement 12, MAE weight 0.7, and average relative error weight 0.3; the parameter range of the random forest module is: number of decision trees [100, 200], tree depth [15, 25], maximum number of features [1, 1], and minimum number of samples per leaf node [1, 2].
[0017] Preferably, in step S3. Constructing the DE-RF model, the improved differential evolution algorithm module adopts an adaptive mutation factor, with the initial mutation factor set to 0.5, which decreases with a decay coefficient of 0.98 as the number of iterations increases.
[0018] Preferably, in step S4. Model training and optimization, the fitness function The calculation formula is as follows:
[0019]
[0020] in The mean absolute error of the validation set, The average relative error of the validation set is used; a greedy selection strategy is employed during training to update the globally optimal parameter combination.
[0021] Preferably, in step S4. Model training and optimization, an RF baseline model is constructed simultaneously, with the parameter range set as follows: number of decision trees [80, 100], tree depth [8, 10], maximum number of features [1, 1], minimum number of leaf nodes [3, 4]. The grid search method is used for optimization, and the average relative error is controlled within 0.1%.
[0022] Preferably, in step S5. Online wear prediction, the accuracy assessment includes calculating the MAE, ARE, maximum relative error, and R² of the improved DE-RF model and the RF baseline model.
[0023] A wear prediction system based on optimized random forest, used to implement the aforementioned wear prediction method based on optimized random forest, is characterized by including:
[0024] The data acquisition module is used to acquire the full-life time-series data of radial wear of hollow pressure-bearing circular tube components with a continuous spiral structure on the inner wall;
[0025] The preprocessing module is used to perform equivalent coefficient conversion and standardization on the data, as well as to handle missing values, outliers, data type conversion, and normalization.
[0026] The model building module constructs a wear prediction model that optimizes the random forest using an improved differential evolution algorithm, namely the DE-RF model.
[0027] The training optimization module is used for model training and parameter optimization, and dynamically saves the optimal model.
[0028] The online prediction module calculates the input online real-time data and outputs the radial wear prediction value.
[0029] The beneficial effects of adopting the above technical solution are as follows:
[0030] This invention constructs a DE-RF radial wear prediction model by organically combining the improved differential evolution algorithm (DE) with random forest (RF). It solves the technical problems of insufficient parameter optimization of random forest, low optimization accuracy of ordinary differential evolution algorithm, and easy premature convergence in the prior art, and realizes the synergistic optimization of parameter optimization and feature extraction.
[0031] This invention, through a complete five-step implementation process—from data acquisition, preprocessing, model building, training optimization to online prediction—forms a comprehensive radial wear prediction solution. The model has low computational complexity, does not rely on complex physical mechanisms, and only requires the collection of total service time and radial wear data. The preprocessing process is simple, and the model training and prediction efficiency is high. It can be directly applied to the radial wear prediction of mechanical components such as pressure-bearing pipe members with helical structures, providing accurate data support for equipment maintenance and repair, and thus possessing greater practicality.
[0032] The improved DE-RF model in this invention combines the parameter optimization advantages of improved DE with the anti-overfitting and feature extraction advantages of RF, which can effectively extract key features of radial wear, reduce noise interference, and improve prediction accuracy. Attached Figure Description
[0033] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0034] Figure 1 This is a flowchart of a wear prediction method based on optimized random forest proposed in this invention;
[0035] Figure 2 This is a flowchart of the dataset preprocessing process in an embodiment of the present invention;
[0036] Figure 3 This is a flowchart of the improved differential evolution algorithm optimization process in an embodiment of the present invention;
[0037] Figure 4 This is a comparison chart of the errors between the DE-RF model and the pure RF model in an embodiment of the present invention;
[0038] Figure 5 This is a flowchart illustrating the model training and optimization process in an embodiment of the present invention.
[0039] Figure 6 This is the fitness loss curve of the DE-RF model in an embodiment of the present invention;
[0040] Figure 7 This is the radial wear prediction result in the embodiment of the present invention;
[0041] Figure 8 This represents the 95% confidence interval for the radial wear prediction results in this embodiment of the invention. Detailed Implementation
[0042] To make the above-mentioned objectives, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] This invention proposes a wear prediction method and system based on optimized random forest, aiming to overcome the shortcomings of existing technologies such as offline detection lag, low prediction accuracy, parameter dependence on human experience, imbalance between global search and local development, weak model generalization ability, and poor engineering practicality. It introduces an improved differential evolution algorithm (DE) to optimize the parameter settings of random forest (RF), constructing a DE-RF model to predict radial wear. By improving the parameter settings of the differential evolution algorithm, the accuracy and efficiency of parameter optimization are enhanced. Combined with the anti-overfitting advantage of random forest, the accuracy of radial wear prediction is improved while reducing computational complexity. Specifically, for the nonlinear temporal characteristics of radial wear in hollow pressure-bearing circular pipe components with a continuous spiral structure on the inner wall, accurate prediction of radial wear is achieved, significantly improving prediction accuracy, model robustness, and engineering practicality, reducing operation and maintenance costs, and ensuring the safe and stable operation of equipment.
[0044] This invention includes a wear prediction method based on optimized random forest, such as... Figure 1 Specifically, the steps include:
[0045] S1. Data acquisition and sample construction: The radial wear of hollow pressure-bearing circular tube components with continuous spiral structure on the inner wall is measured and the full-life time series data are obtained. The radial wear is used as the label to divide the training set and the validation set.
[0046] S2. Data preprocessing: Missing values are removed, outliers are eliminated, and data types are converted for the training and validation sets respectively. MinMax normalization is used to eliminate dimensional differences.
[0047] S3. Construct a DE-RF model, which includes an improved differential evolution algorithm module and a random forest module. The improved differential evolution algorithm module is used to optimize the key parameters of the random forest module, and the random forest module is used to predict radial wear.
[0048] S4. Model training and optimization: The sum of the MAE weighting and the average relative error weighting is used as the fitness function to train and improve the DE-RF model, and the optimal model is saved; at the same time, an RF baseline model is constructed.
[0049] S5. Online wear prediction: Input real-time online data into the trained model, output radial wear prediction values, and evaluate the accuracy of the prediction results.
[0050] S1. Data Acquisition and Sample Construction
[0051] In this embodiment, the radial wear full-life time-series data of a hollow pressure-bearing circular pipe component with a continuous spiral structure on the inner wall in a certain type of booster conveying equipment is obtained by using a laser diameter gauge or an ultrasonic thickness gauge. The corresponding radial wear amount is labeled as a tag, and the dataset is randomly divided into a training set and a validation set in a 7:3 ratio to provide basic data support for model training and evaluation. The input feature parameter is the total service time, and the output target parameter is the radial wear amount, which represents the cumulative wear time and the wear degree, respectively. For the total service time label, the time interval between data is adaptively obtained according to the purpose of the equipment to which the component is composed.
[0052] During data acquisition, the service conditions of the components are strictly recorded, including working pressure, medium temperature, medium flow rate, vibration status, etc., to ensure that the external conditions are consistent within the same sample group. The total number of samples is preferably 800-1200 groups, balancing data richness and training efficiency; in this embodiment, 1000 groups of valid samples are preferred.
[0053] Using radial wear as the label, a stratified random sampling method was employed to divide the dataset into training and validation sets in a 7:3 ratio: 70% of the samples were used as the training set for DE-RF model parameter training, feature learning, and model construction; 30% of the samples were used as the validation set for model performance evaluation, parameter optimization, and generalization capability verification. Strict separation of the training and validation sets prevented data leakage and ensured the reliability of the model's generalization ability and evaluation results.
[0054] S2. Data Preprocessing
[0055] Wear monitoring data is susceptible to environmental noise, equipment interference, measurement errors, and signal loss during the acquisition process, resulting in issues such as missing values, outliers, inconsistent dimensions, and mixed data types. These problems directly affect model training performance and prediction accuracy. Therefore, systematic preprocessing of the training and validation sets is necessary to eliminate noise interference, standardize data formats, eliminate dimensional differences, and improve data quality and model stability.
[0056] In the examples, such as Figure 2 Data preprocessing is performed on both the training and validation sets to eliminate data noise and dimensional differences; specifically including:
[0057] Missing value handling: Traverse the dataset to identify samples with empty input features (total service time) or output labels (radial wear). Use the dropna function to directly delete samples with missing values to avoid interference from missing data in model training and ensure data integrity.
[0058] Outlier Removal: Outliers mainly fall into three categories: incorrect data types, abrupt numerical changes, and values exceeding the physically reasonable range. First, the `pd.to_numeric` function is used to convert the data type, removing non-numerical outliers that cannot be converted to numerical values. Second, based on the 3σ principle, the mean μ and standard deviation σ of radial wear are calculated, removing abrupt outliers exceeding the range [μ-3σ, μ+3σ]. Finally, considering the component material properties and service conditions, unreasonable outliers with radial wear values less than 0 or greater than a preset safety threshold are removed to ensure the physical reasonableness of the data.
[0059] Data normalization: Input features and output labels have different dimensions and large differences in numerical range. Directly inputting them into the model can easily lead to parameter update imbalance, low training efficiency, and slow convergence. The MinMax normalization method is used to linearly map the data to the [0,1] interval, eliminating the difference in dimensions and unifying the data scale. Furthermore, normalization parameters such as mean, standard deviation, minimum, and maximum values are calculated only based on the training set data. The validation set data is directly transformed using the normalization parameters of the training set to avoid data leakage and ensure the model's generalization ability and the objectivity of the evaluation results.
[0060] S3: Constructing the DE-RF model
[0061] To address the characteristics of radial wear data from hollow pressure-bearing circular pipe components with a continuous spiral structure on the inner wall, a hybrid prediction model combining an improved differential evolution algorithm (DE) and random forest (RF) is constructed. The improved DE algorithm adaptively optimizes key RF parameters globally, solving the problems of traditional RF parameters relying on human experience and easily getting trapped in local optima. This fully leverages the advantages of the RF algorithm in resisting overfitting and fitting nonlinear time-series data, accurately capturing the wear evolution pattern. The overall model structure consists of two parts: an improved differential evolution algorithm module and a random forest module. The specific construction process is as follows:
[0062] First, an improved differential evolution algorithm module is constructed to optimize the traditional differential evolution algorithm. The algorithm parameters are set as shown in Table 1. The population size is set according to the training set size to balance population diversity, global search capability, and computational complexity. The maximum number of iterations is set to 80 to balance optimization efficiency and capability. Crossover probability ensures population diversity and reliability of optimization results. The MAE weight is set to 0.7 to focus on controlling absolute error and ensuring prediction accuracy. The average relative error weight is set to 0.3 to balance relative error and control the stability of wear prediction results for different stages. Furthermore, to address the shortcomings of a fixed mutation factor, such as imbalance between global exploration and local development, low convergence accuracy in later stages, and premature convergence, an adaptive mutation factor strategy is adopted. The initial mutation factor is 0.5, which decreases with a decay coefficient of 0.99 as the number of iterations increases. = In the early stages of iteration, the mutation factor is large, which enhances the global exploration capability, expands the search range, and avoids local optima; in the later stages of iteration, the mutation factor is small, which strengthens the local development capability, improves the convergence accuracy, and accelerates the convergence, ensuring a wide optimization range in the early stages and high optimization accuracy in the later stages.
[0063] Table 1. Parameters of the Improved Differential Evolution Algorithm
[0064] Population size 30 Maximum number of iterations 80 Initial variation factor 0.5 Crossover probability 0.9 The number of consecutive iterations without improvement is terminated. 12 MAE weighting 0.7 Average relative error weight 0.3
[0065] like Figure 3 The specific process is as follows:
[0066] 3.1 Population initialization: Within the preset range of random forest parameters, 30 population individuals are randomly generated, and each population individual corresponds to a set of random forest parameter combinations;
[0067] Considering the characteristics of radial wear data of hollow pressure-bearing circular tube components with continuous spiral structure on the inner wall, such as single input features, strong temporal sequence, nonlinearity, and large noise interference, after multiple sets of comparative experiments, a reasonable optimization range of key parameters of the random forest module was set to avoid invalid search and improve optimization efficiency. The range of random forest parameters is shown in Table 2.
[0068] Table 2 Random Forest Parameters
[0069] Number of decision trees [100,200] Tree depth [15,25] Maximum number of features 1 Minimum number of samples for leaf nodes [1,2]
[0070] The improved differential evolution algorithm module acts as the upper-level optimizer, responsible for globally searching for the optimal combination of parameters in a random forest within a preset parameter range; the random forest module acts as the lower-level predictor, responsible for building the model based on the optimal parameters, training and learning, and achieving wear prediction. The two are deeply integrated, and the optimization steps are as follows:
[0071] 3.2 Fitness Function Calculation: For each individual in the population, a random forest model is constructed. The model is trained using the training set, and MAE and ARE are calculated on the validation set according to the formula.
[0072]
[0073] Calculate the fitness value P. The smaller the fitness value, the better the corresponding combination of random forest parameters. W1 and W2 are adjustment weights, which are selected according to the emphasis on the site. In the example, when the emphasis is on absolute error, W1=0.7 and W2=0.3 can be selected.
[0074] 3.3 Mutation Operation: An adaptive mutation factor is used for mutation to ensure a wide optimization range in the early stage and high optimization accuracy in the later stage. The mutation formula is as follows:
[0075]
[0076] in, As a mutated individual, , , Three distinct individuals were randomly selected from the population. The t-th generation adaptive mutation factor enhances population diversity, expands the search range, and avoids local optima.
[0077] 3.4 Crossover operation: A binomial crossover strategy is adopted, with a fixed crossover probability of 0.9, to cross the mutant individuals with the parent individuals to generate experimental individuals, thereby enhancing population diversity and improving search efficiency.
[0078] 3.5 Selection Operation: A greedy selection strategy is adopted. The fitness values of the experimental individuals are compared with those of their parents. Individuals with better fitness values are retained as members of the next generation of the population. The global optimal parameter combination is updated to ensure that the population evolves towards the optimal solution.
[0079] 3.6 Iteration Termination: When the maximum number of iterations (80) is reached or the fitness value has not improved for 12 consecutive generations, the iteration is terminated, and the globally optimal combination of random forest parameters is output.
[0080] Furthermore, a random forest module is constructed, and the globally optimal parameter combination output by the improved DE module is substituted into the random forest model. The model is then trained using the training set to obtain the DE-RF prediction model.
[0081] S4: Model Training and Optimization
[0082] like Figure 5 The fitness function is the sum of the MAE weighted value and the mean relative error (ARE) weighted value. The formula for calculating the fitness function is as follows:
[0083]
[0084] in The mean absolute error of the validation set, The average relative error on the validation set; a greedy selection strategy is used for model training, and the training and validation errors are dynamically monitored, such as... Figure 6 The iteration process terminates when the maximum number of iterations (80) is reached or the fitness value shows no improvement for 12 consecutive generations, and the optimal DE-RF model is saved. Simultaneously, an RF baseline model is constructed, and its parameter range is set as shown in Table 3. Then, the GridSearchCV function is used for grid search, and 5-fold cross-validation is employed, with negative MAE as the evaluation metric, to optimize the parameters of the RF baseline model. Ultimately, the average relative error of the RF baseline model is kept below 0.1%, which is used to compare and verify the advantages of the DE-RF model.
[0085] Table 2 Parameter settings for the RF baseline model
[0086] Number of decision trees [80,100] Tree depth [8,10] Minimum number of samples for leaf nodes [3,4] Maximum number of features 1
[0087] During training, the fitness loss of the training and validation sets was tracked in real time. Both the training and validation losses gradually stabilized, and the final loss value dropped to an extremely low level, indicating that the model training process had good convergence. The trend of the validation loss was highly consistent with that of the training loss, indicating that the model had good generalization ability and effectively learned the inherent laws of radial wear-related data.
[0088] S5: Online Wear Prediction
[0089] The total service time of the hollow pressure-bearing circular tube component with a continuous spiral structure on the inner wall, which is monitored online, is input into the trained DE-RF model. Finally, the radial wear prediction value of the hollow pressure-bearing circular tube component with a continuous spiral structure on the inner wall is output under the current operating environment.
[0090] Results Validation: To comprehensively, objectively, and quantitatively evaluate the predictive performance of the DE-RF model and the RF baseline model, four core evaluation indicators were selected to comprehensively measure model performance from four dimensions: absolute error, relative error, extreme error, and goodness of fit.
[0091] Mean Absolute Error (MAE): Reflects the average absolute deviation between the predicted value and the actual value, in mm. The smaller the value, the higher the prediction accuracy and the smaller the deviation.
[0092] Mean Relative Error (ARE): Reflects the average relative deviation between the predicted value and the actual value, expressed as a percentage. The smaller the value, the higher the prediction accuracy and the better the stability at different wear stages.
[0093] Maximum relative error (MaxRE): Reflects the upper limit of the model's extreme prediction error, expressed as a percentage. The smaller the value, the stronger the model's stability and the higher its reliability under extreme conditions.
[0094] Coefficient of determination (R) 2): Reflects the degree of fit of the model to real wear data, with a value range of [0,1]. The closer to 1, the better the fit and the more accurately the model captures the wear pattern.
[0095] In practical use, both should be evaluated, such as Figure 4 The mean relative error (MRE) of the DE-RF model was 0.03%–0.05%, while that of the RF baseline model was 0.08%–0.10%. The MAE of the DE-RF model was reduced by more than 20% compared to the RF baseline model, and the R² was improved to over 0.999. The predicted curves closely matched the actual curves. Figure 7-8 The predicted values are within the 95% confidence interval, indicating that the uncertainty of the model prediction is controllable and its reliability is high.
[0096] This technical solution acquires data on the total service time and radial wear of hollow pressure-bearing circular pipe components with a continuous spiral structure on the inner wall through actual measurement, providing sample data for wear prediction. By handling missing values, converting data types, and standardizing preprocessing, the data quality is effectively improved and noise interference is reduced, providing more accurate input data for subsequent DE-RF model training. Combined with an improved DE and RF fusion architecture, the system first optimizes the key parameters of RF by improving the DE module, and then extracts radial wear features and makes predictions through the RF module, significantly improving the accuracy of radial wear pattern recognition and prediction.
[0097] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A wear prediction method based on optimized random forest, characterized in that, Includes the following steps: S1. Data acquisition and sample construction: The radial wear of hollow pressure-bearing circular tube components with continuous spiral structure on the inner wall is measured and the full-life time series data are obtained. The radial wear is used as the label to divide the training set and the validation set. S2. Data preprocessing: Missing values are removed, outliers are eliminated, and data types are converted for the training and validation sets respectively. MinMax normalization is used to eliminate dimensional differences. S3. Construct a DE-RF model, which includes an improved differential evolution algorithm module and a random forest module. The improved differential evolution algorithm module is used to optimize the key parameters of the random forest module, and the random forest module is used to predict radial wear. S4. Model training and optimization: The sum of the MAE weighting and the average relative error weighting is used as the fitness function to train and improve the DE-RF model, and the optimal model is saved; at the same time, an RF baseline model is constructed. S5. Online wear prediction: Input real-time online data into the trained model, output radial wear prediction values, and evaluate the accuracy of the prediction results.
2. The wear prediction method based on optimized random forest according to claim 1, characterized in that, In step S1, data acquisition and sample construction, the radial wear amount full-lifetime time-series data includes input feature parameters and output target parameters, wherein the input feature parameter is the total service time and the output target parameter is the radial wear amount.
3. The wear prediction method based on optimized random forest according to claim 1, characterized in that, In step S2, data preprocessing, the normalization process calculates the mean and standard deviation based on the training set, and the validation set is transformed using the same normalization parameters.
4. The wear prediction method based on optimized random forest according to claim 1, characterized in that, In step S3, constructing the DE-RF model, the parameters of the improved differential evolution algorithm module are: population size 30, maximum number of iterations 80, initial mutation factor 0.5, crossover probability 0.9, number of consecutive iterations without improvement 12, MAE weight 0.7, and average relative error weight 0.3; the parameters of the random forest module are: number of decision trees [100, 200], tree depth [15, 25], maximum number of features [1, 1], and minimum number of samples per leaf node [1, 2].
5. The wear prediction method based on optimized random forest according to claim 1, characterized in that, In step S3, constructing the DE-RF model, the improved differential evolution algorithm module adopts an adaptive mutation factor. The initial mutation factor is set to 0.5, and it decreases with a decay coefficient of 0.98 as the number of iterations increases.
6. The wear prediction method based on optimized random forest according to claim 1, characterized in that, In step S4, model training and optimization, the fitness function... The calculation formula is as follows: ; in The mean absolute error of the validation set, The average relative error of the validation set is used; a greedy selection strategy is employed during training to update the globally optimal parameter combination.
7. The wear prediction method based on optimized random forest according to claim 1, characterized in that, In step S4, model training and optimization, an RF baseline model is constructed simultaneously, with the parameters set as follows: number of decision trees [80, 100], tree depth [8, 10], maximum number of features [1, 1], minimum number of leaf nodes [3, 4]. The grid search method is used for optimization, and the average relative error is controlled within 0.1%.
8. The wear prediction method based on optimized random forest according to claim 1, characterized in that, In step S5, online wear prediction, the accuracy assessment includes calculating the MAE, ARE, maximum relative error, and R² of the improved DE-RF model and the RF baseline model.
9. A wear prediction system based on optimized random forest, used to implement the wear prediction method based on optimized random forest as described in claims 1-8, characterized in that, include: The data acquisition module is used to acquire the full-life time-series data of radial wear of hollow pressure-bearing circular tube components with a continuous spiral structure on the inner wall; The preprocessing module is used to perform equivalent coefficient conversion and standardization on the data, as well as to handle missing values, outliers, data type conversion, and normalization. The model building module constructs a wear prediction model that optimizes the random forest using an improved differential evolution algorithm, namely the DE-RF model. The training optimization module is used for model training and parameter optimization, and dynamically saves the optimal model. The online prediction module calculates the input online real-time data and outputs the radial wear prediction value.