Motor train unit wheel polygon prediction method and system based on random forest algorithm
By applying a random forest algorithm in the polygon wear prediction of wheels of EMUs, using operation and maintenance historical data and correlation coefficients to screen features, multiple CART regression trees are constructed, which solves the problems of large data demand and insufficient multi-dimensional feature processing capabilities in the existing technology, and achieves higher prediction accuracy and reliability.
Patent Information
- Application Number
- CN202510277130.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems such as large data demand, insufficient data processing capability for multi-dimensional feature, and limited generalization performance and prediction accuracy in the prediction prediction.
The EMU wheel polygon prediction method based on the random forest algorithm is adopted. By obtaining operation and maintenance historical data, selecting feature indicators with correlation coefficients that meet the requirements, using Bootstrap self-service sampling method to generate multiple random sample sets, and multiple CART regression trees are constructed based on the classification regression algorithm. The output mean of multiple trees is taken as the prediction result, and feature selection is optimized to improve model accuracy.
It improves the accuracy and reliability of wheel polygon state prediction, can effectively process multi-dimensional feature data, and enhances the generalization performance and prediction accuracy of the model.
Smart Images

Figure CN120217316A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of state detection of multiple unit trains. Specifically, it relates to a prediction method and system for wheel polygons of multiple unit trains based on the random forest algorithm. Background Art
[0002] The content of this part only provides background information related to this application, and it may not constitute prior art.
[0003] With the continuous increase in the operating speed of high-speed multiple unit trains, the interaction between wheels and rails has become more intense, and the problem of wheel polygon wear has gradually become a key factor affecting the running safety and comfort of multiple unit trains. Wheel polygon wear not only intensifies the wheel-rail interaction and changes the dynamic performance of the vehicle, but also generates high-frequency vibrations and noises during the high-speed operation of the vehicle, seriously affecting the comfort of passengers. Therefore, accurately predicting and evaluating the state of wheel polygons is of great significance for the maintenance and operation of multiple unit trains.
[0004] In the prior art, Chinese Patent with publication number CN113947130B discloses a training method and a usage method for a wheel polygon wear waveform regression prediction AI model. This method collects vibration detection data during the running process of the wheel, extracts vibration intensity feature vectors and polygon wear condition labels, and uses a one-dimensional convolutional neural network (1-DCNN) for training to achieve the regression prediction of the wheel polygon wear waveform. Although this method can provide waveform details of wheel polygon wear to a certain extent, such as wavelength information and wave depth characteristics, there are still deficiencies. Specifically, it requires a large amount of vibration detection data for model training, which may be limited by data collection conditions in practical applications. Secondly, this method mainly relies on vibration detection data and does not fully consider other factors that may affect wheel polygon wear, such as wheel radius, turning history, etc. In addition, although the one-dimensional convolutional neural network has advantages in processing time series data, there may be certain limitations in generalization performance and prediction accuracy when dealing with complex and changeable wheel polygon wear problems.
[0005] Therefore, there is a need for a prediction and evaluation method and system for wheel polygons of multiple unit trains based on the random forest algorithm, which has good generalization performance, can effectively process multi-dimensional feature data, and improve the accuracy and reliability of prediction. Summary of the Invention
[0006] In order to solve the above technical problems, the purpose of this application is to provide a prediction method and system for wheel polygons of multiple unit trains based on the random forest algorithm. By utilizing the integration advantages of the random forest algorithm and the collaborative work of multiple decision trees, the accuracy and reliability of predicting the state of wheel polygons are improved.
[0007] The object of the present application is achieved through the following technical solutions:
[0008] In a first aspect, the present invention provides a method for predicting the polygon of a multiple unit train wheel based on a random forest algorithm, including:
[0009] Obtain the operation and maintenance historical data of the multiple unit train wheels;
[0010] Select the characteristic indexes with correlation coefficients meeting the requirements from the operation and maintenance historical data as the training sample feature vectors, and randomly divide them into a training set and a test set;
[0011] Use the Bootstrap self-sampling method to randomly extract the feature information of the training set to generate multiple random sample sets;
[0012] Generate CART regression trees for multiple random sample sets based on a classification regression algorithm, and calculate the Gini coefficient of the random sample sets; recursively divide the nodes according to the Gini coefficient until the Gini coefficient is less than the threshold or cannot be further divided to generate multiple CART regression trees;
[0013] Take the output mean of multiple CART regression trees as the prediction result to obtain the initial prediction model of the multiple unit train wheel polygon;
[0014] Input the test set into the initial prediction model, perform multiple random samplings on the feature information of the test set, and put the sampled feature information back into the test set for the next random sampling; repeatedly construct the model, select the feature with the highest accuracy and put it into the optimal feature subset, and repeat the random sampling for the remaining features until all features are traversed and then stop to obtain the final prediction model;
[0015] Input the data to be processed into the final prediction model to obtain the prediction result.
[0016] Further, the step of recursively dividing the nodes according to the Gini coefficient until the Gini coefficient is less than the threshold or cannot be further divided to generate multiple CART regression trees specifically includes:
[0017] In multiple random sample sets where the current node of the regression tree is located, if the Gini coefficients of the multiple random sample sets are less than the Gini coefficient threshold, the classification regression algorithm returns the sub-decision tree and stops recursion; if the Gini coefficients of the multiple random sample sets are greater than or equal to the Gini coefficient threshold, calculate the Gini coefficient of each feature information in the random sample set where the current node is located;
[0018] Select the feature information corresponding to the minimum Gini coefficient as the optimal feature information, divide the random sample set corresponding to the optimal feature information into a first feature set and a second feature set, and use the first feature set and the second feature set as the left child node and the right child node of the current node respectively;
[0019] Calculate the Gini coefficients of the left child node and the right child node. If the Gini coefficients of the left child node and the right child node are greater than or equal to the Gini coefficient threshold, continue to recursively generate the left child node and the right child node of the current node, and calculate the corresponding Gini coefficients; if the Gini coefficients of the left child node and the right child node are less than the Gini coefficient threshold, the classification and regression algorithm returns the sub-decision tree and stops recursion to generate a CART regression tree.
[0020] Further, select the characteristic indexes with correlation coefficients meeting the requirements from the operation and maintenance historical data as the training sample feature vectors, specifically including:
[0021] Adopt the Pearson correlation analysis method to calculate the correlation between each feature and the amplitude of the polygon high-order polygon based on a preset formula to quantitatively reflect the linear relationship between two sets of data.
[0022] Further, the preset formula includes:
[0023]
[0024] Among them, ρ xy represents the Pearson correlation coefficient between variable x and variable y, x i is the x value of the i-th sample, is the sample mean of x, y i is the y value of the i-th sample, is the sample mean of y, and n is the number of samples.
[0025] Further, in the step of selecting the feature with the highest accuracy and putting it into the optimal feature subset, the calculation formula of the accuracy is:
[0026]
[0027] Among them, TP is predicted as the positive class and is actually the positive class; TN is predicted as the negative class and is actually the negative class; FP is predicted as the positive class and is actually the negative class; FN is predicted as the negative class and is actually the positive class.
[0028] Further, the characteristic indexes include the wheel diameter, the machining deviation of the wheel flange thickness, the predicted maximum value of the wheel pair polygon before the previous turning, the machining deviation of the wheel diameter, the vehicle operation mileage, the cumulative operation mileage after the previous turning, the car body type, the axle number, the season, and the presence or absence of the wheel polygon fault.
[0029] Further, after inputting the data to be processed into the final prediction model and obtaining the prediction result, it also includes:
[0030] Display the prediction result in the form of a chart, and the chart includes a trend chart of the predicted wheel polygon amplitude over time;
[0031] Compare the prediction results with a preset safety threshold and mark the prediction results that exceed the safety threshold in the chart.
[0032] In a second aspect, the present invention provides a polygon prediction system for EMU wheels based on the random forest algorithm, including:
[0033] A data acquisition module for acquiring the operation and maintenance historical data of EMU wheels;
[0034] An influencing factor division module for selecting characteristic indicators with a correlation coefficient meeting the requirements from the operation and maintenance historical data as the training sample feature vectors, and randomly dividing them into a training set and a test set;
[0035] A sample generation module for randomly extracting the feature information of the training set by the Bootstrap self-sampling method to generate multiple random sample sets;
[0036] A regression tree generation module for generating CART regression trees for multiple random sample sets based on the classification regression algorithm, and calculating the Gini coefficient of the random sample sets; recursively dividing the nodes according to the Gini coefficient until the Gini coefficient is less than the threshold or cannot be further divided, to generate multiple CART regression trees;
[0037] A prediction model generation module for taking the output mean of multiple CART regression trees as the prediction result to obtain an initial prediction model for the polygon of EMU wheels;
[0038] A model optimization module for inputting the test set into the initial prediction model, performing multiple random samplings on the feature information of the test set, and putting the sampled feature information back into the test set for the next random sampling; repeatedly constructing the model, selecting the features with the highest accuracy and putting them into the optimal feature subset, and repeating the random sampling for the remaining features until all features are traversed and then stopping to obtain the final prediction model;
[0039] A prediction output module for inputting the data to be processed into the final prediction model to obtain the prediction result.
[0040] In a third aspect, the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps corresponding to the method in the first aspect are implemented.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps corresponding to the method in the first aspect are implemented.
[0042] In summary, the technical solutions of the embodiments of the present application at least have the following advantages and beneficial effects:
[0043] The present invention utilizes the operation and maintenance historical data of the wheels of multiple unit trains, constructs training samples by selecting relatively strongly correlated characteristic indicators, and predicts the wheel polygon state using the random forest algorithm. First, key features such as wheel radius, the amplitude of the wheel set polygon before the last wheel turning, etc. are extracted from the operation and maintenance data as the feature vectors of the training samples, and they are randomly divided into a training set and a test set. In the model training stage, the Bootstrap self-sampling method is used to randomly extract multiple sample sets from the training set, construct CART regression trees for each sample set based on the classification regression algorithm, and recursively divide the nodes by calculating the Gini coefficient until the Gini coefficient is less than the set threshold or cannot be divided further, thereby generating multiple CART regression trees. These regression trees together constitute the initial prediction model, and the output result takes the average of the outputs of multiple trees to improve the stability and accuracy of the prediction. Subsequently, the test set is input into the initial prediction model, and the feature information of the test set is randomly sampled multiple times, with replacement each time. The model is repeatedly constructed, the features with the highest accuracy are selected and put into the optimal feature subset, and this process is continued until all features are traversed, and finally the optimized prediction model is obtained. Inputting the data to be processed into the final prediction model can obtain the prediction result of the wheel polygon. Utilizing the integration advantage of the random forest algorithm, through the collaborative work of multiple decision trees, the accuracy and reliability of the prediction of the wheel polygon state are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 FIG. is a flowchart of a method for predicting the polygon of multiple unit train wheels based on the random forest algorithm provided by the present invention;
[0045] Figure 2 FIG. is a schematic diagram of the training process of the random forest prediction model in the present invention;
[0046] Figure 3 FIG. is a structural diagram of a system for predicting the polygon of multiple unit train wheels based on the random forest algorithm provided by the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0048] The following refers to Figure 1 , a method for predicting the polygon of multiple unit train wheels based on the random forest algorithm proposed in the embodiments of the present application, includes:
[0049] S101, obtaining the operation and maintenance historical data of the wheels of multiple unit trains.
[0050] Specifically, the acquisition of operation and maintenance historical data involves the collection of information from multiple aspects. First, data needs to be extracted from the turning records of the wheel lathe. These data include the turning conditions of the wheels at different time points, such as turning time, turning reasons, and changes in wheel dimensions before and after turning. This information can reflect the wear conditions and maintenance history of the wheels during use. Second, data is obtained from the dynamic detection system for faults of EMU wheel sets. These data cover various detection results of the wheels during operation, including the amplitude of wheel polygons, types of tread defects, and defect positions. These detection data can provide the status information of the wheels during actual operation and help analyze the formation and development trends of wheel polygons. In addition, the daily mileage data of the EMU also needs to be collected, including the total mileage of the vehicle and the cumulative mileage after the last turning. These mileage data can reflect the usage intensity and operation time of the vehicle and provide an important reference basis for evaluating the wear degree of the wheels and predicting the development of polygons. When collecting these data, it is necessary to ensure the accuracy and integrity of the data to provide a reliable data basis for subsequent feature selection and model training.
[0051] S102, select characteristic indicators with correlation coefficients meeting the requirements from the operation and maintenance historical data as the training sample feature vectors, and randomly divide them into a training set and a test set.
[0052] The key to this step is to screen out the characteristic indicators that have a significant impact on the prediction of wheel polygons through correlation analysis. Specifically, first, perform a correlation analysis on each characteristic in the operation and maintenance historical data, and calculate the correlation coefficient between each characteristic and the amplitude of the wheel polygon. The Pearson correlation analysis method is used to calculate the correlation between each characteristic and the amplitude of the high-order polygon based on a preset formula to quantitatively reflect the linear relationship between the two sets of data. The value range is between -1 and 1. The closer the absolute value is to 1, the stronger the correlation; the closer it is to 0, the weaker the correlation. The preset formula is:
[0053]
[0054] where ρ xy represents the Pearson correlation coefficient between variable x and variable y, x i is the x value of the i-th sample, is the sample mean of x, y i is the y value of the i-th sample, is the sample mean of y, and n is the number of samples.
[0055] Through correlation analysis, it is possible to determine which features have an important impact on the prediction of wheel polygons. For example, feature indicators such as wheel radius, the amplitude of the wheel set polygon before the last wheel turning, the machining deviation of the flange thickness, the machining deviation of the wheel diameter, and mileage usually have a high correlation with the amplitude of the wheel polygon. These feature indicators can reflect the wear condition, maintenance history, and operating status of the wheel during use, thus providing important information for the prediction of wheel polygons.
[0056] Specifically, when the absolute value of ρ xy ranges from [0.8, 1], variable x and variable y are extremely strongly correlated; when the range is [0.6, 0.8), variable x and variable y are strongly correlated; when the range is [0.4, 0.6), variable x and variable y are moderately correlated; when the range is [0.2, 0.4), variable x and variable y are weakly correlated; when the range is [0, 0.2), variable x and variable y are extremely weakly correlated or uncorrelated. It is calculated that the absolute values of the correlations between the amplitude of the wheel set polygon before the last wheel turning and the wheel radius and the amplitude of the polygon before the current wheel turning are both above 0.4, showing a moderate correlation, and there is a weak correlation between the machining deviation of the flange thickness and the machining deviation of the wheel diameter and the amplitude of the polygon before the current wheel turning. The feature indicators with relatively large correlation coefficients are selected as the feature vectors of the training samples, and training data is randomly sampled to obtain the training set.
[0057] After selecting the feature indicators whose correlation coefficients meet the requirements, these feature indicators need to be used as the feature vectors of the training samples. The construction of feature vectors is the basis for the training of machine learning models, which directly affects the prediction performance of the models. To ensure the generalization ability and prediction accuracy of the models, the feature vectors need to be randomly divided into a training set and a test set. Usually, the training set is used for model training, while the test set is used for model validation and evaluation.
[0058] S103, Using the Bootstrap self-sampling method, randomly extract the feature information of the training set to generate multiple random sample sets.
[0059] Specifically, by randomly sampling the training set with replacement, multiple random sample sets of the same scale as the original training set are generated. The samples in each random sample set are randomly selected from the training set, and the same sample may appear in multiple random sample sets or may not appear in some random sample sets. This method can effectively increase the diversity of samples and reduce the risk of model overfitting.
[0060] In the process of generating a random sample set, first determine the number of samples and the feature dimension of the training set. Then, through a random number generator, randomly select samples from the training set, select one sample each time, and put it back into the training set to ensure that the sample may still be selected in the next selection. Repeat this process until a random sample set with the same scale as the training set is generated. By repeating the above process multiple times, multiple random sample sets can be generated, and these sample sets will be used for subsequent model training.
[0061] S104, as Figure 2 shown, generate CART regression trees for multiple random sample sets based on the classification regression algorithm, and calculate the Gini coefficient of the random sample sets; recursively divide the nodes according to the Gini coefficient until the Gini coefficient is less than the threshold or cannot be further divided, generating multiple CART regression trees.
[0062] Specifically, use the classification regression algorithm to process multiple random sample sets generated by the Bootstrap resampling method before. The classification regression algorithm is a commonly used machine learning algorithm that can perform classification or regression prediction according to the characteristics of the data. Here, mainly use its regression ability to construct a regression tree, that is, a CART regression tree.
[0063] The construction process of the CART regression tree is based on feature selection and data division. During the construction process, calculate the Gini coefficient of each node. The Gini coefficient is an index to measure the purity of a data set. The smaller its value, the higher the purity of the data set, that is, the higher the degree to which the samples in the data set belong to the same category. In the construction of the regression tree, the Gini coefficient is used to measure the impurity of the node. The lower the impurity, the smaller the difference in the target variable (here are the wheel polygon related features) of the samples contained in the node. For example: in the random sample set, define the number of samples as |X| and the number of wheel polygons as |Y|, and the expression of its Gini coefficient Gini(X) is:
[0064]
[0065] According to the Gini coefficient, the algorithm will recursively divide the nodes. Starting from the root node, by selecting the optimal feature and division point, divide the data set into two child nodes, so that the Gini coefficient of each child node is as small as possible, that is, the difference in the target variable of the samples in the child node is as small as possible. This process will be repeated continuously until the stop condition is met. The stop condition is usually that the Gini coefficient is less than the preset threshold, or cannot be further divided (for example, the number of samples in the node reaches the minimum value, or there are no available features for division, etc.). The specific steps are as follows:
[0066] S401. In multiple random sample sets where the current node of the regression tree is located, if the Gini coefficients of the multiple random sample sets are less than the Gini coefficient threshold, the classification regression algorithm returns the sub-decision tree and stops recursion; if the Gini coefficients of the multiple random sample sets are greater than or equal to the Gini coefficient threshold, calculate the Gini coefficient of each feature information in the random sample set where the current node is located.
[0067] This step aims to determine whether to stop further partitioning the random sample set of the current node to generate a sub-decision tree. Specifically, in the multiple random sample sets corresponding to the current node of the regression tree, the Gini coefficient of each random sample set needs to be judged. If the Gini coefficients of the multiple random sample sets are all less than the pre-set Gini coefficient threshold, it indicates that the differences in the data related to the wheel polygon features contained in the current node are small enough, that is, the purity of the data is high. At this time, the classification regression algorithm will return the sub-decision tree and stop the recursive partitioning of this node. The Gini coefficient threshold is a key parameter, which is set according to the actual prediction requirements and data characteristics to control the growth degree of the regression tree and prevent overfitting or underfitting. However, if the Gini coefficients of the multiple random sample sets are greater than or equal to the Gini coefficient threshold, it means that there are still large differences in the data of the current node in the target variable, and there is still room and necessity for further partitioning. At this time, it is necessary to calculate the Gini coefficient of each feature information in the random sample set where the current node is located, so as to screen out the optimal feature information for node partitioning operations in the follow-up, so as to continue to construct the regression tree, further explore the potential laws and relationships in the data, and improve the accuracy of the prediction of the EMU wheel polygon. Through this recursive partitioning strategy based on the Gini coefficient, a CART regression tree with a reasonable structure and good prediction performance can be effectively constructed, laying a solid foundation for finally generating an accurate wheel polygon prediction model.
[0068] S402. Screen out the feature information corresponding to the minimum Gini coefficient as the optimal feature information, divide the random sample set corresponding to the optimal feature information into a first feature set and a second feature set, and use the first feature set and the second feature set as the left and right child nodes of the current node respectively.
[0069] Specifically, first calculate the Gini coefficient of each feature information in the random sample set where the current node is located. During the calculation process, each feature will be evaluated according to formula (2). By comparing the Gini coefficient sizes of different features, screen out the feature information corresponding to the minimum Gini coefficient as the optimal feature information. This is because the minimum Gini coefficient means that when this feature partitions the data set, it can maximize the purity of the child nodes, making the differences in the samples within the child nodes in the target variable as small as possible, so as to better explore the potential laws and relationships in the data and improve the prediction accuracy of the model.
[0070] After determining the optimal feature information, the corresponding random sample set of the optimal feature information needs to be divided into a first feature set and a second feature set. This division process is based on the value range of the optimal feature or different manifestation forms of the feature. For example, if the optimal feature is the wheel radius, it can be divided into two subsets according to the specific value range of the wheel radius. One subset contains wheel samples with smaller radii, and the other subset contains wheel samples with larger radii. The divided first feature set and second feature set serve as the left child node and right child node of the current node respectively, laying the foundation for subsequent recursive division and the construction of the regression tree. In this way, the data set can be further refined, making the samples within each child node more similar, which helps to more accurately predict the relevant features of the wheel polygon and provides strong support for finally generating an accurate wheel polygon prediction model.
[0071] S403. Calculate the Gini coefficients of the left child node and the right child node. If the Gini coefficients of the left child node and the right child node are greater than or equal to the Gini coefficient threshold, continue to recursively generate the left child node and the right child node of the current node and calculate the corresponding Gini coefficients; if the Gini coefficients of the left child node and the right child node are less than the Gini coefficient threshold, the classification regression algorithm returns the sub-decision tree and stops recursion to generate a CART regression tree.
[0072] Specifically, through statistical analysis of the sample data included in the left child node and the right child node, the respective Gini coefficients are calculated according to formula (2). After calculating the Gini coefficients of the left child node and the right child node, it is necessary to judge these two values. If the Gini coefficients of the left child node and the right child node are both greater than or equal to the preset Gini coefficient threshold, this indicates that the data included in these two child nodes still have relatively large differences in the relevant features of the wheel polygon, and there is still room and necessity for further division. At this time, the classification regression algorithm will continue to recursively divide the left child node and the right child node respectively, repeating the processes of steps S401 and S402, that is, calculating the Gini coefficient of each feature information, screening out the optimal feature information, and dividing the data into new child nodes according to the optimal feature information to continue constructing the regression tree.
[0073] However, if the Gini coefficients of the left child node and the right child node are both less than the Gini coefficient threshold, this means that the differences in the relevant features of the wheel polygon of the data included in these two child nodes are small enough, the purity of the data is high, and the requirements for model construction are met. At this time, the classification regression algorithm will return the sub-decision tree and stop the recursive division of this node. Through this recursive division strategy based on the Gini coefficient, a CART regression tree with a reasonable structure and good prediction performance can be effectively constructed.
[0074] S105. Take the average of the outputs of multiple CART regression trees as the prediction result to obtain the initial prediction model for the polygon of the EMU wheel.
[0075] Specifically, for each sample to be predicted, each CART regression tree outputs a prediction value, and these prediction values reflect different characteristics or states of the wheel polygon. Since each regression tree is constructed based on a different random sample set, there may be certain differences in their prediction results. To obtain a more stable and reliable prediction result, the method of taking the average is used for integration.
[0076] First, for each sample in the test set, input it into each CART regression tree to obtain the prediction value of each tree for this sample. These prediction values represent the judgments of different decision trees on the state of the wheel polygon of this sample. Since a certain degree of randomness is introduced in the construction process of each tree, these prediction values may have certain fluctuations.
[0077] Next, calculate the average of the prediction values of all CART regression trees for the same sample. That is, sum the output values of each regression tree and then divide by the total number of regression trees to obtain the final prediction result of this sample. In this way, the possible overfitting or bias problems of a single regression tree can be effectively reduced, and the stability and accuracy of the prediction result can be improved.
[0078] By taking the average of the outputs of multiple CART regression trees, the information of each tree can be fully utilized, the influence of random errors can be reduced, and thus the generalization ability and reliability of the prediction model can be improved. This method can not only effectively integrate the prediction results of multiple regression trees, but also provide a relatively accurate initial prediction model for subsequent model evaluation and optimization.
[0079] S106. Input the test set into the initial prediction model, perform multiple random samplings on the feature information of the test set, and put the sampled feature information back into the test set for the next random sampling; repeatedly construct the model, select the feature with the highest accuracy and put it into the optimal feature subset, and repeat the random sampling for the remaining features until all features are traversed and then stop to obtain the final prediction model.
[0080] Specifically, through the method of random sampling with replacement, part of the samples are drawn from the test set to form multiple sampled sample sets. After each random sampling, the sampled feature information is put back into the test set to ensure that all samples are still likely to be selected during the next random sampling. This method can increase the diversity of samples, reduce the risk of model overfitting, and improve the generalization ability of the model.
[0081] Next, based on these sampled sample sets, the model is repeatedly constructed. Each time the model is constructed, different combinations of features are used to evaluate the impact of each feature on the model's prediction performance. Specifically, by calculating the accuracy rate of each feature in the model, the feature with the highest accuracy rate is selected and placed in the optimal feature subset. The calculation formula for the accuracy rate A is:
[0082]
[0083] Among them, TP is predicted as the positive class and is actually the positive class; TN is predicted as the negative class and is actually the negative class; FP is predicted as the positive class and is actually the negative class; FN is predicted as the negative class and is actually the positive class.
[0084] After screening out the optimal features, the above-mentioned random sampling and model construction processes are repeated for the remaining features until all features are traversed. This ensures that all features are fully evaluated and the features that contribute the most to the model's prediction performance are selected. Finally, based on the optimal feature subset, the final prediction model is constructed. This model has higher accuracy and generalization ability and can more effectively predict the state of the polygon of the EMU wheels, providing an important reference for the maintenance and overhaul of the EMU.
[0085] S107, input the data to be processed into the final prediction model to obtain the prediction result.
[0086] After obtaining the prediction result, it also includes:
[0087] Display the prediction result in the form of a chart. The chart includes a trend chart of the predicted amplitude of the wheel polygon over time; compare the prediction result with a preset safety threshold and mark the prediction results that exceed the safety threshold in the chart.
[0088] Specifically, after obtaining the prediction result, it is first displayed in the form of a chart. The key chart is the trend chart of the predicted amplitude of the wheel polygon over time. Through this visual way, the change situation of the wheel polygon amplitude in a future period of time can be intuitively presented. For example, if the trend chart shows that the wheel polygon amplitude shows a gradually increasing trend, this may mean that the wear of the wheel is intensifying and attention needs to be paid and corresponding maintenance measures need to be taken in a timely manner.
[0089] In addition to visually displaying the prediction results, the prediction results are compared with preset safety thresholds, and the prediction results exceeding the safety thresholds are marked in the chart. The specific implementation of this process is as follows: The system will automatically compare the predicted wheel polygon amplitudes with the preset safety thresholds. Once it is found that the predicted amplitude at a certain time point exceeds the safety threshold, the system will significantly mark that time point in the trend chart, for example, highlighting it with a red mark or other prominent means to remind the maintenance personnel to pay attention. This comparison mechanism can timely detect potential safety hazards and ensure the safe operation of the EMU. Through this comparison and marking mechanism, not only the readability and practicality of the prediction results are improved, but also an important reference basis is provided for the maintenance and overhaul of the EMU, which helps to enhance the safety and reliability of the EMU operation.
[0090] Based on the same inventive concept, as Figure 3 shown, the present invention provides a prediction system for EMU wheel polygons based on a random forest algorithm, including:
[0091] A data acquisition module 201 for acquiring the operation and maintenance historical data of EMU wheels;
[0092] An influencing factor division module 202 for selecting characteristic indicators with correlation coefficients meeting the requirements from the operation and maintenance historical data as training sample feature vectors and randomly dividing them into a training set and a test set;
[0093] A sample generation module 203 for randomly extracting the feature information of the training set by the Bootstrap self-sampling method to generate multiple random sample sets;
[0094] A regression tree generation module 204 for generating CART regression trees for multiple random sample sets based on a classification regression algorithm and calculating the Gini coefficient of the random sample sets; recursively dividing nodes according to the Gini coefficient until the Gini coefficient is less than the threshold or cannot be further divided to generate multiple CART regression trees;
[0095] A prediction model generation module 205 for taking the output mean of multiple CART regression trees as the prediction result to obtain an initial prediction model for EMU wheel polygons;
[0096] A model optimization module 206 for inputting the test set into the initial prediction model, performing multiple random samplings on the feature information of the test set, and putting the sampled feature information back into the test set for the next random sampling; repeatedly constructing models, selecting the features with the highest accuracy and putting them into the optimal feature subset, and repeating the random sampling for the remaining features until all features are traversed and then stopping to obtain the final prediction model;
[0097] A prediction output module 207 for inputting the data to be processed into the final prediction model to obtain the prediction result.
[0098] Based on the same inventive concept, the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, a prediction method for the polygon of the EMU wheel based on the random forest algorithm is implemented.
[0099] Based on the same inventive concept, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, a prediction method for the polygon of the EMU wheel based on the random forest algorithm is implemented.
[0100] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for predicting EMU wheel polygons based on random forest algorithm, characterized in that: include: Obtain the operation and maintenance history data of EMU wheels; From the operation and maintenance historical data, feature indicators with correlation coefficients that meet the requirements are selected as training sample feature vectors, and randomly divided into training sets and test sets; Using the Bootstrap self-service sampling method, the feature information of the training set is randomly extracted to generate multiple random sample sets; Generate CART regression trees of multiple random sample sets based on the classification regression algorithm, and calculate the Gini coefficient of the random sample set; recursively divide the nodes according to the Gini coefficient until the Gini coefficient is less than the threshold or cannot be further divided, and generate multiple CART regression trees; The output mean of multiple CART regression trees is taken as the prediction result to obtain the initial prediction model of EMU wheel polygons; Input the test set into the initial prediction model, perform multiple random sampling on the feature information of the test set, and put the sampled feature information back into the test set for the next random sampling; repeatedly build the model, select the feature with the highest accuracy and put it into the optimal feature subset, and repeat the random sampling of the remaining features until all features are traversed and stop, so as to obtain the final prediction model; The data to be processed is input into the final prediction model to obtain a prediction result.
2. The method for predicting EMU wheel polygons based on random forest algorithm according to claim 1, characterized in that: The step of recursively dividing nodes according to the Gini coefficient until the Gini coefficient is less than a threshold or cannot be further divided, and generating multiple CART regression trees specifically includes: In the multiple random sample sets where the current node of the regression tree is located, if the Gini coefficients of the multiple random sample sets are less than the Gini coefficient threshold, the classification regression algorithm returns to the sub-decision tree and stops recursion; if the Gini coefficients of the multiple random sample sets are greater than or equal to the Gini coefficient threshold, the Gini coefficient of each feature information in the random sample set where the current node is located is calculated; Select feature information corresponding to the minimum Gini coefficient as optimal feature information, divide the random sample set corresponding to the optimal feature information into a first feature set and a second feature set, and use the first feature set and the second feature set as the left child node and the right child node of the current node respectively; Calculate the Gini coefficients of the left child node and the right child node. If the Gini coefficients of the left child node and the right child node are greater than or equal to the Gini coefficient threshold, continue to recursively generate the left child node and the right child node of the current node, and calculate the corresponding Gini coefficients; if the Gini coefficients of the left child node and the right child node are less than the Gini coefficient threshold, the classification regression algorithm returns to the child decision tree and stops recursion to generate a CART regression tree.
3. The method for predicting EMU wheel polygons based on random forest algorithm according to claim 1, characterized in that: From the operation and maintenance historical data, feature indicators whose correlation coefficients meet the requirements are selected as training sample feature vectors, including: The Pearson correlation analysis method was used to calculate the correlation between each feature and the amplitude of the high-order polygon based on the preset formula to quantitatively reflect the linear relationship between the two sets of data.
4. The method for predicting EMU wheel polygons based on random forest algorithm according to claim 3 is characterized in that: The preset formula includes: Among them, ρ xy represents the Pearson correlation coefficient between variables x and y, x i is the x value of the ith sample, is the sample mean of x, y i is the y value of the i-th sample, is the sample mean of y, and n is the number of samples.
5. The method for predicting EMU wheel polygons based on random forest algorithm according to claim 1, characterized in that: In the step of selecting the feature with the highest accuracy and putting it into the optimal feature subset, the accuracy calculation formula is: Among them, TP means the prediction is positive and the actual is positive; TN means the prediction is negative and the actual is negative; FP means the prediction is positive and the actual is negative; FN means the prediction is negative and the actual is positive.
6. The method for predicting EMU wheel polygons based on random forest algorithm according to claim 1, characterized in that: The characteristic indicators include wheel diameter, wheel rim thickness processing deviation, predicted maximum value of wheel polygon before the last turning, wheel diameter processing deviation, total vehicle mileage, cumulative mileage after the last turning and repair, carriage type, axle number, season and the presence or absence of wheel polygon failure.
7. The method for predicting EMU wheel polygons based on random forest algorithm according to claim 1, characterized in that: After inputting the data to be processed into the final prediction model to obtain the prediction result, the method further includes: The prediction result is presented in a graph form, wherein the graph includes a graph showing a trend of the predicted wheel polygon amplitude over time; The prediction results are compared with a preset safety threshold, and the prediction results exceeding the safety threshold are marked in the chart.
8. A wheel polygon prediction system for EMU based on random forest algorithm, characterized in that: include: A data acquisition module is used to obtain the operation and maintenance history data of the EMU wheels; The influencing factor division module is used to select characteristic indicators with correlation coefficients that meet the requirements from the operation and maintenance historical data as training sample feature vectors, and randomly divide them into training sets and test sets; A sample generation module is used to randomly extract feature information of the training set to generate multiple random sample sets using the Bootstrap self-service sampling method; The regression tree generation module is used to generate CART regression trees of multiple random sample sets based on the classification regression algorithm and calculate the Gini coefficient of the random sample set; recursively divide the nodes according to the Gini coefficient until the Gini coefficient is less than the threshold or cannot be further divided, and generate multiple CART regression trees; The prediction model generation module is used to take the output mean of multiple CART regression trees as the prediction result to obtain the initial prediction model of the EMU wheel polygon; The model optimization module is used to input the test set into the initial prediction model, perform multiple random sampling on the feature information of the test set, and put the sampled feature information back into the test set for the next random sampling; repeatedly build the model, select the feature with the highest accuracy and put it into the optimal feature subset, and repeat the random sampling of the remaining features until all features are traversed and stop, so as to obtain the final prediction model; The prediction output module is used to input the data to be processed into the final prediction model to obtain the prediction result.
9. An electronic device, characterized in that: The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for predicting EMU wheel polygons based on a random forest algorithm as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements a method for predicting EMU wheel polygons based on a random forest algorithm as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Methods and equipment for training AI models for wheel polygonal wear waveform regression prediction
CN113947130B
Cited By
Motor train unit intelligent operation and maintenance evaluation method and system based on credibility engineering
CN120471611A
A method and system for evaluating intelligent operation and maintenance of EMUs based on reliability engineering
CN120471611B