Automobile preventive maintenance intelligent optimization system based on random forest algorithm

By constructing a multi-decision tree model using the random forest algorithm and combining it with vehicle sensor and GPS data, the maintenance cycle is dynamically adjusted, which solves the shortcomings of traditional car maintenance methods, realizes accurate and personalized maintenance suggestions, and improves the system's predictive ability and user experience.

CN120875837APending Publication Date: 2025-10-31YUNNAN CAR TRAFFICKING NETWORK TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510991582.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-22
Filing Date
2025-07-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional car maintenance methods rely on fixed times or mileages, which cannot take into account individual vehicle differences, leading to over- or under-maintenance issues. Existing intelligent maintenance systems are inadequate in terms of data processing capabilities, accuracy, and user experience.

Method used

An intelligent optimization system for preventive vehicle maintenance based on the random forest algorithm is adopted. By collecting data from onboard sensors and GPS, and combining multi-dimensional data preprocessing and model training, a multi-decision tree model is constructed to dynamically adjust the maintenance cycle and provide personalized suggestions.

Benefits of technology

It improves the accuracy and robustness of maintenance cycle prediction, reduces maintenance costs, enhances user experience, and adapts to changes in different vehicles and driving environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875837A_ABST
    Figure CN120875837A_ABST
Patent Text Reader

Abstract

A vehicle preventive maintenance intelligent optimization system based on a random forest algorithm comprises a data acquisition module, a data preprocessing module, a model training and prediction module, a maintenance period optimization module and a user interface module, and the data acquisition module acquires real-time data of a vehicle through a vehicle-mounted sensor and a GPS device; the data preprocessing module is used for cleaning, denoising and standardizing data, the model training and predicting module is used for constructing an integrated model of multiple decision-making trees by utilizing a random forest algorithm based on the preprocessed data, and the maintenance period optimizing module is used for dynamically adjusting a recommendation period in combination with historical maintenance information of a user and a prediction result. And the suggestion period and the related analysis result are displayed to the user through the user interface module, so that the service life of the vehicle is prolonged, the maintenance cost is reduced, and the system has the advantages of being efficient, accurate and capable of achieving real-time feedback and is suitable for the customized maintenance requirement of the intelligent vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent services and data processing, specifically to an intelligent optimization system for preventive maintenance of automobiles based on the random forest algorithm. Background Technology

[0002] With the development of the automotive industry, automobiles have become an indispensable means of transportation in modern life. Routine maintenance and repair are crucial for maintaining normal operation, extending lifespan, and ensuring driving safety. However, traditional car maintenance methods typically rely on fixed times or mileage to determine maintenance cycles. For example, manufacturers or service providers often recommend maintenance every certain mileage (e.g., 5000 kilometers) or after a certain period (e.g., 6 months). While simple, this approach doesn't consider the actual differences in each car's usage, including driver habits, road conditions, climate, and the vehicle's condition. This "one-size-fits-all" maintenance method often leads to over-maintenance or under-maintenance. Over-maintenance increases the user's time and money costs, while under-maintenance can lead to decreased vehicle performance and even safety hazards. Therefore, how to intelligently and personalizedly determine car maintenance cycles has become an important technical challenge.

[0003] With the rapid development of information technology and big data analytics, data-driven intelligent prediction technologies have been widely applied in various fields. Especially driven by the Internet of Things (IoT) and Vehicle-to-Everything (V2X), modern automobiles are now able to collect and transmit massive amounts of sensor data in real time, such as engine temperature, fuel consumption, tire pressure, and driving speed. This real-time data provides the technological possibility for dynamically assessing vehicle status and predicting vehicle maintenance needs. However, to extract meaningful features from large amounts of data and analyze specific vehicle maintenance needs, machine learning algorithms are still needed. In recent years, machine learning algorithms, especially random forests and neural network data-driven models, have demonstrated high accuracy and robustness in prediction and classification, and have been gradually introduced into intelligent vehicle management and health management systems. Random forest algorithms, as an ensemble learning-based algorithm, have good generalization ability and anti-overfitting ability, and can effectively handle high-dimensional data and complex nonlinear relationships. They are particularly suitable for multi-factor comprehensive assessment of vehicle status and maintenance cycle prediction.

[0004] Current intelligent maintenance systems have gradually acquired certain predictive capabilities, but several key challenges remain. First, the diversity and real-time nature of data necessitate powerful data processing capabilities. Significant differences in driving habits and environmental conditions among different vehicle models and users result in highly diverse data sources, requiring the system to quickly analyze and process this heterogeneous data. Second, accuracy and robustness are core requirements for intelligent maintenance cycle optimization. Traditional simple prediction methods struggle to meet the needs of practical applications. Extracting key features from complex, multi-dimensional data and continuously optimizing the prediction model in the face of dynamic data changes is a current research focus. Finally, user experience and acceptance must also be considered. Existing intelligent maintenance systems often feature cluttered user interfaces, making it difficult for users to clearly understand the maintenance recommendations. Therefore, to achieve accurate maintenance cycle prediction, personalized maintenance recommendations, and an easy-to-use user interface, it is necessary to construct an intelligent maintenance cycle optimization system that can adapt to diverse data inputs, possesses high predictive performance, and meets user experience requirements. Summary of the Invention

[0005] To address the aforementioned problems, this invention aims to provide an intelligent optimization system for preventive automotive maintenance based on the random forest algorithm, thereby resolving the issues raised in the background section.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] This invention provides an intelligent optimization system for preventive vehicle maintenance based on the random forest algorithm. The system includes a data acquisition module, a data preprocessing module, a model training and prediction module, a maintenance cycle optimization module, and a user interface module. The data acquisition module obtains real-time vehicle data through onboard sensors, GPS devices, etc. The data preprocessing module cleans, denoises, and standardizes the data. The model training and prediction module uses the preprocessed data to construct an ensemble model with multiple decision trees using the random forest algorithm, thereby improving prediction accuracy and robustness. The system inputs new data into the model and predicts suitable maintenance cycles based on the driver's real-time driving behavior and vehicle health. The maintenance cycle optimization module dynamically adjusts the recommended cycle based on the user's historical maintenance information and prediction results. The user interface module displays the suggested cycle and related analysis results to the user, enabling them to perform maintenance at appropriate times, thereby extending vehicle lifespan and reducing maintenance costs.

[0008] Furthermore, the data acquisition module consists of an onboard sensor integration unit, a Global Positioning System (GPS) device, an external data interface, and a data storage unit. The onboard sensor integration unit is the data acquisition module, which collects real-time internal operating data of the vehicle through engine sensors, fuel consumption sensors, temperature sensors, tire pressure sensors, brake wear sensors, battery power sensors, and speed sensors. Each sensor corresponds to a different collection task. The GPS device collects real-time vehicle location information, driving route, and driving speed data through the Global Positioning System. The system can analyze the vehicle's driving conditions. The data collected by the GPS device will be input into the system along with the sensor data to analyze the impact of environmental factors on the vehicle, thereby more accurately predicting maintenance cycles. The data acquisition module also includes an external data interface for receiving information from external data sources. Vehicle operation is not only affected by its own state but also significantly affected by the external environment. Through the external data interface, the system can receive weather data, real-time traffic flow information, and road construction and road condition information. This data is obtained through cooperation with weather data providers, traffic information platforms, and road management systems and is updated in real time. The introduction of external data can not only supplement environmental factors that are difficult for onboard sensors to capture but also enable the system to understand the vehicle's operating environment.

[0009] Furthermore, the collected data is stored in the data storage unit for subsequent analysis and prediction. The data storage unit has efficient storage and management functions, and can perform hierarchical management of real-time and historical data. Real-time data is used for current maintenance cycle prediction, while historical data is used for model training and long-term trend analysis. During storage, the data will be grouped and labeled according to time, data type, and source device dimensions to facilitate subsequent rapid retrieval and analysis.

[0010] Furthermore, the data preprocessing module performs cleaning, denoising, and standardization preprocessing operations on the multi-source data collected by the data acquisition module. For numerical extreme values, it uses three times the standard deviation for detection; for outliers, it replaces them with the average value; for missing data, the system fills in the missing data through interpolation to avoid the missing data causing bias in model training. Then, the data is denoised by optimizing the moving average method and averaging continuous data points to smooth the data, thereby reducing abnormal fluctuations. The denoised data is then standardized and normalized. Standardization adjusts the data to a standard normal distribution with the same mean and variance, while normalization scales the data proportionally to the [0,1] interval.

[0011] Furthermore, the model training and prediction module trains the model using the random forest algorithm, as detailed below:

[0012] Define the training set as x i =[xi1 ,x i2 ,...x im ], where D is the training set, x i Let y be the feature vector of the i-th group, and N represent the number of samples. i For the target maintenance cycle, x i1 ,x i2 ,...x im Let N' be the first feature vector of the i-th sample, N' be the second feature vector of the i-th sample, and N' be the m-th feature vector of the i-th sample. N' samples are drawn from the training set with replacement, where N' ≤ N, and K distinct subsets are generated, each with length N'. When constructing each decision tree, FLOOR(sqrtm) features are randomly selected from the m features for splitting, where FLOOR is the downward fetching function. For each node of the k-th decision tree, the optimal splitting condition for all possible split points is calculated. For the current sample set of the k-th decision tree... Information gain (Gain) is expressed as follows:

[0013]

[0014] Where, p n The information gain (Gini index) represents the proportion of the nth sample in the sample set D. It measures the improvement in data purity before and after the split, calculated based on the information entropy using the Shannon formula. index ,have:

[0015]

[0016] The Gini index, calculated based on a probability distribution, measures data impurity. The mean squared error (MSE) is defined to measure the deviation between predicted and actual values.

[0017]

[0018] in, for Maintenance cycle, For computation in the random forest algorithm The maintenance cycle, the greater the information gain, the greater the improvement in the purity of the feature partitioning nodes. If we define maximization as forward optimization and minimization as reverse optimization, in order to unify the evaluation index, we reverse optimize the information gain as... Constructing indicator structure For K decision trees, there is a K*3 dimensional standardized evaluation matrix A*:

[0019]

[0020] in, This refers to the information gain inverse element in the standardized evaluation matrix of the first decision tree. This refers to the information gain inverse element in the standardized evaluation matrix of the second decision tree. Gini represents the information gain inverse element in the standardized evaluation matrix of the Kth decision tree. index (1) represents the Gini index element in the standardized evaluation matrix of the first decision tree. index (2) represents the Gini index element in the standardized evaluation matrix of the second decision tree. index (K) represents the Gini index element in the standardized evaluation matrix of the Kth decision tree. The mean squared error element in the standardized evaluation matrix of the first decision tree. The mean squared error element in the standardized evaluation matrix of the second decision tree. Let A* be the mean squared error element in the standardized evaluation matrix of the Kth decision tree, and define the maximum value vector of the standardized matrix as A*. + :

[0021]

[0022] Define the normalized matrix minimum vector A* -

[0023]

[0024] Then the evaluation metric SCORE of the k-th decision tree k The calculation is as follows:

[0025]

[0026] By calculating SCORE k This is used to standardize Gain, Gini_index, and MSE to eliminate the influence of dimensions. The relative distance between each index and the optimal / worst solution is calculated using Euclidean distance to quantify the overall performance of the decision tree. The final score reflects the overall effectiveness of the decision tree and is used to guide model optimization.

[0027]

[0028] Furthermore, prediction performance is improved by constructing a node loss function, as shown below:

[0029]

[0030] Where Loss(k) is the loss function of the k-th decision tree. To minimize the weighted loss of the entire decision tree, the loss increment ΔLoss is calculated. k :

[0031]

[0032] Among them, D k< D k The left child subset of the decision tree under the splitting rule π, D k> D k The right child subset of the decision tree under the splitting rule π has {D} k< ∩D k>}=D k Loss(k<) is D k< The loss function, Loss(k>), is D. k> The loss function, to prevent overfitting, introduces a regularization term and is as follows:

[0033]

[0034] Where Loss is the loss function of the random forest, α is the model complexity penalty coefficient, which constrains the number of samples contained in a node to prevent the tree structure from expanding indefinitely, and its value ranges from [0.01, 0.1]. If α is large, the node complexity penalty accounts for a higher proportion, and the system tends to remove child nodes containing fewer samples and prioritize retaining important nodes; if α is small, the penalty for node complexity is weakened, and the system tends to retain more child nodes, allowing for a more complex tree structure. β is the feature selection penalty coefficient, reflecting the importance distribution of nodes to different features, and its value ranges from [0.01, 0.1]. This is achieved through feature variance... To measure this, increasing β increases the penalty for feature variance, prompting the system to be more inclined to remove split nodes associated with low-importance features. Decreasing β weakens the penalty for feature selection, allowing the random forest to utilize more features for complex splits. m The feature weights have a value range of [0,1] and represent the importance of the feature to the model's prediction.

[0035] Furthermore, the maintenance cycle optimization module optimizes the maintenance cycle predicted by the model to provide accurate and personalized maintenance suggestions. By dynamically adjusting and optimizing the initially predicted maintenance cycle, the maintenance suggestions are made to meet actual needs. The maintenance cycle optimization module includes a historical data analysis submodule, a dynamic adjustment submodule, and a personalized suggestion generation submodule.

[0036] Furthermore, the historical data analysis submodule identifies factors affecting maintenance cycles based on the vehicle's historical operating data and maintenance records, thereby making personalized adjustments to the prediction results; the dynamic adjustment submodule dynamically adjusts the predicted maintenance cycle based on the vehicle's real-time data and current status; while making dynamic adjustments, the personalized suggestion generation submodule generates specific maintenance suggestions based on all information for user reference, integrating the results from the prediction model, rule base, historical analysis, and dynamic adjustment to provide users with maintenance suggestions that better meet their actual needs.

[0037] Furthermore, the user interface module provides users with vehicle status monitoring, maintenance cycle suggestions, and historical data query functions through a visual interface and interactive features, thereby enhancing user experience and system usability.

[0038] Furthermore, the user interface module features data display capabilities, transforming the system's analyzed and predicted data into easily understandable visual information. It integrates real-time vehicle operating data, maintenance cycle recommendations, and historical maintenance records, displaying them through charts, graphical dashboards, and graphs. For maintenance cycle recommendations, the system displays specific recommended maintenance times or mileages, lists the components and systems requiring maintenance, and provides explanations for each recommendation. If a component is nearing its maintenance cycle, a notification will appear on the interface, reminding the user to pay attention.

[0039] The beneficial effects of this invention are as follows: This invention integrates multi-dimensional data sources, including vehicle sensor data, historical maintenance records, driving environment data, and driving habit data, to fully explore factors affecting maintenance cycles. Traditional maintenance models typically rely on single mileage or time indicators, failing to accurately reflect the vehicle's true wear and tear and maintenance needs. This system, through comprehensive analysis of multiple data sources, can more accurately determine the actual health status of various vehicle components, significantly improving the accuracy of maintenance cycle prediction. This invention introduces a random forest algorithm model. The random forest algorithm is an ensemble learning-based algorithm that uses multiple decision trees to achieve voting decisions on prediction results. This invention utilizes the random forest algorithm for intelligent prediction of maintenance cycles, effectively addressing the shortcomings of traditional models in terms of data nonlinearity and high dimensionality. The random forest algorithm exhibits high robustness in handling complex, nonlinear data and can continuously adjust prediction results based on dynamic changes in input data. Unlike traditional random forest algorithms, this invention extends this by adding three indicators: Gain, Gini, and Gini. index The three metrics for evaluating random forests—Gain, Gini coefficient, and MSE—were standardized and aligned. indexThe three metrics, MSE, and β, contain the feature information of each decision tree in the random forest and are quantified into a K-row, 3-column matrix, greatly reducing the computational load of the system participating in the random forest. The deviation of each metric is calculated based on Euclidean distance to synthesize the missile score of each decision tree. This score information is associated with the loss function, which includes a regularization term to prevent overfitting. Unlike traditional regularization, this invention introduces a model complexity penalty coefficient and a feature selection penalty coefficient. α is the model complexity penalty coefficient, which constrains the number of samples a node contains, preventing unlimited expansion of the tree structure. If α is large, the node complexity penalty is high, and the system tends to remove child nodes with fewer samples, prioritizing the retention of important nodes. If α is small, the penalty for node complexity is weakened, and the system tends to retain more child nodes, allowing for a more complex tree structure. β is the feature selection penalty coefficient, reflecting the importance distribution of nodes for different features, and is calculated using feature variance. To measure this, increasing β increases the penalty for feature variance, prompting the system to be more inclined to remove split nodes associated with low-importance features. Conversely, decreasing β weakens the penalty for feature selection, allowing the random forest to utilize more features for complex splits. Furthermore, the random forest algorithm has strong resistance to outlier data, reducing the risk of overfitting in a single model. The introduction of this predictive model enables the system to maintain high prediction accuracy and applicability under different vehicles, driving environments, and driving habits, providing users with more precise maintenance cycle recommendations. This invention's intelligent optimization system for vehicle maintenance cycles based on the random forest algorithm has significant innovation and practicality. Through multiple technological innovations, including multi-dimensional data sources, intelligent predictive models, dynamic optimization, personalized recommendations, and remote monitoring, it provides accurate maintenance cycle recommendations for vehicles, optimizing user experience and maintenance costs. It meets the intelligent and personalized needs of modern vehicle management and has broad application prospects in future intelligent vehicle maintenance management. It not only improves vehicle safety and reliability but also provides car owners with more convenient and economical maintenance solutions, possessing significant market value and social benefits. Attached Figure Description

[0040] The invention will be further illustrated with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the invention. For those skilled in the art, other drawings can be obtained based on the following drawings without any creative effort.

[0041] Figure 1 This is a schematic diagram of the structure of the present invention. Detailed Implementation

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] Please see Figure 1 The present invention will be further described in conjunction with the following examples.

[0044] See Figure 1 This invention aims to provide an intelligent optimization system for preventive vehicle maintenance based on the random forest algorithm. The system includes a data acquisition module, a data preprocessing module, a model training and prediction module, a maintenance cycle optimization module, and a user interface module. The data acquisition module obtains real-time vehicle data through onboard sensors, GPS devices, etc. The data preprocessing module cleans, denoises, and standardizes the data. The model training and prediction module uses the preprocessed data to construct an ensemble model with multiple decision trees using the random forest algorithm, thereby improving prediction accuracy and robustness. The system inputs new data into the model and predicts suitable maintenance cycles based on the driver's real-time driving behavior and vehicle health. The maintenance cycle optimization module dynamically adjusts the recommended cycle based on the user's historical maintenance information and prediction results. The user interface module displays the suggested cycle and related analysis results to the user, enabling them to perform maintenance at appropriate times, thereby extending vehicle lifespan and reducing maintenance costs.

[0045] Specifically, the data acquisition module consists of an on-board sensor integration unit, a global positioning system (GPS) device, an external data interface, and a data storage unit. The on-board sensor integration unit is the data acquisition module, which collects the vehicle's internal operating data in real time through engine sensors, fuel consumption sensors, temperature sensors, tire pressure sensors, brake wear sensors, battery power sensors, and speed sensors. Each sensor corresponds to a different data collection task.

[0046] Specifically, engine sensors collect engine temperature, speed, and workload parameters to help monitor engine operating status; fuel consumption sensors record fuel consumption per unit time, providing a basis for analyzing vehicle energy efficiency; temperature sensors measure the temperature of the vehicle and its key components, promptly identifying abnormal temperatures; tire pressure sensors collect pressure data from each tire to enable the system to analyze tire wear and safety; brake wear sensors monitor the wear of the braking system, ensuring its effectiveness and safety; and battery charge and speed sensors record changes in battery charge and speed, respectively. Through the coordinated work of these multiple sensors, the on-board sensor integration unit can comprehensively and meticulously reflect the vehicle's health and operating status.

[0047] Specifically, GPS devices use the Global Positioning System to collect real-time vehicle location information, driving routes, and driving speed data, enabling the system to analyze road conditions.

[0048] Specifically, the vehicle's driving conditions include prolonged driving on congested roads, highways, or rugged mountain roads. Different road conditions will have different effects on the vehicle's wear and tear. In particular, the vehicle's maintenance needs will change depending on the frequency of starts and stops, long-term high-speed driving, or harsh road conditions. Therefore, the data collected by the GPS device will be input into the system together with the sensor data to comprehensively analyze the impact of environmental factors on the vehicle, thereby more accurately predicting the maintenance cycle.

[0049] Specifically, the data acquisition module also includes an external data interface for receiving information from external data sources. Vehicle operation is not only affected by its own condition but also significantly influenced by the external environment, including climate conditions, traffic flow, and road conditions. Through the external data interface, the system can receive weather data, real-time traffic flow information, and road construction and road condition information. This data is obtained and updated in real time through cooperation with weather data providers, traffic information platforms, and road management systems. The introduction of external data can not only supplement environmental factors that are difficult for onboard sensors to capture but also enable the system to understand the operating environment of the vehicle. Vehicles that drive in high humidity, low temperature, or frequent rainfall environments for a long time are more prone to aging or rusting of their internal parts, which leads to a need to appropriately shorten the maintenance cycle.

[0050] Specifically, the collected data is stored in the data storage unit for subsequent analysis and prediction. The data storage unit has efficient storage and management functions, and can manage real-time data and historical data in a hierarchical manner. Real-time data is used for current maintenance cycle prediction, while historical data is used for model training and long-term trend analysis. During storage, the data will be grouped and labeled according to time, data type, and source device dimensions to facilitate rapid retrieval and analysis later.

[0051] Specifically, the data preprocessing module performs cleaning, denoising, and standardization preprocessing operations on the multi-source data collected by the data acquisition module. For numerical extreme values, it uses three times the standard deviation for detection; for outliers, it replaces them with the average value; for missing data, the system fills in the missing data through interpolation to avoid the missing data causing bias in model training. Then, it completes data denoising by optimizing the moving average method and averaging continuous data points to smooth the data, thereby reducing abnormal fluctuations. The denoised data is then standardized and normalized. Standardization adjusts the data to a standard normal distribution with the same mean and variance, while normalization scales the data proportionally to the [0,1] interval.

[0052] Specifically, the model training and prediction module trains the model using the random forest algorithm, as follows:

[0053] Define the training set as Where D is the training set, x i Let y be the feature vector of the i-th group, and N represent the number of samples. i For the target maintenance cycle, x i1 ,x i2 ,...x im Let N' be the first feature vector of the i-th sample, N' be the second feature vector of the i-th sample, and N' be the m-th feature vector of the i-th sample. N' samples are drawn from the training set with replacement, where N' ≤ N, and K distinct subsets are generated, each with length N'. When constructing each decision tree, FLOOR(sqrtm) features are randomly selected from the m features for splitting, where FLOOR is the downward fetching function. For each node of the k-th decision tree, the optimal splitting condition for all possible split points is calculated. For the current sample set of the k-th decision tree... Information gain (Gain) is expressed as follows:

[0054]

[0055] Where, p n The information gain (Gini index) represents the proportion of the nth sample in the sample set D. It measures the improvement in data purity before and after the split, calculated based on the information entropy using the Shannon formula. index ,have:

[0056]

[0057] The Gini index, calculated based on a probability distribution, measures data impurity. The mean squared error (MSE) is defined to measure the deviation between predicted and actual values.

[0058]

[0059] in, for Maintenance cycle, For computation in the random forest algorithm The maintenance cycle, the greater the information gain, the greater the improvement in the purity of the feature partitioning nodes. If we define maximization as forward optimization and minimization as reverse optimization, in order to unify the evaluation index, we reverse optimize the information gain as... Constructing indicator structure For K decision trees, there is a K*3 dimensional standardized evaluation matrix A*:

[0060]

[0061] in, This refers to the information gain inverse element in the standardized evaluation matrix of the first decision tree. This refers to the information gain inverse element in the standardized evaluation matrix of the second decision tree. Gini represents the information gain inverse element in the standardized evaluation matrix of the Kth decision tree. index (1) represents the Gini index element in the standardized evaluation matrix of the first decision tree. index (2) represents the Gini index element in the standardized evaluation matrix of the second decision tree. index (K) represents the Gini index element in the standardized evaluation matrix of the Kth decision tree. The mean squared error element in the standardized evaluation matrix of the first decision tree. The mean squared error element in the standardized evaluation matrix of the second decision tree. Let A* be the mean squared error element in the standardized evaluation matrix of the Kth decision tree, and define the maximum value vector of the standardized matrix as A*. + :

[0062]

[0063] Define the normalized matrix minimum vector A* -

[0064]

[0065] Then the evaluation metric SCORE of the k-th decision tree k The calculation is as follows:

[0066]

[0067] By calculating SCORE kThis is used to standardize Gain, Gini_index, and MSE to eliminate the influence of dimensions. The relative distance between each index and the optimal / worst solution is calculated using Euclidean distance to quantify the overall performance of the decision tree. The final score reflects the overall effectiveness of the decision tree and is used to guide model optimization.

[0068]

[0069] Specifically, prediction performance is improved by constructing a node loss function, as shown below:

[0070]

[0071] Where Loss(k) is the loss function of the k-th decision tree. To minimize the weighted loss of the entire decision tree, the loss increment ΔLoss is calculated. k :

[0072]

[0073] Among them, D k< D k The left child subset of the decision tree under the splitting rule π, D k> D k The right child subset of the decision tree under the splitting rule π has {D} k< ∩D k>}=D k Loss(k<) is D k< The loss function, Loss(k>), is D. k> The loss function, to prevent overfitting, introduces a regularization term and is as follows:

[0074]

[0075] Where Loss is the loss function of the random forest, α is the model complexity penalty coefficient, which constrains the number of samples contained in a node to prevent the tree structure from expanding indefinitely, and its value ranges from [0.01, 0.1]. If α is large, the node complexity penalty accounts for a higher proportion, and the system tends to remove child nodes containing fewer samples and prioritize retaining important nodes; if α is small, the penalty for node complexity is weakened, and the system tends to retain more child nodes, allowing for a more complex tree structure. β is the feature selection penalty coefficient, reflecting the importance distribution of nodes to different features, and its value ranges from [0.01, 0.1]. This is achieved through feature variance... To measure this, increasing β increases the penalty for feature variance, prompting the system to be more inclined to remove split nodes associated with low-importance features. Decreasing β weakens the penalty for feature selection, allowing the random forest to utilize more features for complex splits. mThe feature weights have a value range of [0,1] and represent the importance of the feature to the model's prediction.

[0076] Specifically, the maintenance cycle optimization module optimizes the maintenance cycle predicted by the model to provide accurate and personalized maintenance suggestions. By dynamically adjusting and optimizing the initially predicted maintenance cycle, the maintenance suggestions are made to meet actual needs. The maintenance cycle optimization module includes a historical data analysis submodule, a dynamic adjustment submodule, and a personalized suggestion generation submodule.

[0077] Specifically, the historical data analysis submodule identifies factors affecting maintenance cycles based on the vehicle's historical operating data and maintenance records, thereby making personalized adjustments to the prediction results; the dynamic adjustment submodule dynamically adjusts the predicted maintenance cycle based on the vehicle's real-time data and current status; while making dynamic adjustments, the personalized suggestion generation submodule generates specific maintenance suggestions based on all information for user reference, integrating the results from the prediction model, rule base, historical analysis, and dynamic adjustment to provide users with maintenance suggestions that better meet their actual needs.

[0078] Specifically, the user interface module provides users with vehicle status monitoring, maintenance cycle suggestions, and historical data query functions through a visual interface and interactive features, thereby improving user experience and system usability.

[0079] Specifically, the user interface module has a data display function, which transforms the data analyzed and predicted by the system into visual information that is easy for users to understand. It integrates the vehicle's real-time operating data, maintenance cycle suggestions, and historical maintenance records, and displays them in the form of charts, graphical dashboards, and graphs. For maintenance cycle suggestions, the system will display the specific recommended maintenance time or mileage, list the parts and systems that need maintenance, and provide an explanation for each suggestion. If a part is about to reach its maintenance cycle, a prompt will appear on the interface to remind the user to pay attention in time.

[0080] The beneficial effects of this invention are as follows: This invention integrates multi-dimensional data sources, including vehicle sensor data, historical maintenance records, driving environment data, and driving habit data, to fully explore factors affecting maintenance cycles. Traditional maintenance models typically rely on single mileage or time indicators, failing to accurately reflect the vehicle's true wear and tear and maintenance needs. This system, through comprehensive analysis of multiple data sources, can more accurately determine the actual health status of various vehicle components, significantly improving the accuracy of maintenance cycle prediction. This invention introduces a random forest algorithm model. The random forest algorithm is an ensemble learning-based algorithm that uses multiple decision trees to achieve voting decisions on prediction results. This invention utilizes the random forest algorithm for intelligent prediction of maintenance cycles, effectively addressing the shortcomings of traditional models in terms of data nonlinearity and high dimensionality. The random forest algorithm exhibits high robustness in handling complex, nonlinear data and can continuously adjust prediction results based on dynamic changes in input data. Unlike traditional random forest algorithms, this invention extends this by adding three indicators: Gain, Gini, and Gini. index The three metrics for evaluating random forests—Gain, Gini coefficient, and MSE—were standardized and aligned. index The three metrics, MSE, and β, contain the feature information of each decision tree in the random forest and are quantified into a K-row, 3-column matrix, greatly reducing the computational load of the system participating in the random forest. The deviation of each metric is calculated based on Euclidean distance to synthesize the missile score of each decision tree. This score information is associated with the loss function, which includes a regularization term to prevent overfitting. Unlike traditional regularization, this invention introduces a model complexity penalty coefficient and a feature selection penalty coefficient. α is the model complexity penalty coefficient, which constrains the number of samples a node contains, preventing unlimited expansion of the tree structure. If α is large, the node complexity penalty is high, and the system tends to remove child nodes with fewer samples, prioritizing the retention of important nodes. If α is small, the penalty for node complexity is weakened, and the system tends to retain more child nodes, allowing for a more complex tree structure. β is the feature selection penalty coefficient, reflecting the importance distribution of nodes for different features, and is calculated using feature variance. To measure this, increasing β increases the penalty for feature variance, prompting the system to be more inclined to remove split nodes associated with low-importance features. Conversely, decreasing β weakens the penalty for feature selection, allowing the random forest to utilize more features for complex splits. Furthermore, the random forest algorithm has strong resistance to outlier data, reducing the risk of overfitting in a single model. The introduction of this predictive model enables the system to maintain high prediction accuracy and applicability under different vehicles, driving environments, and driving habits, providing users with more precise maintenance cycle recommendations. This invention's intelligent optimization system for vehicle maintenance cycles based on the random forest algorithm has significant innovation and practicality. Through multiple technological innovations, including multi-dimensional data sources, intelligent predictive models, dynamic optimization, personalized recommendations, and remote monitoring, it provides accurate maintenance cycle recommendations for vehicles, optimizing user experience and maintenance costs. It meets the intelligent and personalized needs of modern vehicle management and has broad application prospects in future intelligent vehicle maintenance management. It not only improves vehicle safety and reliability but also provides car owners with more convenient and economical maintenance solutions, possessing significant market value and social benefits.

[0081] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make substitutions for some of the technical features. Any modifications, substitutions, or improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An intelligent optimization system for preventive vehicle maintenance based on a random forest algorithm, comprising a data acquisition module, a data preprocessing module, a model training and prediction module, a maintenance cycle optimization module, and a user interface module. The data acquisition module acquires real-time vehicle data through onboard sensors, GPS devices, etc. The data preprocessing module cleans, denoises, and standardizes the data. The model training and prediction module constructs an ensemble model with multiple decision trees based on the preprocessed data using the random forest algorithm, thereby improving prediction accuracy and robustness. The system inputs new data into the model and predicts a suitable maintenance cycle based on the driver's real-time driving behavior and the vehicle's health condition. The maintenance cycle optimization module dynamically adjusts the recommended cycle by combining the user's historical maintenance information and prediction results. The user interface module displays the suggested cycle and related analysis results to the user, enabling the user to perform maintenance at the appropriate time, thereby extending the vehicle's service life and reducing maintenance costs.

2. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 1, characterized in that, The data acquisition module consists of an onboard sensor integration unit, a Global Positioning System (GPS) device, an external data interface, and a data storage unit. The onboard sensor integration unit is the data acquisition module, which collects real-time internal operating data of the vehicle through engine sensors, fuel consumption sensors, temperature sensors, tire pressure sensors, brake wear sensors, battery power sensors, and speed sensors. Each sensor corresponds to a different collection task. The GPS device collects real-time vehicle location information, driving route, and driving speed data through the Global Positioning System. The system can analyze the vehicle's driving conditions. The data collected by the GPS device is input into the system along with the sensor data to analyze the impact of environmental factors on the vehicle, thereby more accurately predicting maintenance cycles. The data acquisition module also includes an external data interface for receiving information from external data sources. Vehicle operation is not only affected by its own state but also significantly influenced by the external environment. Through the external data interface, the system can receive weather data, real-time traffic flow information, and road construction and road condition information. This data is obtained through cooperation with weather data providers, traffic information platforms, and road management systems and is updated in real time. The introduction of external data not only supplements environmental factors that are difficult for onboard sensors to capture but also allows the system to understand the vehicle's operating environment.

3. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 2, characterized in that, The collected data is stored in the data storage unit for subsequent analysis and prediction. The data storage unit has efficient storage and management functions, and can manage real-time data and historical data in a hierarchical manner. Real-time data is used for current maintenance cycle prediction, while historical data is used for model training and long-term trend analysis. During storage, the data will be grouped and labeled according to time, data type, and source device dimensions to facilitate rapid retrieval and analysis later.

4. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 1, characterized in that, The data preprocessing module performs cleaning, denoising, and standardization preprocessing operations on the multi-source data collected by the data acquisition module. For numerical extreme values, it uses three times the standard deviation for detection; for outliers, it replaces them with the average value; for missing data, the system fills in the missing data through interpolation to avoid the missing data causing bias in model training. Then, it completes data denoising by optimizing the moving average method and averaging continuous data points to smooth the data, thereby reducing abnormal fluctuations. The denoised data is then standardized and normalized. Standardization adjusts the data to a standard normal distribution with the same mean and variance, while normalization scales the data proportionally to the [0,1] interval.

5. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 1, characterized in that, The model training and prediction module trains the model using the random forest algorithm, as detailed below: Define the training set as x i =[x i1 ,x i2 ,...x im ], where D is the training set, x i Let y be the feature vector of the i-th group, and N represent the number of samples. i For the target maintenance cycle, x i1 ,x i2 ,...x im Let N' be the first feature vector of the i-th sample, N' be the second feature vector of the i-th sample, and N' be the m-th feature vector of the i-th sample. N' samples are drawn from the training set with replacement, where N' ≤ N, and K distinct subsets are generated, each with length N'. When constructing each decision tree, FLOOR(sqrtm) features are randomly selected from the m features for splitting, where FLOOR is the downward fetching function. For each node of the k-th decision tree, the optimal splitting condition for all possible split points is calculated. For the current sample set of the k-th decision tree... Information gain (Gain) is expressed as follows: Where, p n The information gain (Gini index) represents the proportion of the nth sample in the sample set D. It measures the improvement in data purity before and after the split, calculated based on the information entropy using the Shannon formula. index ,have: The Gini index, calculated based on a probability distribution, measures data impurity. The mean squared error (MSE) is defined to measure the deviation between predicted and actual values. in, for Maintenance cycle, For computation in the random forest algorithm The maintenance cycle, the greater the information gain, the greater the improvement in the purity of the feature partitioning nodes. If we define maximization as forward optimization and minimization as reverse optimization, in order to unify the evaluation index, we reverse optimize the information gain as... Constructing indicator structure For K decision trees, there is a K*3 dimensional standardized evaluation matrix A*: in, This refers to the information gain inverse element in the standardized evaluation matrix of the first decision tree. This refers to the information gain inverse element in the standardized evaluation matrix of the second decision tree. Gini represents the information gain inverse element in the standardized evaluation matrix of the Kth decision tree. index (1) represents the Gini index element in the standardized evaluation matrix of the first decision tree. index (2) represents the Gini index element in the standardized evaluation matrix of the second decision tree. index (K) represents the Gini index element in the standardized evaluation matrix of the Kth decision tree. The mean squared error element in the standardized evaluation matrix of the first decision tree. The mean squared error element in the standardized evaluation matrix of the second decision tree. Let A* be the mean squared error element in the standardized evaluation matrix of the Kth decision tree, and define the maximum value vector of the standardized matrix as A*. + : Define the normalized matrix minimum vector A* - Then the evaluation metric SCORE of the k-th decision tree k The calculation is as follows: By calculating SCORE k This is used to standardize Gain, Gini_index, and MSE to eliminate the influence of dimensions. The relative distance between each index and the optimal / worst solution is calculated using Euclidean distance to quantify the overall performance of the decision tree. The final score reflects the overall effectiveness of the decision tree and is used to guide model optimization.

6. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 5, characterized in that, The prediction performance is improved by constructing a node loss function, as shown below: Where Loss(k) is the loss function of the k-th decision tree. To minimize the weighted loss of the entire decision tree, the loss increment ΔLoss is calculated. k : Among them, D k< D k The left child subset of the decision tree under the splitting rule π, D k> D k The right child subset of the decision tree under the splitting rule π has {D} k< ∩D k> }=D k Loss(k<) is D k< The loss function, Loss(k>), is D. k> The loss function, to prevent overfitting, introduces a regularization term and is as follows: Where Loss is the loss function of the random forest, α is the model complexity penalty coefficient, which constrains the number of samples contained in a node to prevent the tree structure from expanding indefinitely, and its value ranges from [0.01, 0.1]. If α is large, the node complexity penalty accounts for a higher proportion, and the system tends to remove child nodes containing fewer samples and prioritize retaining important nodes; if α is small, the penalty for node complexity is weakened, and the system tends to retain more child nodes, allowing for a more complex tree structure. β is the feature selection penalty coefficient, reflecting the importance distribution of nodes to different features, and its value ranges from [0.01, 0.1]. This is achieved through feature variance... To measure this, increasing β increases the penalty for feature variance, prompting the system to be more inclined to remove split nodes associated with low-importance features. Decreasing β weakens the penalty for feature selection, allowing the random forest to utilize more features for complex splits. m The feature weights have a value range of [0,1] and represent the importance of the feature to the model's prediction.

7. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 1, characterized in that, The maintenance cycle optimization module optimizes the maintenance cycle predicted by the model to provide accurate and personalized maintenance suggestions. By dynamically adjusting and optimizing the initially predicted maintenance cycle, the maintenance suggestions are made to meet actual needs. The maintenance cycle optimization module includes a historical data analysis submodule, a dynamic adjustment submodule, and a personalized suggestion generation submodule.

8. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 7, characterized in that, The historical data analysis submodule identifies factors affecting maintenance cycles based on the vehicle's historical operating data and maintenance records, thereby making personalized adjustments to the prediction results. The dynamic adjustment submodule dynamically adjusts the predicted maintenance cycle based on the vehicle's real-time data and current status. While making dynamic adjustments, the personalized suggestion generation submodule generates specific maintenance suggestions based on all information for user reference. By integrating the results from the prediction model, rule base, historical analysis, and dynamic adjustment, the submodule provides users with maintenance suggestions that better meet their actual needs.

9. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 1, characterized in that, The user interface module provides users with vehicle status monitoring, maintenance cycle suggestions, and historical data query functions through a visual interface and interactive features, thereby improving user experience and system usability.

10. The intelligent optimization system for preventive maintenance of automobiles based on random forest algorithm according to claim 1, characterized in that, The user interface module features data display capabilities, transforming system-analyzed and predicted data into easily understandable visual information. It integrates real-time vehicle operating data, maintenance cycle recommendations, and historical maintenance records, displaying them through charts, graphical dashboards, and graphs. For maintenance cycle recommendations, the system displays specific recommended maintenance times or mileages, lists the components and systems requiring maintenance, and provides explanations for each recommendation. If a component is nearing its maintenance cycle, a notification will appear on the interface to remind the user to pay attention.

Citation Information

Patent Citations

  • Process modular scheme evaluation method for large-scale customization

    CN111985787A

  • Photovoltaic power prediction error evaluation method considering cloud cluster influence

    CN118761040A

  • Vehicle part service life prediction method based on user driving habits

    CN118981697A

  • Automobile preventive maintenance intelligent optimization system based on random forest algorithm

    CN119624420A

  • Smart industrial safety monitoring system using IoT and machine learning

    IN202041040672A

Cited By

  • Vehicle sensor cleaning method and device, electronic equipment and storage medium

    CN121302148A