Reliability prediction method and device based on vehicle after-sales data and electronic equipment

By preprocessing vehicle after-sales data from multiple data sources and training machine learning algorithms, a reliability prediction model is constructed, which solves the problems of incomplete data coverage and weak model generalization ability in traditional methods, and achieves accurate prediction of vehicle reliability and fault warning.

CN121684979APending Publication Date: 2026-03-17CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511762879.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Traditional vehicle reliability prediction methods rely on expert experience and limited historical data, resulting in incomplete data coverage and weak generalization ability of prediction models, making it difficult to meet the needs of accurate prediction.

Method used

By acquiring and preprocessing vehicle after-sales data from multiple data sources, predictive features related to vehicle reliability are extracted, and a reliability prediction model is constructed using machine learning algorithms, including data cleaning, encoding, normalization, feature selection, and model training. Hyperparameters are optimized to improve prediction accuracy.

Benefits of technology

It enables accurate prediction of vehicle reliability, provides information on potential fault types, probability of occurrence, and preventive measures, and improves the level of reliability management throughout the entire vehicle lifecycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121684979A_ABST
    Figure CN121684979A_ABST
Patent Text Reader

Abstract

The invention provides a reliability prediction method and device based on vehicle after-sales data and electronic equipment, and relates to the field of vehicle reliability prediction.The method comprises the steps that the vehicle after-sales data of multiple data sources are preprocessed, and standard vehicle after-sales data are obtained; extracting prediction features related to vehicle reliability from the standard vehicle after-sales data; training the prediction features by adopting a machine learning algorithm to obtain a reliability prediction model; and processing the current vehicle after-sales data of the target vehicle by adopting the trained reliability prediction model, and outputting a reliability prediction result of the target vehicle. Through multi-source vehicle after-sales data preprocessing, prediction features are accurately extracted, and a reliability prediction model is trained in combination with a machine learning algorithm, so that the potential fault type, occurrence probability and prevention measures of the target vehicle can be accurately output, and the problem of low prediction precision caused by experience or single data traditionally is solved; and the full-life-cycle reliability management level of the automobile is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automotive reliability prediction, and more specifically, to a reliability prediction method, apparatus, and electronic device based on vehicle after-sales data. Background Technology

[0002] Vehicle reliability is a core indicator influencing consumers' car-buying decisions and manufacturers' market competitiveness. Accurate prediction of vehicle reliability is crucial for reducing after-sales costs and improving user experience. Traditional vehicle reliability prediction methods rely heavily on expert experience and limited historical data, resulting in incomplete data coverage and weak generalization ability of prediction models, making it difficult to meet actual needs in terms of prediction accuracy.

[0003] With the development of big data and machine learning technologies, reliability prediction methods based on vehicle after-sales data have gradually become a research hotspot. Vehicle after-sales data contains key information such as vehicle model, fault type, and maintenance records, directly reflecting the reliability status of vehicles in actual use. However, existing technologies have failed to fully explore the deeper value of vehicle after-sales data, exhibiting shortcomings in areas such as data redundancy removal, key feature selection, and model adaptive optimization. This makes it difficult to construct prediction models that balance accuracy and universality, and hinders their effective adaptation to the reliability prediction needs of different brands and models of vehicles. Summary of the Invention

[0004] The purpose of this application is to provide a reliability prediction method, device, and electronic device based on vehicle after-sales data, which solves the above-mentioned problems existing in the prior art, can make full use of vehicle after-sales data, build a prediction model through machine learning, and improve the accuracy of vehicle reliability prediction.

[0005] Firstly, a reliability prediction method based on vehicle after-sales data is provided, which may include: Obtain vehicle after-sales data from multiple data sources for the same model as the target vehicle during historical periods, and preprocess the vehicle after-sales data from the multiple data sources to obtain standard vehicle after-sales data. Extract predictive features related to vehicle reliability from the standard vehicle after-sales data; A reliability prediction model is obtained by training the predicted features using machine learning algorithms. The trained reliability prediction model is used to process the current after-sales data of the target vehicle and output the reliability prediction result of the target vehicle.

[0006] In one possible implementation, the vehicle after-sales data from the multiple data sources is preprocessed to obtain standard vehicle after-sales data, including: For vehicle after-sales data from any data source, perform data cleaning on the vehicle after-sales data to obtain initial vehicle after-sales data; The initial vehicle after-sales data is transformed to obtain target vehicle after-sales data; wherein, the data transformation includes encoding non-numerical features and normalizing numerical features. The standard vehicle after-sales data is obtained by associating and integrating target vehicle after-sales data from different data sources.

[0007] In one possible implementation, predictive features related to vehicle reliability are extracted from the standard vehicle after-sales data, including: Based on the configured vehicle domain data, a candidate feature set is determined from the standard vehicle after-sales data; The contribution of each feature in the candidate feature set to the reliability prediction is calculated using a feature importance assessment method. The candidate features are sorted according to their contribution, and features with a ranking higher than a preset threshold are selected as the predicted features.

[0008] In one possible implementation, the feature importance evaluation method is either the Random Forest algorithm or the XGBoost algorithm.

[0009] In one possible implementation, a machine learning algorithm is used to train the predicted features to obtain a reliability prediction model, including: The dataset containing the predicted features is divided into a training set and a test set; The selected machine learning algorithm is trained using the training set, and the performance of the target model is verified using the test set to obtain the verification results. Based on the verification results, the hyperparameters of the target model are tuned to obtain the reliability prediction model.

[0010] In one possible implementation, the machine learning algorithm is a gradient boosting decision tree model; The performance of the target model is verified by at least one of the following metrics: accuracy, recall, and F1 score.

[0011] In one possible implementation, after outputting the reliability prediction result of the target vehicle, the method further includes: According to the configured collection cycle, newly added vehicle after-sales data is collected, and the newly added vehicle after-sales data is preprocessed and feature extracted to obtain new samples. The newly added samples are used to incrementally learn or periodically retrain the deployed reliability prediction model to obtain the target reliability prediction model.

[0012] Secondly, a reliability prediction device based on vehicle after-sales data is provided, the device may include: The acquisition unit is used to acquire vehicle after-sales data from multiple data sources of the same model as the target vehicle during a historical period, and to preprocess the vehicle after-sales data from the multiple data sources to obtain standard vehicle after-sales data. An extraction unit is used to extract predictive features related to vehicle reliability from the standard vehicle after-sales data. The training unit is used to train the prediction features using machine learning algorithms to obtain a reliability prediction model; The processing unit is used to process the current after-sales data of the target vehicle using the trained reliability prediction model, and output the reliability prediction result of the target vehicle.

[0013] Thirdly, an electronic device is provided, which includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When a processor executes a program stored in memory, it implements any of the steps described in the first aspect above.

[0014] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the steps of any of the methods described in the first aspect above.

[0015] This application provides a reliability prediction method, apparatus, and electronic device based on vehicle after-sales data. The method includes: acquiring vehicle after-sales data from multiple data sources of the same model as the target vehicle during historical periods; preprocessing the multi-source vehicle after-sales data to obtain standard vehicle after-sales data; extracting predictive features related to vehicle reliability from the standard vehicle after-sales data; training the predictive features using a machine learning algorithm to obtain a reliability prediction model; and using the trained reliability prediction model to process the current vehicle after-sales data of the target vehicle and outputting the reliability prediction result of the target vehicle. This application, through multi-source vehicle after-sales data preprocessing, accurate feature extraction, and machine learning algorithm training of the reliability prediction model, can accurately output the potential fault types, probabilities of occurrence, and preventive measures for the target vehicle. It solves the problem of low prediction accuracy caused by traditional reliance on experience or single data, providing users with more reliable vehicle usage guarantees and significantly improving the level of reliability management throughout the entire vehicle lifecycle. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 A system architecture diagram for a reliability prediction method based on vehicle after-sales data, provided in an embodiment of this application; Figure 2 A flowchart illustrating a reliability prediction method based on vehicle after-sales data provided in this application embodiment; Figure 3 A flowchart illustrating the collection of vehicle after-sales data provided in this application embodiment; Figure 4 A schematic diagram illustrating the decision-making process of the machine learning model provided in this application embodiment; Figure 5 A schematic diagram illustrating the model training process provided in this application embodiment; Figure 6 A schematic diagram of a reliability prediction device based on vehicle after-sales data provided in this application embodiment; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0019] The reliability prediction method based on vehicle after-sales data provided in this application embodiment can be applied to... Figure 1 In the system architecture shown, such as Figure 1As shown, the system may include a server and a terminal. The server can be a physical server, a server cluster consisting of multiple physical servers, or a distributed system. It can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal may be a user equipment (UE) such as a mobile phone, smartphone, laptop, digital radio receiver, personal digital assistant (PDA), tablet computer (PAD), handheld device, in-vehicle device, wearable device, computing device, or other processing device connected to a wireless modem, mobile station (MS), mobile terminal, etc. The terminal and server can be directly or indirectly connected via wired or wireless communication methods, which is not limited herein.

[0020] The terminal is used to acquire vehicle after-sales data from multiple data sources for the same model as the target vehicle during historical periods, and to send the vehicle after-sales data from multiple data sources to the server. A server is used to receive vehicle after-sales data from multiple data sources to execute a reliability prediction method based on vehicle after-sales data provided in this application.

[0021] Vehicle reliability is a core indicator influencing consumers' car-buying decisions and manufacturers' market competitiveness. Accurate prediction of vehicle reliability is crucial for reducing after-sales costs and improving user experience. Traditional vehicle reliability prediction methods rely heavily on expert experience and limited historical data, resulting in incomplete data coverage and weak generalization ability of prediction models, making it difficult to meet actual needs in terms of prediction accuracy.

[0022] With the development of big data and machine learning technologies, reliability prediction methods based on vehicle after-sales data have gradually become a research hotspot. Vehicle after-sales data contains key information such as vehicle model, fault type, and maintenance records, directly reflecting the reliability status of vehicles in actual use. However, existing technologies have failed to fully explore the deeper value of vehicle after-sales data, exhibiting shortcomings in areas such as data redundancy removal, key feature selection, and model adaptive optimization. This makes it difficult to construct prediction models that balance accuracy and universality, and hinders their effective adaptation to the reliability prediction needs of different brands and models of vehicles.

[0023] Therefore, this application provides a reliability prediction method based on vehicle after-sales data to solve the above-mentioned problems in the prior art. It can make full use of vehicle after-sales data, build a prediction model through machine learning, and improve the accuracy of vehicle reliability prediction.

[0024] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0025] Figure 2 This is a flowchart illustrating a reliability prediction method based on vehicle after-sales data, provided as an embodiment of this application. Figure 2 As shown, the method may include: Step S210: Obtain vehicle after-sales data from multiple data sources for the same model as the target vehicle during historical periods, and preprocess the vehicle after-sales data from multiple data sources to obtain standard vehicle after-sales data.

[0026] Among them, combined Figure 3 As shown, the process for collecting vehicle after-sales data (including vehicle claim data, after-sales repair record data, etc.) is as follows: Data collection follows a complete chain from fault occurrence, claim triggering, data upload, and data aggregation: When a vehicle of the same model experiences a fault after being sold, it determines whether the user has initiated a claim. If a claim is initiated, corresponding vehicle claim data is generated. After the data is uploaded to the vehicle manufacturer's database, the repair service center's business system, or the customer feedback platform, it is downloaded and collected through authorized interfaces. The multiple data sources collected include, but are not limited to, the vehicle manufacturer's vehicle claim database, the repair record system of authorized repair service centers, and claim complaint data from the customer feedback platform. The data content covers key fields such as Vehicle Identification Number (VIN), vehicle model, time and location of fault occurrence, fault type, repair items, repair duration, and parts replacement information, ensuring that the data comprehensively reflects the historical reliability status of the same vehicle model.

[0027] This includes preprocessing vehicle after-sales data from multiple data sources to obtain standard vehicle after-sales data, including: For vehicle after-sales data from any data source, data cleaning is performed to obtain initial vehicle after-sales data. Specifically, this includes: comparing the Vehicle Identification Number (VIN) with vehicle model information to filter out claim records corresponding only to the same model as the target vehicle; removing redundant data submitted repeatedly (such as multiple uploads of the same claim); correcting obvious errors (such as logical inconsistencies in fault occurrence time, mismatches between repair items and fault types); and using linear interpolation to complete a few missing key fields (such as missing repair time, which can be filled in based on the average repair time of similar faults). The final result is initial vehicle after-sales data that meets quality standards and has complete information, laying the foundation for subsequent processing.

[0028] The initial vehicle after-sales data is transformed to obtain the target vehicle after-sales data. This transformation includes encoding non-numerical features and normalizing numerical features. Specifically, for non-numerical features such as fault type, repair item, and part name, one-hot encoding or label encoding is used for quantification (e.g., converting fault types such as engine fault and transmission fault into unique numerical identifiers) to adapt them to the input requirements of machine learning algorithms. For numerical features such as repair time, fault interval mileage, and part price, Min-Max normalization is used to map them to the [0,1] interval, eliminating the interference of different units on model training and ensuring that the weights of each feature are balanced during model training.

[0029] After-sales data of target vehicles from different data sources are correlated and integrated to obtain standard vehicle after-sales data. Specific operations include: matching and correlating target vehicle after-sales data from manufacturer databases, service centers, and customer feedback platforms using Vehicle Identification Number (VIN) and fault occurrence time as core correlation fields; standardizing data formats and field naming rules (e.g., unifying fault date and repair report time as fault occurrence time); removing duplicate and redundant records after correlation; and supplementing missing information across data sources (e.g., supplementing missing parts specification fields in service center records using parts information from the manufacturer database).

[0030] The resulting standard vehicle after-sales data features a unified format, complete fields, and consistent logic, and can be directly used for subsequent feature selection and machine learning model training, providing high-quality data support for reliability prediction of the same vehicle model.

[0031] Step S220: Extract predictive features related to vehicle reliability from standard vehicle after-sales data.

[0032] Specifically, step 1 involves determining a candidate feature set from standard vehicle after-sales data based on the configured vehicle domain data. This process involves filtering and determining the candidate feature set from the acquired standard vehicle after-sales data, ensuring that the candidate features comprehensively cover the key dimensions affecting vehicle reliability. The configured vehicle domain data includes recognized reliability correlation indicators in the automotive engineering field (such as failure mode classification standards, maintenance complexity grading rules, and vehicle usage intensity evaluation dimensions). The core fields of the standard vehicle after-sales data (which have been standardized in format through the preceding preprocessing) include vehicle model, failure type, failure occurrence time, failure interval mileage, repair items, repair duration, parts replacement type, parts replacement frequency, failure recurrence frequency, and post-repair failure interval time. The identified candidate feature set specifically includes: A) Fault-related features: fault type (e.g., quantified and encoded features such as engine faults, transmission faults, etc.), fault occurrence frequency (number of faults of the same vehicle model per unit time), fault recurrence frequency (number of repairs for the same fault in the same vehicle), and fault interval mileage (mileage traveled between two consecutive faults); B) Repair-related features: repair time (repair time for a single fault), repair complexity (quantified indicators based on the number and technical difficulty of repair items, such as high complexity for engine replacement and low complexity for filter replacement), parts replacement type (classification features of critical / non-critical parts), and parts replacement frequency (number of times the same part is replaced per unit time); C) Usage-related features: vehicle usage time (time from vehicle manufacturing to the first fault), cumulative mileage at the time of fault occurrence (total mileage traveled by the vehicle when the fault is triggered), and fault interval time after repair (time from repair completion to the next fault). The above candidate feature set comprehensively reflects the correlation between vehicle reliability and dimensions such as fault mode, repair effectiveness, and wear and tear, and all are derived from valid fields of standard vehicle after-sales data.

[0033] Step 2: Employ a feature importance assessment method to calculate the contribution of each feature in the candidate feature set to reliability prediction. The feature importance assessment method can be either the Random Forest algorithm or the XGBoost algorithm. Calculate the contribution of each feature in the candidate feature set to vehicle reliability prediction and quantify the predictive value of each feature. The specific implementation process is as follows: Method 1: Configure the core parameters of the Random Forest (e.g., set the number of decision trees to 100-200, the maximum tree depth to 10-15 layers, and the feature subset selection method to sqrt). Combine the feature data corresponding to the candidate feature set with vehicle reliability labels (e.g., failure probability, fault-free operating time, etc., based on fault record annotations in standard vehicle after-sales data) to form a training dataset. Method 2: Configure the parameters of XGBoost (e.g., set the learning rate to 0.1-0.3, the maximum tree depth to 6-10 layers, and the regularization coefficient λ to 1-3). Use the same training dataset as the Random Forest to ensure that the evaluation benchmarks of the two algorithms are consistent. Next, the Random Forest algorithm evaluates feature importance using out-of-bag (OOB) error: during the training of each decision tree, some features are randomly removed, and the increase in OOB error after removing a feature is calculated. The larger the increase in error, the higher the contribution of that feature to reliability prediction. Alternatively, the XGBoost algorithm evaluates contribution using feature split gain: the information gain (e.g., gain based on the Gini coefficient or mean squared error) brought by each feature during all decision tree splits is calculated. The higher the cumulative gain, the stronger the feature's contribution to reliability prediction. Finally, the configured Random Forest or XGBoost algorithm is run, outputting a quantified contribution value (e.g., a normalized value in the 0-1 range, with larger values ​​indicating higher contribution) for each feature in the candidate feature set, forming a correspondence table between features and their contributions.

[0034] Step 3: Sort candidate features according to their contribution and select features ranked higher than the configured preset threshold as predicted features. Specifically, sort candidate features in descending order and, based on the configured preset threshold, select features ranked higher than the threshold as the final predicted features. This ensures that the predicted features are both representative and concise, avoiding redundant features that could affect model training efficiency. Based on the feature-contribution correspondence table, sort candidate features from highest to lowest contribution (e.g., the sorting result might be: frequency of failure > time between failures after repair > frequency of parts replacement > repair complexity > type of failure > vehicle usage time, etc.), and generate a feature ranking list; preset feature contribution thresholds (this threshold can be configured based on actual prediction needs, such as setting it to the top 60% of contribution or an absolute contribution value ≥ 0.05; the threshold can be increased to improve prediction accuracy, and appropriately decreased to consider model generalization ability); compare the contribution of each feature in the feature ranking list with the preset threshold, and select features ranked higher than the threshold as predicted features. For example, if the preset threshold is the top 60% of the contribution, then the top 60% of the features are selected from the sorting list, and low-contribution features ranked lower (such as the vehicle's manufacturing year, the area of ​​the repair shop, etc., which have little impact on reliability prediction) are removed.

[0035] In summary, the final selected predictive features not only cover the core influencing factors of vehicle reliability prediction, but also eliminate redundant information through contribution evaluation. These features can be directly input into subsequent machine learning models, providing accurate and efficient feature input for reliability prediction of the same vehicle model, and ensuring the model's prediction accuracy and generalization ability.

[0036] Step S230: Use machine learning algorithms to train the prediction features to obtain a reliability prediction model.

[0037] Specifically, to ensure the accuracy and adaptability of the vehicle reliability prediction model, it is necessary to combine the characteristics of standard vehicle after-sales data with prediction requirements, and follow... Figure 4 The model selection process shown proceeds sequentially through judgment and selection, with the specific steps as follows: First, we analyze the characteristics of standard vehicle after-sales data, with the core judgment dimension being whether it contains time-to-event information. Among them, time-to-event information refers to the time span data from a certain key node (such as the new car leaving the factory or the last repair being completed) to the occurrence of the fault event (for example, the vehicle's fault-free running time can be calculated from the fault occurrence time and the vehicle's factory time, and the fault-free running time after repair can be calculated from the fault occurrence time and the repair completion time).

[0038] In this application, standard vehicle after-sales data includes key fields such as the time of failure, the vehicle's manufacturing time, and the time of repair completion. Time-to-event information can be clearly obtained through field association and calculation.

[0039] The core advantage of the survival analysis model lies in its ability to effectively handle censored data (i.e., some vehicles did not experience malfunctions during the observation period, and only the status data of vehicles that were still malfunction-free at the end of the observation period is recorded). Further assessment is needed to determine whether the vehicle claim data in this application contains censored data.

[0040] In vehicle reliability prediction scenarios, some vehicles will inevitably not experience any failures during the observation period (such as samples from new car launches or some samples from high-reliability models). Their vehicle claim data only records the observation cutoff time and not the failure occurrence time, which is typical censored data. Therefore, this application requires processing the censored data.

[0041] Prioritize survival analysis models (such as the Cox proportional hazards model, Weibull distribution model, etc.), while optimizing the selection based on other factors such as problem nature, model interpretability, model performance, and computational resources. Problem nature fit: This application focuses on the changing pattern of vehicle failure probability over time. The survival analysis model can directly output the risk function (i.e., the probability of a vehicle failing at a certain moment), which is highly consistent with the prediction requirements. Model interpretability assurance: Taking the Cox proportional hazards model as an example, it can output the risk coefficient of each predicted feature (such as failure frequency and maintenance complexity), intuitively explaining the degree of influence of the feature on the occurrence of failure (risk coefficient > 1 indicates that the feature will increase the failure risk), which is convenient for technicians to understand and verify; Model performance advantages: The survival analysis model has significantly higher fitting accuracy for censored data than traditional machine learning models, and can effectively capture the nonlinear relationship between fault-free runtime and fault risk; Computational resource compatibility: Survival analysis models such as the Cox proportional hazards model have low computational complexity and can be quickly trained on ordinary servers or in cloud environments, adapting to the engineering deployment requirements of this solution.

[0042] In some embodiments, if the data characteristics do not include time-to-event information (e.g., only predicting "whether a failure will occur" without considering the time dimension), then the process proceeds to the "no" branch, considering traditional machine learning models (such as random forests, XGBoost, support vector machines, etc.). In this case, further consideration of the problem nature, model interpretability, model performance, and computational resources is needed for selection. If a balance between accuracy and interpretability is required, the random forest model can be chosen, which can output high prediction accuracy and interpret the impact of each factor on the failure through feature importance assessment. If extremely high accuracy is required and sufficient computing resources are available, the XGBoost model can be selected, which can fully explore the complex correlations in car claim data through gradient boosting strategy. If the data has extremely high dimensionality and linear interpretability is required, a support vector machine model can be chosen. This model captures non-linear relationships through kernel function mapping while maintaining the interpretability of the model.

[0043] In summary, based on the characteristics of vehicle claim data, which includes time-to-event information and contains censored data, this application prioritizes the Cox proportional hazards model as the core prediction model, while retaining traditional machine learning models as alternatives, to ensure accurate vehicle reliability prediction under different data scenarios.

[0044] Step S230 specifically includes: dividing the dataset containing the prediction features into a training set and a test set; that is, dividing the dataset containing the prediction features into a training set and a test set according to the principle of data distribution consistency. In practice, a 7:3 or 8:2 ratio can be used (for example, 70% of the data is used for training and 30% for testing). During the division process, it is necessary to ensure that the two datasets are consistent in terms of vehicle model, fault type, and key prediction feature distribution to avoid insufficient model generalization ability due to data skew. For example, if a certain model accounts for 20% of the overall data, then the proportion of that model in the training set and the test set should also be close to 20% to ensure the prediction stability of the model in all scenarios.

[0045] The selected machine learning algorithm is trained using a training set, and the performance of the target model is validated using a test set to obtain validation results. A gradient boosting decision tree model (such as XGBoost) is selected as the core machine learning algorithm, and the model is trained using the training set. During training, the model learns the non-linear mapping relationship between predicted features (such as failure frequency, maintenance complexity, and parts replacement frequency) and vehicle reliability labels (such as failure probability and fault-free operating time). Taking XGBoost as an example, it iteratively constructs multiple decision trees, with each tree fitting based on the residual of the previous tree, gradually improving the model's prediction accuracy for vehicle reliability.

[0046] Based on the validation results, the hyperparameters of the target model are tuned to obtain a reliability prediction model.

[0047] The machine learning algorithm used is a gradient boosting decision tree model. The performance of the target model is validated by at least one of the following metrics: accuracy, recall, and F1 score.

[0048] The specific process of performance verification is as follows: Accuracy is the proportion of correctly predicted samples out of the total number of samples in the test set, reflecting the overall accuracy of the model's predictions. The formula is: Accuracy = Number of correctly predicted samples / Total number of samples in the test set; Recall is the proportion of actual fault samples correctly identified by the model out of all actual fault samples in the test set, reflecting the model's ability to identify fault samples. The formula is: Recall rate = Number of correctly identified faulty samples / Number of actual faulty samples in the test set; The F1 score is the harmonic mean of precision and recall, comprehensively evaluating the model's balance between accuracy and recall. The formula is:

[0049] For example, if the model accurately predicts 85% of the samples in the test set (85% accuracy) and identifies 90% of the actual faulty samples (90% recall), then the F1 score can quantify the model's overall performance in avoiding false positives and false negatives, providing a comprehensive reference for model performance.

[0050] Based on the performance verification results, the hyperparameters of the gradient boosting decision tree model are tuned to obtain the optimal reliability prediction model. The hyperparameters to be tuned include, but are not limited to: Learning rate: controls how well each decision tree fits the residual. A smaller learning rate can improve model stability, but it will increase training time. Number of decision trees: This refers to the total number of trees used in iterative training. More trees can improve the model's fitting ability, but may lead to overfitting. Maximum tree depth: limits the complexity of a single decision tree. Deeper trees can capture finer-grained feature relationships, but are prone to overfitting. Regularization coefficients: L1 / L2 regularization is used to constrain model complexity and prevent overfitting.

[0051] The optimization process employs an iterative verification method: for example, if the model's accuracy is insufficient, the number of decision trees can be increased or the maximum tree depth can be appropriately increased to enhance the fitting ability; if the model is overfitting (the performance on the test set is much lower than that on the training set), the regularization coefficient can be increased or the learning rate can be decreased. Through multiple rounds of parameter adjustment and performance verification iterations, the reliability prediction model that performs best in terms of accuracy, recall, and F1 score is finally obtained, ensuring that it has high accuracy and strong generalization ability in the task of predicting automotive reliability.

[0052] In summary, by appropriately partitioning the dataset, training the gradient boosting decision tree model, validating the performance of multiple metrics, and optimizing hyperparameters, this approach can construct an accurate vehicle reliability prediction model, providing strong technical support for reliability analysis and fault warning of the same vehicle model.

[0053] Step S240: Using the trained reliability prediction model, process the current after-sales data of the target vehicle and output the reliability prediction result of the target vehicle.

[0054] The reliability prediction results include, but are not limited to: the potential fault types of the target vehicle (such as specific categories like engine faults and transmission faults), the probability of potential fault occurrence (represented by a value in the range of 0-1, with higher values ​​indicating higher fault risk), and preventive measures for potential faults (such as recommended maintenance suggestions like checking transmission fluid every 5,000 kilometers), providing accurate decision-making basis for automakers or repair service providers.

[0055] After outputting the reliability prediction results for the target vehicle, the method further includes: Combination Figure 5 As shown, new vehicle after-sales data is collected according to the configured collection cycle. This new data is then preprocessed and its features extracted to obtain new samples. Specifically, new vehicle after-sales data is collected according to the configured collection cycle (which can be set weekly, monthly, etc.). The new vehicle after-sales data undergoes the same preprocessing and feature extraction process as historical data: first, data cleaning, transformation, and integration are performed to obtain standard vehicle after-sales data; then, based on vehicle domain data and feature importance assessment methods, predictive features related to reliability are extracted, ultimately yielding new samples.

[0056] In some embodiments, if the data fluctuation is temporary (such as abnormal data caused by short-term extreme usage scenarios of the target vehicle), the system proceeds to the judgment stage of whether to continue prediction; if the data distribution changes over a long period of time (such as changes in reliability patterns caused by vehicle model iteration or parts process updates) or the model's generalization ability is insufficient, the system proceeds to the data retraining stage.

[0057] If further predictions are needed (e.g., there is still a need for reliability predictions for vehicles of the same model), the system returns to the new data input stage and uses the current model to continue processing new vehicle after-sales data; if further predictions are not needed or the model needs to be optimized, the system enters the data retraining stage.

[0058] Next, the deployed reliability prediction model is incrementally learned or periodically retrained using the new samples to obtain the target reliability prediction model. If a gradient boosting decision tree model (such as XGBoost) is used, its incremental training feature can be leveraged to use the new samples as new training data, iteratively training the model on the existing basis. This allows the model parameters to adapt to the new data distribution, reducing computational resource consumption and shortening the update cycle. If the number of new samples is large or the data distribution changes significantly, the new samples are merged with historical samples, and a full retraining process is performed according to dataset partitioning, model training, performance verification, and hyperparameter tuning. This ensures the model fully learns the features of both new and old data, maintaining stable prediction accuracy.

[0059] Through the aforementioned iterative update mechanism, the reliability prediction model can be continuously optimized, maintaining its ability to accurately predict vehicle reliability and providing long-term technical support for vehicle reliability management throughout its entire lifecycle.

[0060] This application provides a reliability prediction method based on vehicle after-sales data. The method includes: acquiring vehicle after-sales data from multiple data sources for the same vehicle model as the target vehicle during historical periods; preprocessing the multi-source vehicle after-sales data to obtain standard vehicle after-sales data; extracting predictive features related to vehicle reliability from the standard vehicle after-sales data; training the predictive features using a machine learning algorithm to obtain a reliability prediction model; and using the trained reliability prediction model to process the current vehicle after-sales data of the target vehicle, outputting the reliability prediction result for the target vehicle. This application, through multi-source vehicle after-sales data preprocessing, accurate feature extraction, and training a reliability prediction model using machine learning algorithms, can accurately output the potential fault types, probabilities of occurrence, and preventative measures for the target vehicle. This solves the problem of low prediction accuracy caused by traditional reliance on experience or single data, providing users with more reliable vehicle usage assurance and significantly improving the level of reliability management throughout the entire vehicle lifecycle.

[0061] Corresponding to the above method, embodiments of this application also provide a reliability prediction device based on vehicle after-sales data, such as... Figure 6 As shown, the device includes: The acquisition unit 610 is used to acquire vehicle after-sales data from multiple data sources of the same model as the target vehicle in a historical period, and to preprocess the vehicle after-sales data from the multiple data sources to obtain standard vehicle after-sales data. Extraction unit 620 is used to extract predictive features related to vehicle reliability from the standard vehicle after-sales data; The training unit 630 is used to train the prediction features using a machine learning algorithm to obtain a reliability prediction model; The processing unit 640 is used to process the current after-sales data of the target vehicle using the trained reliability prediction model, and output the reliability prediction result of the target vehicle.

[0062] The functions of each unit in the reliability prediction device based on vehicle after-sales data provided in the above embodiments of this application can be implemented through the above methods and steps. Therefore, the specific working process and beneficial effects of each unit in the reliability prediction device based on vehicle after-sales data provided in the embodiments of this application will not be repeated here.

[0063] This application also provides an electronic device, such as... Figure 7As shown, it includes a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740.

[0064] Memory 730 is used to store computer programs; When the processor 710 executes the program stored in the memory 730, it performs the following steps: Obtain vehicle after-sales data from multiple data sources for the same model as the target vehicle during historical periods, and preprocess the vehicle after-sales data from the multiple data sources to obtain standard vehicle after-sales data. Extract predictive features related to vehicle reliability from the standard vehicle after-sales data; A reliability prediction model is obtained by training the predicted features using machine learning algorithms. The trained reliability prediction model is used to process the current after-sales data of the target vehicle and output the reliability prediction result of the target vehicle.

[0065] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0066] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0067] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0068] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0069] The implementation methods and beneficial effects of the various components of the electronic device in the above embodiments for solving the problem can be found in [reference needed]. Figure 2 The steps in the illustrated embodiments are used to implement the electronic device. Therefore, the specific working process and beneficial effects of the electronic device provided in this application will not be repeated here.

[0070] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the 1111 methods described in the above embodiments.

[0071] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the 1111 methods described in the above embodiments.

[0072] Those skilled in the art will understand that the embodiments in this application can be provided as methods, systems, or computer program products. Therefore, the embodiments in this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the embodiments in this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0073] This application describes embodiments of methods, apparatus (systems), and computer program products according to embodiments of this application with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0074] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0075] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0076] Unless otherwise defined, the technical or scientific terms used in this application shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected," "coupled," or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0077] Although preferred embodiments have been described in this application, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the embodiments in this application are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments in this application.

[0078] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the embodiments of this application and their equivalents, then these modifications and variations are also intended to be included in the embodiments of this application.

Claims

1. A reliability prediction method based on vehicle after-sales data, characterized by, The method comprises: acquiring vehicle after-sales data of multiple data sources of the same vehicle model as the target vehicle in a historical stage, and preprocessing the vehicle after-sales data of the multiple data sources to obtain standard vehicle after-sales data; extracting prediction features related to vehicle reliability from the standard vehicle after-sales data; training the prediction features by using a machine learning algorithm to obtain a reliability prediction model; processing current vehicle after-sales data of the target vehicle by using the trained reliability prediction model, and outputting a reliability prediction result of the target vehicle.

2. The method of claim 1, wherein, The preprocessing of the vehicle after-sales data of the multiple data sources to obtain the standard vehicle after-sales data comprises: performing data cleaning on vehicle after-sales data of any data source to obtain initial vehicle after-sales data; performing data conversion on the initial vehicle after-sales data to obtain target vehicle after-sales data; wherein the data conversion comprises encoding processing on non-numeric features and normalization processing on numeric features; associating and integrating target vehicle after-sales data of different data sources to obtain the standard vehicle after-sales data.

3. The method of claim 1, wherein, The extraction of prediction features related to vehicle reliability from the standard vehicle after-sales data comprises: determining a candidate feature set from the standard vehicle after-sales data based on configured vehicle domain data; calculating the contribution of each feature in the candidate feature set to reliability prediction by using a feature importance evaluation method; sorting the candidate features according to the contribution, and selecting features ranked higher than a configured preset threshold as the prediction features.

4. The method of claim 3, wherein, The feature importance evaluation method is a random forest algorithm or an XGBoost algorithm.

5. The method of claim 1, wherein, The training of the prediction features by using a machine learning algorithm to obtain a reliability prediction model comprises: dividing a data set containing the prediction features into a training set and a test set; training a selected machine learning algorithm using the training set, and verifying the performance of a target model by using the test set to obtain a verification result; based on the verification result, performing hyperparameter tuning on the target model to obtain the reliability prediction model.

6. The method of claim 5, wherein, The machine learning algorithm is a gradient boosting decision tree model. The performance of the target model is verified by at least one of accuracy, recall rate, and F1 score.

7. The method of claim 1, wherein, After outputting the reliability prediction result of the target vehicle, the method further comprises: collecting newly added vehicle after-sales data according to a configured collection period, and preprocessing and extracting features from the newly added vehicle after-sales data to obtain new samples; performing incremental learning or periodic retraining on the deployed reliability prediction model by using the new samples to obtain a target reliability prediction model.

8. A reliability prediction device based on vehicle after-sales data, characterized by, The device comprises: an acquisition unit configured to acquire vehicle after-sales data of multiple data sources of the same vehicle model as a target vehicle in a historical stage, and preprocess the vehicle after-sales data of the multiple data sources to obtain standard vehicle after-sales data; an extraction unit configured to extract prediction features related to vehicle reliability from the standard vehicle after-sales data; The training unit is configured to train the prediction feature by using a machine learning algorithm to obtain a reliability prediction model; The processing unit is configured to process the current vehicle after-sales data of the target vehicle by using the trained reliability prediction model, and output a reliability prediction result of the target vehicle.

9. An electronic device, comprising: The electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is configured to store a computer program; The processor is configured to execute the program stored on the memory to implement the method steps of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the method steps of any one of claims 1-7.