A new energy automobile product reliability analysis method based on defect investigation data
By using a reliability analysis method for new energy vehicle products based on defect survey data, and employing a log-normal distribution and parameter adaptive random forest algorithm, suspected defective products can be quickly identified and judged. This solves the problem of long judgment time in new energy vehicle supervision technology and improves supervision efficiency and product reliability.
Patent Information
- Application Number
- CN202411741783.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing new energy vehicle regulatory technologies cannot meet the requirements for regulating large-scale products. They mainly rely on accident investigation information after an accident occurs, resulting in long timeframes for determining defective products and a heavy reliance on professional experience.
Based on defect investigation data, by segmenting complaint record data, a failure probability model based on log-normal distribution and a parameter adaptive random forest network algorithm are constructed to identify suspected defective products, and a defective product discrimination model is constructed for accurate discrimination.
It enables rapid reliability analysis of new energy vehicles, identifies potential failure modes, improves safety performance, optimizes production processes, reduces defect rates, lowers maintenance and repair costs, and extends service life.
Smart Images

Figure CN119669901B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of defective product management, in particular to a new energy vehicle product reliability analysis method and system based on defect investigation data. BACKGROUND
[0002] Defective products are an important form of product quality problems, and are the responsibility and focus of quality supervision departments. In particular, in the field of new energy vehicles, defects in power battery systems, braking systems, acceleration systems, etc. of products can greatly affect people's personal and property safety.
[0003] China is a major country in new energy vehicles, and the number of new energy vehicles in use has exceeded 20 million by 2023. However, due to the rapid pace of technological iteration, some new energy vehicle technologies are put into the market before they have been extensively verified, and their product defect probability and reliability are questionable. Therefore, it is necessary to comprehensively and reliably supervise the safety of new energy vehicle products in operation, recall related defective new energy vehicle products, improve the reliability of new energy vehicle operation, and further expand the market acceptance of new energy vehicles.
[0004] Existing new energy vehicle supervision technologies mainly rely on accident investigation information after accidents occur, resulting in long defect product determination time and serious dependence on professional experience for product defect determination. In the context of the increasing number of new energy vehicles in China, traditional new energy vehicle supervision technologies cannot meet the requirements of large-scale product supervision due to long cycle and slow response. SUMMARY
[0005] To solve the problem that existing new energy vehicle supervision technologies cannot meet the requirements of large-scale product supervision, the present application proposes a new energy vehicle product reliability analysis method based on defect investigation data. After dividing the complaint record data, a failure probability model based on a lognormal distribution is constructed using a heuristic algorithm based on historical product complaint record data, and the actual complaint probability value of the current product is compared with the current product failure probability value calculated by the failure probability model under the same conditions. Then, suspected defective products are determined, and defect investigation is carried out accordingly. A defect product discrimination model based on a parameter adaptive random forest network algorithm is constructed to discriminate defective products, and then reliability analysis and evaluation are realized to provide a basis for vehicle defect recall. The present application also relates to a new energy vehicle product reliability analysis system based on defect investigation data.
[0006] The technical solution of the present application is as follows:
[0007] A new energy vehicle product reliability analysis method based on defect investigation data, characterized by the following steps,
[0008] The complaint record data division and failure probability model construction step divides the complaint record data of several new energy vehicles of the same model according to the time of leaving factory into historical product complaint record data and current product complaint record data, and constructs a failure probability model based on a lognormal distribution including data deletion, mileage interval, product complaint times, and product leaving factory time according to the historical product complaint record data by using a heuristic algorithm.
[0009] The suspected defective product determination step inputs the current product complaint record data into the constructed failure probability model based on a lognormal distribution to calculate a current product failure probability value, compares the actual complaint probability value of the current product calculated and processed according to the current product complaint record data with the current product failure probability value calculated by the failure probability model under the same condition, and determines that the current product is a suspected defective product when the actual complaint probability value of the current product exceeds the set threshold of the current product failure probability value calculated by the failure probability model.
[0010] The defective product discrimination model construction step uses the same model recall vehicle data and the same model non-recall vehicle data of the suspected defective product leaving factory time and mileage to construct a training data set, and further constructs a defective product discrimination model of a parameter adaptive random forest network algorithm including product leaving factory time, mileage interval, and actual complaint probability of the current product.
[0011] The defective product determination step inputs the leaving factory time, mileage interval, and product complaint times of the suspected defective product into the constructed defective product discrimination model to discriminate the defective product, and further realizes reliability analysis.
[0012] Preferably, in the complaint record data division and failure probability model construction step, after the failure probability model based on a lognormal distribution is constructed, the parameters of the two-dimensional lognormal distribution failure probability model are fitted by a bat algorithm, the mean square deviation function of the complaint probability, vehicle production time, and mileage corresponding to the fitting parameters is used as the objective function, the initial position, speed, loudness, and pulse emission frequency of each bat are initialized, and the position and speed of the bat are adjusted to iteratively search the parameter space, each bat evaluates the objective function according to its position, and adjusts the search strategy according to the evaluation result.
[0013] Preferably, in the complaint record data division and failure probability model construction step, the historical product complaint record data and the current product complaint record data both include several complaint types in the power system, braking system, electrical system, and entertainment system of the new energy vehicle.
[0014] Preferably, in the complaint record data division and failure probability model construction step, a heuristic algorithm including grey wolf optimization algorithm, bat algorithm, simulated annealing algorithm, and / or ant colony algorithm is used to construct the failure probability model based on lognormal distribution.
[0015] Preferably, in the defect product discrimination model construction step, when constructing the defect product discrimination model of the parameter adaptive random forest network algorithm, a decision tree algorithm is used, features are randomly selected, a mathematical description of the decision tree prediction is generated, and then a particle swarm-simulated annealing method is used for random forest network parameter adaptive identification; the random forest network parameter adaptive identification includes setting the initial temperature and cooling rate of simulated annealing, particle fitness evaluation, particle swarm speed update, temperature regulation of simulated annealing, and acceptance criteria.
[0016] Preferably, in the defect product discrimination model construction step, features are randomly selected, in the splitting process of each decision tree node, a number of feature subsets are randomly selected from all features to find the best split, for a given input, the decision tree prediction starts from the root node, recursively applies the decision function until a leaf node is reached, a mathematical description of the decision tree prediction is generated, and the final prediction of the random forest is based on the aggregation of the prediction results of all decision trees.
[0017] A new energy vehicle product reliability analysis system based on defect investigation data, characterized by comprising complaint record data division and failure probability model construction module, suspected defect product judgment module, defect product discrimination model construction module, and defect product judgment module connected in sequence,
[0018] The complaint record data division and failure probability model construction module divides the complaint record data of a number of new energy vehicles of the same model into historical product complaint record data and current product complaint record data according to the time of leaving the factory, and uses a heuristic algorithm to construct a failure probability model based on lognormal distribution including data deletion, driving mileage interval, product complaint frequency, and product leaving time according to the historical product complaint record data;
[0019] The suspected defect product judgment module inputs the current product complaint record data into the constructed failure probability model based on lognormal distribution to calculate the current product failure probability value, compares the actual complaint probability value of the current product calculated and processed according to the current product complaint record data with the current product failure probability value calculated by the failure probability model under the same condition, and determines that the current product is a suspected defect product when the actual complaint probability value of the current product exceeds the set threshold of the current product failure probability value calculated by the failure probability model.
[0020] The defect product discrimination model construction module uses the same model recall vehicle data and the same model non-recall vehicle data as the suspected defect product factory time and mileage to construct a training data set, and then constructs a defect product discrimination model including product factory time, mileage interval, and actual complaint probability of the current product.
[0021] The defect product discrimination model is used to input the factory time, mileage interval, and product complaint times of the suspected defect product to discriminate the defect product, and then the reliability analysis is realized.
[0022] Preferably, in the complaint record data division and failure probability model construction module, after the failure probability model based on the lognormal distribution is constructed, the failure probability model parameter fitting of the two-dimensional lognormal distribution is performed by using the bat algorithm, the mean square deviation function of the complaint probability, vehicle production time, and mileage corresponding to the fitting parameters is used as the objective function, the initial position, speed, loudness, and pulse emission frequency of each bat are initialized, and the position and speed of the bat are adjusted to iteratively search the parameter space, each bat evaluates the objective function according to its position, and adjusts the search strategy according to the evaluation result.
[0023] Preferably, in the complaint record data division and failure probability model construction module, the heuristic algorithm including the grey wolf optimization algorithm, the bat algorithm, the simulated annealing algorithm, and / or the ant colony algorithm is used to construct the failure probability model based on the lognormal distribution.
[0024] Preferably, in the defect product discrimination model construction module, when the defect product discrimination model of the parameter adaptive random forest network algorithm is constructed, the decision tree algorithm is used, the features are randomly selected, the mathematical description of the decision tree prediction is generated, and then the random forest network parameter adaptive identification based on the particle swarm-simulated annealing method is performed; the random forest network parameter adaptive identification includes setting the initial temperature and cooling rate of simulated annealing, particle fitness evaluation, particle swarm speed update, temperature regulation of simulated annealing, and acceptance criteria.
[0025] The application has the following beneficial effects:
[0026] The application provides a new energy vehicle product reliability analysis method based on defect investigation data, in the complaint record data division and failure probability model construction step, a plurality of new energy vehicle complaint record data of the same vehicle model are divided into historical product complaint record data (as long-term data) and current product complaint record data (as short-term data) according to the factory time, and a failure probability model based on a logarithmic normal distribution including data deletion, mileage interval, product complaint times and product factory time is constructed by using a heuristic algorithm according to the historical product complaint record data (long-term data), the model describes the relationship between the product factory batch and the product mileage and the product complaint rate, and the failure probability of the new energy vehicle is represented by logarithm in the model. The related parameters to be fitted in the model can be identified by using a heuristic algorithm, so that the establishment of the double-parameter logarithmic normal distribution failure probability model is realized; in the suspected defect product determination step, the short-term data is processed by referring to the long-term data processing method, the actual complaint probability value of the product factory time and the driving range corresponding to the short-term data considering the data deletion is calculated, and the current product failure probability value calculated by the double-parameter logarithmic normal distribution failure probability model under the same condition is compared, if the actual complaint probability value exceeds the current product failure probability value calculated by the model by a certain threshold, the batch product corresponding to the short-term data is considered as a suspected defect product, and the defect investigation is carried out accordingly; in the defect product discrimination model construction step, a product defect discrimination model based on a parameter adaptive random forest algorithm is constructed, the same type of recall vehicle data and the same type of non-recall vehicle data with the same type of suspected defect product factory time and driving range are used as model training data, the data includes the complaint rate of the whole vehicle after deletion processing, and the complaint rate of each subsystem of the whole vehicle, the parameters of the random forest network model can be identified by using particle swarm simulated annealing, and the output of the model is whether the current input data corresponds to a defective product; finally, in the defect product determination step, the factory time, driving range interval and product complaint times of the suspected defect product are input into the constructed defect product discrimination model for accurate discrimination of the defect product, and then the reliability analysis and evaluation are realized. In the massive new energy vehicle supervision data, the suspected defect product is found and the defect investigation is carried out accordingly, the new energy vehicle reliability is evaluated combined with the defect investigation result, the existing new energy vehicle supervision technology mainly depends on the accident investigation information after the accident, the problem that the defect product determination time is long and the product defect determination seriously depends on professional experience is solved, the requirement of large-batch product supervision is met, and the basis for vehicle defect recall is provided.
[0027] The present application is based on new energy vehicle product reliability analysis based on defect investigation data, through the division of different data types, and the construction of failure probability model based on logarithmic normal distribution and product defect discrimination model based on parameter adaptive random forest algorithm respectively by using different algorithm techniques, the determination of suspected defective products is realized, and the defect investigation is carried out accordingly, the multi-level discrimination of defective products is realized, through analysis and test, the potential failure mode is identified, so as to improve the safety performance of the vehicle and reduce the risk of accident occurrence; through collecting and analyzing data, enterprises can optimize production process, improve production efficiency, reduce the rate of unqualified products, and improve market competitiveness; through reliability analysis, design can be optimized, failure rate can be reduced, so as to reduce the maintenance and maintenance cost in later period, and through reliability analysis, the deficiencies in design can be found and solved, so as to prolong the service life of new energy vehicles.
[0028] The present application also relates to a new energy vehicle product reliability analysis system based on defect investigation data, which corresponds to the new energy vehicle product reliability analysis method based on defect investigation data described above, and can be understood as a system for realizing the new energy vehicle product reliability analysis method based on defect investigation data described above, the system comprises complaint record data division and failure probability model construction module, suspected defective product determination module, defect product discrimination model construction module and defect product determination module connected in turn, each module works with each other, after the complaint record data is divided, the failure probability model based on logarithmic normal distribution is constructed by using heuristic algorithm according to historical product complaint record data, and the actual complaint probability value of the current product is compared with the current product failure probability value calculated by the failure probability model under the same condition, and then the suspected defective product is determined, and the defect investigation is carried out accordingly, the defect product discrimination model of parameter adaptive random forest network algorithm is constructed to discriminate the defect product, and then the reliability analysis and evaluation are realized, and the basis for vehicle defect recall is provided. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 It is the flowchart of the new energy vehicle product reliability analysis method based on defect investigation data of the present application.
[0030] Figure 2 It is the structure block diagram of the new energy vehicle product reliability analysis system based on defect investigation data of the present application. DETAILED DESCRIPTION
[0031] The present application will be described below in combination with the drawings.
[0032] The application relates to a new energy automobile product reliability analysis method based on defect investigation data. First, complaint record data is divided, and a failure probability model based on a double-parameter logarithmic normal distribution is established according to historical product complaint records, and a heuristic optimization algorithm is used to identify model parameters. Secondly, according to the current product complaint record, it is judged whether the current product failure probability conforms to the failure probability model, and the product seriously deviating from the implementation probability is determined as a suspected defect product. Thirdly, a parameter adaptive random forest network algorithm defect product discrimination model considering product shipment time, mileage and complaint rate is constructed. Finally, the shipment time, mileage and complaint rate of the suspected defect product are input into the defect product discrimination model to realize the discrimination of the defect product. The application can play an important supporting role for new energy automobile defect recall. The flowchart of the method is shown in Figure 1 The method comprises the following steps.
[0033] S1: complaint record data division and failure probability model construction step, a plurality of new energy automobile complaint record data of the same vehicle model is divided into historical product complaint record data and current product complaint record data according to the shipment time, and a failure probability model based on a logarithmic normal distribution including data deletion, mileage interval, product complaint times and product shipment time is constructed by using a heuristic algorithm according to the historical product complaint record data. The model describes the relationship between the product shipment batch and the product mileage and the product complaint rate, and the failure probability of the new energy automobile is expressed by using a logarithm in the model. The divided historical product complaint record data and current product complaint record data both include complaint types of a power system, a braking system, an electrical system, an entertainment system and the like of the new energy automobile. Further, the step can construct the failure probability model based on the logarithmic normal distribution by using a heuristic algorithm including a grey wolf optimization algorithm, a bat algorithm, a simulated annealing algorithm and / or an ant colony algorithm.
[0034] S11, data division
[0035] Based on the existing past complaint data of a vehicle, the data is divided into product complaints within three months (current product complaint record data, also short-term data) and product complaints outside three months (historical product complaint record data, also long-term data) with three months as the dividing line.
[0036] S12, data processing
[0037] The long-term data complaint is inductively processed, and the repeated complaint of the same vehicle in a certain mileage range to a certain system is deleted, that is, only one complaint record data is reserved for the complained system in a certain mileage range, and is summarized into a table as shown in Table 1:
[0038] Table 1
[0039]
[0040] S13, constructing a failure probability model based on a lognormal distribution
[0041] First, the product failure complaint probability is calculated, and the calculation method is as follows:
[0042]
[0043] In the above formula, P m,L represents the complaint rate of the product within the range of the factory time m and the range of the mileage L; x m,L represents the total number of complaints of the product within the range of the factory time m and the range of the mileage L; N m,L represents the total number of samples within the range of the factory time m and the range of the mileage L.
[0044] Based on the three-month product complaint probability calculated in the above formula, a two-dimensional parameter lognormal distribution failure probability model based on vehicle production time and mileage is constructed. The lognormal distribution failure probability model has the following characteristics:
[0045] ① Handle non-negative data: Lognormal distribution is used for model positive data (such as price, income, length, etc.), which cannot be negative naturally.
[0046] ② Handle skewed data: For right-skewed or long-tailed distribution data, lognormal distribution provides better fitting. It can effectively handle data sets that have mostly small values but also contain some extreme values.
[0047] ③ Multiplicative effect: Lognormal distribution is suitable for cases where multiple independent factors multiply to affect the result, such as some types of economic, biological or engineering data.
[0048] ④ Linearization effect of logarithmic transformation: Through logarithmic transformation, the multiplicative relationship can be converted into additive relationship, simplifying the complexity of the model.
[0049] Assuming that the complaint probability of the product is P, and the complaint rate is positively correlated with the average vehicle factory time x1 and the average mileage x2, then the two-parameter lognormal distribution mathematical model construction method is as follows:
[0050] Define the natural logarithm of the complaint probability P, x1, and x2 as:
[0051]
[0052] Where, μ P , is the mean of the log variable ln(P), ln(x1), and ln(x2), and For the corresponding variance, N is the normal distribution.
[0053] Since P depends on the vehicle production time x1 and the driving mileage x2, then: P = f(x1, x2)
[0054] After taking the logarithm, we get: ln(P) = g(ln(x1), ln(x2))
[0055] Where g() represents a linear function, then the above formula can be written as the following lognormal distribution failure probability model:
[0056] ln(P) = a + b·ln(x1) + c·ln(x2) + d·ln(x1)·ln(x2)
[0057] Where a, b, c, d are the parameters to be fitted, and parameter d is the multiplicative parameter between x1 and x2.
[0058] S14, bat algorithm parameter fitting
[0059] After constructing the failure probability model based on the lognormal distribution, the bat algorithm is used to fit the two-dimensional lognormal distribution failure probability model parameters. The mean square deviation function of the complaint probability, vehicle production time and driving mileage corresponding to the to-be-fitted parameters is used as the objective function. The initial position, speed, loudness and pulse emission frequency of each bat are initialized, and the position and speed of the bat are adjusted to iteratively search the parameter space. Each bat evaluates the objective function according to its position, and adjusts its search strategy according to the evaluation result.
[0060] Bat Algorithm (BA) is a swarm intelligence-based optimization algorithm, which is inspired by the behavior of bats using echolocation when hunting at night. The bat algorithm is mainly used to solve optimization problems. The bat algorithm optimization process mainly includes the following steps:
[0061] ① Define the objective function: the objective function is a function used to evaluate the parameter identification effect, which is usually the difference between the actual output and the model output (such as mean square error).
[0062] ② Initialization: set the size of the bat colony, the initial position (i.e. the initial estimate of the parameter) and the speed of each bat, as well as related algorithm parameters such as frequency range, loudness and pulse emission rate.
[0063] ③ Iterative search: in each iteration, the position and speed of the bat are adjusted to search the parameter space. Each bat evaluates the objective function according to its position, and adjusts its search strategy according to the result.
[0064] ④ Update parameters: update the parameter estimate according to the position of the bat. Usually, the position with the best objective function value is selected as the current optimal solution.
[0065] V. Termination condition: the algorithm stops when a preset number of iterations is reached or the objective function reaches a certain threshold.
[0066] VI. Output result: the optimal estimate of the parameters.
[0067] The parameter fitting method for the two-dimensional lognormal distribution failure probability model based on the bat algorithm is as follows:
[0068] I. Define the optimization function using the mean square error:
[0069]
[0070] In the above formula, N represents the number of points to be fitted; P i ,x 1,i ,x 2,i represent the complaint probability and vehicle production time and mileage corresponding to the i-th parameter to be fitted, respectively.
[0071] II. Bat parameter initialization:
[0072] Assume that the bat colony size is M, i.e., there are M bats participating in parameter optimization;
[0073] The position of each bat is Pos j = (a j ,b j ,c j ,d j ), and the speed is v j = (v aj ,v bj ,v cj ,v dj ), where j represents the j-th bat; the loudness A j and the pulse emission frequency r j of each bat are initialized, where the loudness A j is a four-dimensional vector, and the pulse emission frequency r j is a number.
[0074] III. Iterative search
[0075] i) Update the speed and position of each bat:
[0076]
[0077] In the above formula, v represents the speed of the j-th bat at time t; represents the speed at time t-1; Pos best represents the position of the bat closest to the optimal solution at present; and These represent the positions of the j-th bat at time t and time t-1, respectively. f represents the sound wave frequency of bat j at time t; min f max These represent the maximum and minimum frequencies of the sound wave, respectively, where β∈[0,1].
[0078] ii) Generate random numbers and perform a local search:
[0079] If the random number ε ∈ [-1, 1] is greater than the pulse transmission frequency r, then... j If so, the bat enters a local search and its position is updated, as shown in the following update equation:
[0080]
[0081] iii) Update loudness and pulse generation frequency:
[0082] If bat j finds a better solution, then update the loudness A of bat j. j and pulse emission frequency r j :
[0083]
[0084] In the above formula, α and γ represent the loudness and the attenuation constant of the pulse frequency, respectively. This is the initial value of the pulse frequency.
[0085] iv)Pos best Update until the convergence condition is met.
[0086] S2: Suspected defective product determination step: Input the current product complaint record data into the constructed failure probability model based on log-normal distribution to calculate the current product failure probability value, and compare the actual complaint probability value of the current product calculated based on the current product complaint record data with the current product failure probability value calculated by the failure probability model under the same conditions. If the actual complaint probability value of the current product exceeds the set threshold of the current product failure probability value calculated by the failure probability model, the current product is determined to be a suspected defective product.
[0087] This step essentially involves initial screening of suspected defective products based on a failure probability model using a log-normal distribution. Specifically, product complaints within the past three months can be calculated on a monthly basis, using the average vehicle manufacturing time x within this dataset. 1,3 and average mileage x 2,3 , and x 1,3 ,x 2,3 The corresponding actual complaint probability P3 value. Let x 1,3 ,x 2,3into the failure probability model based on the lognormal distribution established in S1, and the P value corresponding to x 1,3 ,x 2,3 P3 is greater than P by a certain threshold, the product is judged as a suspected defective product, and batch investigation is carried out for the product.
[0088] S3: Defective product discrimination model construction step, using the same model recall vehicle data and the same model non-recall vehicle data with the same time and mileage of the suspected defective product to construct the training data set, and then constructing the defective product discrimination model including product factory time, mileage interval and actual complaint probability of the product. The product defect discrimination algorithm model includes DBSCAN, k-means, decision tree algorithm, etc.
[0089] This step is used to establish a defective product discrimination model based on random forest algorithm. Random forest is an ensemble learning method based on machine learning model (non-linear tree-based model). In the 1980s, Breiman et al. invented classification tree algorithm, which greatly reduced the calculation amount by repeatedly dividing the data for classification or regression. In 2001, Breiman combined classification trees into random forests, that is, randomization was performed on the use of variables (columns) and data (rows), and many classification trees were generated, and the results of classification trees were summarized. Random forest improves prediction accuracy without significantly increasing computational complexity. Random forest is not sensitive to multicollinearity, and the results are relatively stable for missing data and unbalanced data. It can well predict the effect of up to thousands of explanatory variables, and is known as one of the best algorithms currently.
[0090] The defective product discrimination model based on random forest algorithm is constructed and the training data set is as follows:
[0091] 1. Constructing training data set:
[0092] Using the same model recall vehicle data and the same model non-recall vehicle data with the same time and mileage of the suspected defective product to construct the training data set, the data content contained in the data set is shown in Table 2:
[0093] Table 2
[0094]
[0095]
[0096] 2. Bootstrap sampling:
[0097] Bootstrap sampling is a crucial step in the random forest algorithm, used to generate the dataset for training each tree. Mathematically, bootstrap sampling can be described by the following process:
[0098] The original dataset D contains N independent and identically distributed samples, i.e., D = {(x1, y1), (x2, y2), ..., (x... M ,y M )}, where each sample contains a feature vector x i and the corresponding label y i For each decision tree T k Use the following steps to generate the training dataset D k :
[0099] ①. Initialization, let Dk be an empty set: D k ={}
[0100] ②. Repeated sampling: For i = 1 to M (usually, D k (Set the size to be the same as the original dataset D):
[0101] i) Randomly select a sample (x) from D. * ,y * This allows the same sample to be selected multiple times (i.e., sampling with replacement).
[0102] ii) Select the sample (x) * ,y * Put it into D k middle.
[0103] 3. Randomly select features:
[0104] Randomly selecting features when constructing each decision tree in a random forest is a key step in improving model diversity. This process ensures that each tree does not simply replicate the decision paths of other trees, thereby enhancing the overall model's adaptability and generalization ability to patterns not explicitly shown in the data. In this case, the features used to construct the defective product discrimination model based on random forest include the number of sample vehicles N, the total number of complaints X, and the number of braking system complaints x. B Number of complaints about the power system (x) A Number of complaints about human-computer interaction systems (x) H .
[0105] The step can adopt a decision tree algorithm when constructing a defective product discrimination model of a parameter adaptive random forest network algorithm, randomly selects features, and generates a mathematical description of a decision tree prediction, randomly selects features, in each decision tree node splitting process, randomly selects a feature subset from all features to find the best split, for a given input, the decision tree prediction starts from the root node, recursively applies the decision function until a leaf node is reached, generates a mathematical description of the decision tree prediction, the final prediction of the random forest is based on the aggregation of the prediction results of all decision trees; the random forest network parameter adaptive identification based on the particle swarm-simulated annealing method; the random forest network parameter adaptive identification includes setting the initial temperature and cooling rate of simulated annealing, particle fitness evaluation, particle swarm speed update, temperature regulation of simulated annealing, and acceptance criteria.
[0106] In each decision tree node splitting process, m feature subsets are randomly selected from all possible features to find the best split. Let the original feature set be: F = {f1, f2,..., f p}
[0107] Where f p represents the pth feature in the original data, p = 5 in this case, and m ≤ p.
[0108] Then for any node, the selected feature subset F s can be represented as:
[0109] The selection of F s is independent for each node.
[0110] 4. Decision tree prediction:
[0111] The prediction process of a decision tree can be regarded as a mapping from the input feature space to the class label. For a given input data x, the prediction function h k (x) of the decision tree T k can be divided into a series of applications of decision rules, which guide the input vector to a specific leaf node in the tree based on the feature values of the input vector. At the leaf node, the decision tree outputs a prediction value, which is determined according to the distribution of samples in the leaf node during training.
[0112] ① Mathematical description of decision tree prediction
[0113] Let N be the set of nodes in the decision tree, and each node n ∈ N is associated with a decision function d n (x). This function determines which child node to move to next based on the feature values of the input x. The decision function can be formalized as:
[0114]
[0115] In the above formula, n left represents the selection of the left node, n right represents the selection of the right node; x f represents the corresponding value of the feature f of the input vector x, and θ is a predefined threshold value.
[0116] ② Prediction process
[0117] For a given input x, the prediction process of a decision tree is a process of starting from the root node and recursively applying the decision function d n (x) until a leaf node is reached. Then, the prediction value of the leaf node is returned as the prediction output y l of the entire decision tree, where the prediction function h k (x) can be represented as: h k (x) = y l
[0118] 5. Aggregated prediction:
[0119] The final prediction of a random forest is based on the aggregation of the prediction results of all its decision trees. For classification problems, the majority vote method is usually used:
[0120]
[0121] In the above formula represents the voting result, the mode function returns the most frequent value in a set of values (i.e. the majority vote), and K is the total number of decision trees.
[0122] 6. Random forest network parameter adaptive identification based on particle swarm-simulated annealing method
[0123] ①. Algorithm initialization:
[0124] Randomly initialize the position and velocity of the particle swarm, and set the initial temperature and cooling rate of simulated annealing.
[0125] ②. Particle fitness evaluation:
[0126] i) Define the fitness function:
[0127] According to the actual situation and the detection result, the value scoring method for each detection result of the DBSCAN algorithm is shown in Table 3:
[0128] Table 3
[0129]
[0130] ii) Based on the accuracy score of each sample, the algorithm comprehensive accuracy scoring is carried out, i.e. the algorithm fitness function definition, and the scoring calculation formula is as follows:
[0131]
[0132] where N represents the total number of sample detections performed, score i represents the accuracy score of the i-th sample detection, score avg represents the comprehensive accuracy score of the algorithm, and the smaller the value, the better the accuracy of the algorithm.
[0133] ③ Particle swarm velocity update:
[0134] The velocity of each particle in the particle swarm is updated, and the update formula is:
[0135]
[0136] In the above formula, is the velocity of particle i in dimension d at time t; is the position of particle i in dimension d at time t, w is the inertia weight used to control the persistence of particle velocity, c1 and c2 are learning factors used to adjust the movement of particles to the individual optimal position and global optimal position, and r1 and r2 are random numbers between 0 and 1 used to increase the randomness of the search.
[0137] ④ Temperature regulation and acceptance criteria of simulated annealing:
[0138] In PSO-SA, the temperature regulation and acceptance criteria of simulated annealing (SA) are used in the process of updating the velocity and position of the particle, so that there is a certain probability to accept the new position when the fitness of the new position is poor:
[0139] Temperature regulation formula: T (t+1) = α·T (t)
[0140] Metropolis acceptance criterion: P(ΔE, T) = exp(-ΔE / (k·T))
[0141] The above formula T (t) is the temperature at time t, α is the cooling rate, which is 0 to 1, ΔE is the difference in fitness between the new position and the current position, k is the Boltzmann constant, and P(ΔE, T) is the probability of accepting a solution with poor fitness at temperature T.
[0142] S4: Defective product determination step, input the suspected defective product's time of leaving factory, driving mileage interval and product complaint times into the constructed defective product discrimination model to discriminate the defective product, and then realize reliability analysis.
[0143] This step involves statistically organizing the suspected defective batches of products output in S2, and then using the statistical results as input to the defective product discrimination model based on the random forest algorithm in S3, thereby achieving accurate identification of defective products. The data statistical organization format is shown in Table 4:
[0144] Table 4
[0145]
[0146] This invention also relates to a reliability analysis system for new energy vehicle products based on defect survey data. Corresponding to the aforementioned reliability analysis method for new energy vehicle products based on defect survey data, it can be understood as a system that implements the aforementioned reliability analysis method for new energy vehicle products based on defect survey data. Figure 2 As shown, the system includes a complaint record data segmentation and failure probability model construction module, a suspected defective product judgment module, a defective product discrimination model construction module, and a defective product judgment module connected in sequence. Specifically, the complaint record data segmentation and failure probability model construction module segments complaint record data for several new energy vehicles of the same model into historical product complaint record data and current product complaint record data according to their manufacturing date. Based on the historical product complaint record data, it uses a heuristic algorithm to construct a failure probability model based on a log-normal distribution, including data censoring, mileage range, number of product complaints, and product manufacturing date. The suspected defective product judgment module inputs the current product complaint record data into the constructed log-normal distribution-based failure probability model to calculate the current product failure probability value, and processes the current product complaint record data to obtain the current product failure probability value. The actual complaint probability of a product is compared with the failure probability of the current product calculated by the failure probability model under the same conditions. When the actual complaint probability of the current product exceeds a set threshold calculated by the failure probability model, the current product is determined to be a suspected defective product. The defective product discrimination model construction module uses data from recalled vehicles of the same model with the same manufacturing time and mileage as the suspected defective product, as well as data from non-recalled vehicles of the same model, to construct a training dataset. This dataset is then used to construct a defective product discrimination model using a parameterized adaptive random forest network algorithm that includes the product's manufacturing time, mileage range, and the actual complaint probability of the current product. The defective product judgment module inputs the manufacturing time, mileage range, and number of product complaints of the suspected defective product into the constructed defective product discrimination model to judge the defective product, thereby achieving reliability analysis.
[0147] Preferably, in the complaint record data division and failure probability model construction module, after the failure probability model based on the lognormal distribution is constructed, the failure probability model parameter fitting of two-dimensional lognormal distribution is performed through the bat algorithm, the mean square deviation function of the complaint probability, the vehicle production time and the driving mileage corresponding to the fitting parameter is taken as the objective function, the initial position, the speed, the loudness and the pulse emission frequency of each bat are initialized, and the parameter space is iteratively searched by adjusting the position and the speed of the bat, and each bat evaluates the objective function according to the position thereof and adjusts the search strategy according to the evaluation result.
[0148] Preferably, in the complaint record data division and failure probability model construction module, the failure probability model based on the lognormal distribution is constructed by using a heuristic algorithm including the grey wolf optimization algorithm, the bat algorithm, the simulated annealing algorithm and / or the ant colony algorithm.
[0149] Preferably, in the defect product discrimination model construction module, when the defect product discrimination model of the parameter adaptive random forest network algorithm is constructed, the decision tree algorithm is adopted, the features are randomly selected, the mathematical description of the decision tree prediction is generated, and the random forest network parameter adaptive identification based on the particle swarm-simulated annealing method is performed; the random forest network parameter adaptive identification includes setting the initial temperature and the cooling rate of the simulated annealing, the particle fitness evaluation, the particle swarm speed updating, the temperature regulation of the simulated annealing and the acceptance criterion.
[0150] The new energy automobile product reliability analysis method and system based on defect investigation data provided by the application sequentially perform complaint record data division, failure probability model construction, suspected defect product determination, defect product discrimination model construction and defect product determination, find suspected defect products in massive new energy automobile supervision data, and carry out defect investigation in a targeted manner, evaluate the reliability of the new energy automobile in combination with the defect investigation result, solve the problem that the existing new energy automobile supervision technology mainly depends on the accident investigation information after the occurrence of an accident, leading to a long defect product determination time and a serious dependence of product defect determination on professional experience, and meet the requirements of large-batch product supervision and provide a basis for vehicle defect recall.
[0151] It should be noted that the above specific embodiments can enable those skilled in the art to more fully understand the present application, but in no way limit the present application. Therefore, although the present application has been described in detail with reference to the drawings and examples, those skilled in the art should understand that the present application can still be modified or equivalently replaced, in short, all technical solutions and improvements that do not deviate from the spirit and scope of the present application should be covered in the protection scope of the patent of the present application.
Claims
1. A reliability analysis method for new energy vehicle products based on defect survey data, characterized in that, Includes the following steps, The steps for dividing complaint record data and constructing a failure probability model are as follows: Based on the manufacturing time, the complaint record data of several new energy vehicles of the same model are divided into historical product complaint record data and current product complaint record data. Based on the historical product complaint record data, a failure probability model based on log-normal distribution is constructed using a heuristic algorithm, including data censoring, driving mileage range, number of product complaints, and product manufacturing time. The suspected defective product determination process involves inputting the current product complaint record data into a failure probability model based on a log-normal distribution to calculate the current product failure probability value. The actual complaint probability value of the current product calculated based on the current product complaint record data is then compared with the failure probability value of the current product calculated by the failure probability model under the same conditions. If the actual complaint probability value of the current product exceeds a set threshold for the failure probability value of the current product calculated by the failure probability model, the current product is determined to be a suspected defective product. The steps for building a defective product discrimination model are as follows: a training dataset is constructed using data from recalled vehicles of the same model with the same manufacturing time and mileage as the suspected defective product, as well as data from non-recalled vehicles of the same model. Then, a defective product discrimination model based on a parameter adaptive random forest network algorithm is constructed, which includes the product manufacturing time, mileage range, and the actual probability of complaints against the current product. The defective product identification process involves inputting the manufacturing date, mileage range, and number of product complaints of suspected defective products into a constructed defective product identification model to identify defective products and thus achieve reliability analysis.
2. The reliability analysis method for new energy vehicle products according to claim 1, characterized in that, In the steps of data segmentation of complaint records and construction of failure probability model, after constructing a failure probability model based on log-normal distribution, the parameters of the two-dimensional log-normal distribution failure probability model are fitted by the bat algorithm. The standard deviation function of the complaint probability, vehicle production time and driving mileage corresponding to the parameters to be fitted is used as the objective function. The initial position, speed, loudness and pulse emission frequency of each bat are initialized. The parameter space is iteratively searched by adjusting the position and speed of the bat. Each bat evaluates the objective function according to its position and adjusts the search strategy according to the evaluation results.
3. The reliability analysis method for new energy vehicle products according to claim 1, characterized in that, In the steps of classifying complaint record data and constructing the failure probability model, the classified historical product complaint record data and current product complaint record data both include several complaint types in the power system, braking system, electrical system, and entertainment system of new energy vehicles.
4. The reliability analysis method for new energy vehicle products according to claim 1, characterized in that, In the steps of segmenting complaint record data and constructing the failure probability model, a failure probability model based on the log-normal distribution is constructed using heuristic algorithms including the gray wolf optimization algorithm, the bat algorithm, the simulated annealing algorithm, and / or the ant colony algorithm.
5. The reliability analysis method for new energy vehicle products according to any one of claims 1 to 4, characterized in that, In the defective product discrimination model construction step, when constructing the defective product discrimination model using the parameter adaptive random forest network algorithm, a decision tree algorithm is used to randomly select features and generate a mathematical description of the decision tree prediction. Then, the random forest network parameters are adaptively identified based on the particle swarm-simulated annealing method. The adaptive identification of random forest network parameters includes setting the initial temperature and cooling rate of simulated annealing, particle fitness evaluation, particle swarm velocity update, temperature control of simulated annealing, and acceptance criteria.
6. The reliability analysis method for new energy vehicle products according to claim 5, characterized in that, In the defective product discrimination model construction steps, features are randomly selected. During the splitting process of each decision tree node, several feature subsets are randomly selected from all features to find the optimal split. For a given input, the decision tree prediction starts from the root node and recursively applies the decision function until a leaf node is reached, generating a mathematical description of the decision tree prediction. The final prediction of the random forest is based on the aggregation of the prediction results of all its decision trees.
7. A reliability analysis system for new energy vehicle products based on defect survey data, characterized in that, This includes a sequentially connected module for classifying complaint record data and constructing a failure probability model, a module for identifying suspected defective products, a module for constructing a defective product discrimination model, and a module for determining defective products. The complaint record data segmentation and failure probability model construction module divides the complaint record data of several new energy vehicles of the same model into historical product complaint record data and current product complaint record data according to the manufacturing time. Based on the historical product complaint record data, a failure probability model based on log-normal distribution is constructed using a heuristic algorithm, including data censoring, driving mileage range, number of product complaints and product manufacturing time. The suspected defective product determination module inputs the current product complaint record data into the constructed failure probability model based on log-normal distribution to calculate the current product failure probability value, and compares the actual complaint probability value of the current product calculated based on the current product complaint record data with the current product failure probability value calculated by the failure probability model under the same conditions. When the actual complaint probability value of the current product exceeds the set threshold of the current product failure probability value calculated by the failure probability model, the current product is determined to be a suspected defective product. The defective product discrimination model construction module uses data from recalled vehicles of the same model with the same manufacturing time and mileage as the suspected defective product, as well as data from non-recalled vehicles of the same model, to construct a training dataset. This dataset is then used to construct a defective product discrimination model based on a parameter adaptive random forest network algorithm, which includes the product manufacturing time, mileage range, and the actual probability of complaints against the current product. The defective product identification module inputs the manufacturing time, mileage range, and number of product complaints of suspected defective products into the constructed defective product identification model to identify defective products, thereby achieving reliability analysis.
8. The new energy vehicle product reliability analysis system according to claim 7, characterized in that, In the complaint record data segmentation and failure probability model construction module, after constructing a failure probability model based on a log-normal distribution, the parameters of the two-dimensional log-normal distribution failure probability model are fitted using the bat algorithm. The standard deviation function of the complaint probability, vehicle production time, and driving mileage corresponding to the parameters to be fitted is used as the objective function. The initial position, speed, loudness, and pulse emission frequency of each bat are initialized. The parameter space is iteratively searched by adjusting the position and speed of the bats. Each bat evaluates the objective function according to its position and adjusts the search strategy according to the evaluation results.
9. The new energy vehicle product reliability analysis system according to claim 7, characterized in that, In the complaint record data partitioning and failure probability model construction module, a failure probability model based on log-normal distribution is constructed using heuristic algorithms including the gray wolf optimization algorithm, bat algorithm, simulated annealing algorithm, and / or ant colony algorithm.
10. The new energy vehicle product reliability analysis system according to any one of claims 7 to 9, characterized in that, In the defective product discrimination model construction module, when constructing the defective product discrimination model based on the parameter adaptive random forest network algorithm, a decision tree algorithm is used to randomly select features and generate a mathematical description of the decision tree prediction. Then, the random forest network parameters are adaptively identified based on the particle swarm-simulated annealing method. The adaptive identification of random forest network parameters includes setting the initial temperature and cooling rate of simulated annealing, particle fitness evaluation, particle swarm velocity update, temperature control of simulated annealing, and acceptance criteria.
Citation Information
Patent Citations
Nuclear power quality defect cause analysis method
CN114969267A
Easy-to-complaint user prediction method and system and storage medium
CN118364336A