Automobile key component performance degradation intelligent diagnosis large model and prediction method and device
By using Pearson correlation coefficient and grey relational analysis to screen features, and combining multi-head attention mechanism to separate and fuse features, the problem of feature extraction and model selection in lithium-ion battery capacity prediction is solved, thereby improving the accuracy and reliability of prediction.
Patent Information
- Application Number
- CN202511003679.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies for predicting lithium-ion battery capacity suffer from insufficient feature extraction, neglect of multi-dimensional correlations, and difficulty in capturing nonlinear patterns through model selection, resulting in insufficient prediction accuracy and reliability.
We use Pearson correlation coefficient and grey relational analysis to screen multi-dimensional features, construct a low-redundancy and highly interpretable feature set, and combine multi-head attention mechanism to separate and fuse trend and seasonal features to build a capacity prediction model.
It improves the accuracy and reliability of battery capacity prediction, enhances the model's generalization ability, and can better adapt to different operating conditions and types of battery data.
Smart Images

Figure CN120910440A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of Internet big data and new generation information technology, and in particular to a vehicle key component performance degradation intelligent diagnosis large model and a prediction method and device. BACKGROUND
[0002] In order to reduce dependence on oil and pollution to the environment, many countries have accelerated the development of electric vehicles. Lithium-ion batteries (LIBS) have become the main energy storage device in electric vehicle applications due to their high energy density, reasonable cycle life and recyclability. However, the internal components of lithium-ion batteries inevitably degrade to varying degrees during the continuous charging and discharging process, resulting in a decrease in performance. In order to ensure the safe operation, reliability and efficiency of lithium-ion batteries, it is extremely important to accurately assess their state of health (SOH).
[0003] Battery SOH can be defined by cycle life, internal resistance increase and capacity loss. Capacity directly determines the energy that can be stored and released by the battery, and it is more important for the management and maintenance of the battery system. Therefore, the definition of SOH as the ratio of the current actual capacity to the initial rated capacity has received more attention. However, the measurement and calculation of battery capacity depend on a complete charging and discharging process, which is challenging for batteries in actual use. In recent years, a large number of studies have been conducted to predict battery capacity degradation, and can be roughly divided into model-based and data-driven methods.
[0004] Model-based methods require the establishment of a mathematical model to describe the degradation of battery capacity, typically including empirical models and electrochemical models. The capacity loss of the battery is fitted using an exponential function or a polynomial function, and then the battery capacity is extrapolated according to calendar time, cycle number or driving distance to obtain an empirical model. However, a large number of orthogonal experiments are usually required to consider the effects of different influencing factors, and the use of pre-defined parameter functions to fit the degradation curve is limited under variable operating conditions. Electrochemical models can describe the degradation process inside the battery at the microscopic level, and typical degradation mechanisms include the growth of the solid electrolyte interface (SEI) and the plating of lithium. After the degradation mechanism of the battery under study is determined, the future capacity degradation of the battery can be derived through simulation. However, due to the coupling of side reactions inside the battery, the complex degradation process is difficult to determine and quantify. In addition, as the battery ages, changes in the parameters of the electrochemical model can result in large capacity prediction errors.
[0005] In recent years, with the generation and accumulation of a large amount of battery data and the penetration of artificial intelligence technology, data-driven methods have been widely applied in battery capacity prediction. Unlike model-based methods, data-driven methods can draw the relationship between several features and battery health indicators without needing to know the mechanism and propagation accurately, thereby avoiding the complex modeling process.
[0006] However, the existing data-driven methods still have some problems in feature extraction and model selection. In terms of feature extraction, most methods only consider single-dimensional features of battery charging data, ignoring the correlation and comprehensiveness between multi-dimensional features, resulting in insufficient representation of battery capacity changes and affecting the accuracy of subsequent prediction. In terms of model selection, some models are difficult to capture the nonlinear rules of battery capacity changes over time, have limited ability to learn long-distance dependencies when processing time series data, and are prone to overfitting or underfitting, which cannot achieve long-term and accurate prediction of battery capacity.
[0007] In summary, there is an urgent need to design a more scientific and effective intelligent diagnosis large model and prediction method for performance degradation of key components of vehicles to improve the accuracy and reliability of battery capacity prediction. SUMMARY
[0008] To solve the above problems of the prior art, the technical problem to be solved by the present application is to provide an intelligent diagnosis large model and prediction method for performance degradation of key components of vehicles, which realizes feature selection through Pearson correlation coefficient and grey correlation degree, constructs a low-redundancy and high-interpretable feature set, and improves the representation ability of the feature set for battery performance degradation. At the same time, the prediction model divides non-time and time features and fuses them, separates trend and seasonal features, and combines a multi-head attention mechanism to realize capacity sequence prediction of future continuous time steps, thereby improving the accuracy and reliability of battery capacity prediction.
[0009] To solve the above technical problems, the present application adopts the following technical solutions:
[0010] The intelligent diagnosis large model and prediction method for performance degradation of key components of vehicles comprises:
[0011] S1: obtaining standard capacities corresponding to each group of battery charging data in a historical charging data set;
[0012] S2: extracting multi-dimensional statistical features of each group of battery charging data, and selecting an optimal feature set based on the combination of Pearson correlation coefficient and grey correlation degree and labeled capacity;
[0013] S3: inputting the optimal feature set into a trained capacity prediction model to output a predicted capacity sequence of future continuous time steps;
[0014] The processing steps of the capacity prediction model include:
[0015] S301: dividing the optimal feature set into non-time features and time features, and generating initial fusion features by fusion;
[0016] S302: separate the initial fusion feature into trend item feature and seasonal item feature;
[0017] S303: model the trend item feature and the seasonal item feature respectively, and fuse to generate trend-seasonal fusion feature;
[0018] S304: extract the local feature of the trend-seasonal fusion feature, and calculate the dependency relationship of each time step to other time steps through the multi-head attention mechanism to obtain the attention fusion feature;
[0019] S305: after the attention fusion feature is processed by pooling, it is input into a fully connected layer for prediction to obtain the predicted capacity sequence of the future continuous time steps.
[0020] Preferably, in step S1, the following steps are specifically included:
[0021] S101: calculate the labeled capacity based on the battery charging data through the improved current integration method;
[0022] The formula is:
[0023]
[0024] In the formula: C represents the labeled capacity; Δt represents the fixed sampling interval; I(t) represents the charging current; t1 and t2 represent the charging start and end times; SOC t1 , SOC t2 represent the state of charge at the start and end of charging;
[0025] S102: use the sliding window statistical method to process the mean or median of the labeled capacity of each battery charging data, and exclude the abnormality caused by SOC error or data noise;
[0026] The formula is:
[0027]
[0028] In the formula: Ck represents the labeled capacity of the kth battery charging data; C i represents the kth labeled capacity in the statistical window; N is the total amount of data in the statistical window; k represents the statistical period number.
[0029] Preferably, in step S2, the following steps are specifically included:
[0030] S201: perform multi-dimensional statistical analysis on the feature sequence and the labeled capacity sequence of the battery charging data in a period, and calculate the basic statistics including arithmetic mean, cumulative total and standard deviation;
[0031] S202: Calculate the Pearson correlation coefficient ρ of the feature and the labeling capacity based on the ratio of the covariance and the standard deviation of the feature sequence and the labeling capacity sequence xi ;
[0032] The formula is expressed as:
[0033]
[0034] In the formula: ρ xi represents the Pearson correlation coefficient of the i-th feature x i and the labeling capacity; c represents the labeling capacity; and are the average values of the feature sequence and the labeling capacity sequence respectively;
[0035] S203: Calculate the gray correlation coefficient r of the feature and the labeling capacity by constructing the correlation function of the feature sequence and the labeling capacity sequence i ;
[0036] The formula is expressed as:
[0037]
[0038] In the formula: r i represents the gray correlation coefficient of the i-th feature x i and the labeling capacity; n represents the number of features, ξ i (k) is the gray correlation coefficient of the i-th feature x i at the k-th moment; ρ is the discrimination coefficient; x i (k) represents the i-th feature at the k-th moment; c i (k) is the i-th labeling capacity at the k-th moment;
[0039] S204: Select the feature with a Pearson correlation coefficient greater than the first coefficient with the labeling capacity, or a gray correlation coefficient greater than the second coefficient with the labeling capacity as the optimal feature.
[0040] Preferably, step S301 specifically comprises the following steps:
[0041] S3011: Divide the optimal feature set into non-time features X feature and time features X time ;
[0042] S3012: Linearly project the non-time features to obtain the first features F embed ;
[0043] The formula is expressed as:
[0044] F embed = W f X feature+b f ;
[0045] wherein: represents a learnable projection matrix; bias vector; X feature represents a non-temporal feature;
[0046] S3013: Time2Vec encodes the temporal feature to obtain a second feature T embed ;
[0047] The formula is:
[0048]
[0049] wherein: is a linear term parameter; is a periodic term parameter; X time represents a temporal feature;
[0050] S3014: The first feature F embed and the second feature T embed are fused to obtain an initial fusion feature X embd ;
[0051] The formula is:
[0052] X embd = F embed + T embed .
[0053] Preferably, in step S302, the following steps are specifically included:
[0054] S3021: One-dimensional average pooling is used to smooth the initial fusion feature X embd , and the sequence length is preserved by symmetric padding to extract a trend item feature X trend representing long-term gradual changes.
[0055] The formula is:
[0056] X trend = AvgPool1D(Padding(X embd ));
[0057] wherein: AvgPool1D represents one-dimensional average pooling; Padding represents symmetric padding operation;
[0058] S3022: The trend item feature X embd is subtracted from the initial fusion feature X trend to obtain a seasonal item feature X season reflecting periodic fluctuations and short-term details.
[0059] The formula is represented as:
[0060] X season = X embd - X trend .
[0061] Preferably, in step S303, the following steps are specifically included:
[0062] S3031: modeling the trend item feature X trend through a Transformer encoder to obtain the feature Z trend ;
[0063] The formula is represented as:
[0064]
[0065] S3032: modeling the seasonal item feature X season through a Transformer encoder to obtain the feature Z season ;
[0066] The formula is represented as:
[0067]
[0068] S3033: fusing the feature Z trend and the feature Z season to obtain the trend-seasonal fusion feature Z enc ;
[0069] The formula is represented as:
[0070] Z enc = Z trend + Z season .
[0071] Preferably, in step S304, the following steps are specifically included:
[0072] S3041: extracting the local feature Z enc of the trend-seasonal fusion feature Z k using one-dimensional convolution with a kernel size set S = {1, 3, 5} in parallel;
[0073] The formula is represented as:
[0074] Z k = Conv1D k (Z enc ), k ∈ S;
[0075] wherein Conv1D k represents one-dimensional convolution;
[0076] S3042: take the local feature Z k with trend-seasonal fusion feature Z enc fusion, get multi-scale fusion feature Z multi ;
[0077] The formula is:
[0078] Z multi = Avg(Z k + Z enc );
[0079] S3043: input the multi-scale fusion feature Z multi into the multi-head attention mechanism, calculate the dependency relationship of each time step to other time steps: self-attention performs weighted summation on each time step of the input sequence, capturing the long-distance dependency relationship of the sequence; then add the attention output to the original feature to get the attention fusion feature Z attn ;
[0080] The formula is:
[0081] Z attn = LayerNorm(MultiHeadAttention(Z multi ) + Z multi ).
[0082] Preferably, in step S305, the following steps are specifically included:
[0083] S3051: respectively perform average pooling and maximum pooling on the attention fusion feature Z attn , and splice to get the global feature h;
[0084] The formula is:
[0085] h = Concat(AvgPool(Z attn ), MaxPool(Z attn ));
[0086] In the formula: h represents the global feature; Concat represents the splicing operation;
[0087] S3052: input the global feature h into the fully connected layer for prediction, to get the predicted capacity sequence
[0088] The formula is:
[0089]
[0090] In the formula: represents the predicted capacity sequence; W o , b odenote the learnable bias parameters in the fully connected layer.
[0091] Preferably, in step S3, the loss function when training the capacity prediction model is as follows:
[0092]
[0093] wherein: H denotes the number of data points in the predicted capacity sequence and the labeled capacity sequence; y h and denote the hth labeled capacity and predicted capacity in the labeled capacity sequence Y and the predicted capacity sequence , h∈{1,…,H}.
[0094] A computer device, comprising: one or more processors;
[0095] The processor is configured to store one or more programs.
[0096] When the one or more programs are executed by the one or more processors, the automobile key component performance degradation intelligent diagnosis large model and prediction method according to any one of claims 1-9 are implemented.
[0097] Compared with the prior art, the automobile key component performance degradation intelligent diagnosis large model and prediction method has the following advantages
[0098] Beneficial effects:
[0099] The present application can fully mine various information contained in the battery charging data by extracting multi-dimensional statistical features, which reflect the charging state and performance characteristics of the battery from different angles. Combined with the Pearson correlation coefficient and the grey correlation degree, the features highly related to the battery capacity labeling can be accurately selected to construct a low-redundancy and high-interpretability feature set, improve the representation ability of the feature set for battery performance degradation, and enable the subsequent prediction model to learn based on more representative features, thereby improving the accuracy of battery capacity prediction and accurately evaluating the state of health (SOH) of the battery. At the same time, the optimal feature set extracted can better adapt to data of different working conditions and different types of batteries, so that the capacity prediction model can still maintain good prediction performance when facing unknown battery charging data, thereby enhancing the generalization ability of the model.
[0100] The application divides the optimal feature set into non-time features and time features, the non-time features generally reflect the static properties of the battery, such as the model of the battery, the initial capacity and the like, and the time features embody the dynamic changes of the battery in the charging process, such as the change of the charging current with time and the like, the division and fusion of the two types of features can comprehensively utilize the static and dynamic information of the battery, more comprehensively describe the performance state of the battery, provide more abundant input information for the subsequent prediction model, and further improve the accuracy of the battery capacity prediction.
[0101] The application separates the initial fusion features into trend item features and seasonal item features, the trend item features reflect the long-term trend that the battery capacity gradually decreases with time, and the seasonal item features embody the fluctuation law of the battery performance in certain periods, modeling the trend item features and the seasonal item features respectively can more accurately capture the two change laws, avoid the mutual interference of the trend and the seasonal factors, and make the model more in-depth and accurate in understanding the degradation of the battery performance. Meanwhile, the trend-seasonal fusion features are generated by modeling and fusing the trend item and the seasonal item, which can comprehensively consider the long-term change and short-term fluctuation of the battery performance, so that the prediction model can better adapt to the performance changes of different time scales and reduce the prediction error caused by ignoring the seasonal factors or the trend changes.
[0102] The local features of the trend-seasonal fusion features are extracted, which can capture the local change mode of the battery performance at different time points, and the multi-head attention mechanism can calculate the dependency relationship of each time step to other time steps, automatically learn the importance weight between different positions in the time sequence, this mechanism makes the model better handle the long-distance dependency problem in the time sequence data, pay attention to the historical information important to the current prediction, and enhances the modeling capability of the model for the degradation time sequence of the battery performance, thereby improving the accuracy of the battery capacity prediction. Meanwhile, the attention fusion features are subjected to pooling processing, which can reduce the dimension of the features and reduce the calculation amount; the pooled features are input into the full connection layer for prediction, which can quickly obtain the prediction capacity sequence of the future continuous time steps under the premise of ensuring the prediction accuracy, thereby improving the practicability and reliability of the whole prediction method. BRIEF DESCRIPTION OF DRAWINGS
[0103] In order to make the purpose, technical scheme and advantages of the application clearer, the application will be further described in detail below with reference to the drawings, in which:
[0104] Figure 1 The logic block diagram of the intelligent diagnosis large model and prediction method for the performance degradation of the key components of the automobile.
[0105] Figure 2 The battery pack of the electric vehicle is calculated for capacity.
[0106] Figure 3Results of the correlation analysis of battery capacity and features.
[0107] Figure 4 Battery capacity prediction model.
[0108] Figure 5 Input embedding module.
[0109] Figure 6 Multi-scale convolutional attention module.
[0110] Figure 7 Battery capacity prediction results for different feature sets.
[0111] Figure 8 Capacity prediction results for different models.
[0112] Figure 9 Capacity prediction results for different backtracking windows. DETAILED DESCRIPTION
[0113] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments but not all embodiments of the present application. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative work based on the embodiments in the present application belong to the scope of protection of the present application.
[0114] It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the figures, or the orientation or positional relationship commonly used when the product is in use. They are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance. In addition, the terms "horizontal," "vertical," etc., do not mean that the component is required to be absolutely horizontal or suspended, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted. In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0115] The following detailed explanation illustrates the specific implementation methods:
[0116] Example 1:
[0117] This embodiment discloses a large-scale intelligent diagnostic model and prediction method for performance degradation of key automotive components.
[0118] like Figure 1 As shown, a large-scale intelligent diagnostic model and prediction method for performance degradation of key automotive components includes:
[0119] S1: Obtain the standard capacity corresponding to each group of battery charging data in the historical charging dataset;
[0120] S2: Extract multi-dimensional statistical features of each group of battery charging data, and select the optimal feature set based on the Pearson Correlation Coefficient (PCC) and Grey Relational Grade (GRG) combined with the labeled capacity;
[0121] In this embodiment, the multi-dimensional statistical features can include total charging capacity, CC stage time, CV stage voltage, current variance, voltage kurtosis, CC-CV conversion capacity, voltage platform length, incremental capacity peak position, differential voltage valley depth, capacity attenuation rate, internal resistance growth rate, capacity attenuation trend slope, sliding window mean, etc. The optimal features can include total charging capacity, CC stage time, CV stage voltage, current variance, voltage kurtosis, etc.
[0122] S3: input the optimal feature set into the trained capacity prediction model, and output the predicted capacity sequence of the future continuous time steps;
[0123] The capacity prediction model can effectively capture the long-term and short-term dependence relationship of capacity attenuation through time series decomposition, multi-scale attention mechanism and double encoder architecture, and can accurately predict the future capacity degradation trajectory based on different historical backtracking window lengths. The processing steps of the capacity prediction model include:
[0124] S301: divide the optimal feature set into non-time features and time features, and fuse to generate initial fusion features;
[0125] S302: separate the initial fusion features into trend item features and seasonal item features;
[0126] S303: model the trend item features and the seasonal item features respectively, and fuse to generate trend-seasonal fusion features;
[0127] S304: extract local features of the trend-seasonal fusion features, and calculate the dependence relationship of each time step to other time steps through a multi-head attention mechanism to obtain attention fusion features;
[0128] S305: after the attention fusion features are processed by pooling, input into a fully connected layer for prediction to obtain the predicted capacity sequence of the future continuous time steps.
[0129] The present application can comprehensively mine various information contained in the battery charging data by extracting multi-dimensional statistical features, which reflect the charging state and performance characteristics of the battery from different angles. Combined with the Pearson correlation coefficient and the grey correlation degree, the features highly related to the battery capacity label can be accurately selected to construct a feature set with low redundancy and high interpretability, improve the representation ability of the feature set for battery performance degradation, and enable the subsequent prediction model to learn based on more representative features, thereby improving the accuracy of battery capacity prediction and accurately evaluating the state of health (SOH) of the battery. At the same time, the optimal feature set extracted can better adapt to the data of different working conditions and different types of batteries, so that the capacity prediction model can still maintain good prediction performance when facing unknown battery charging data, thereby enhancing the generalization ability of the model.
[0130] The application divides the optimal feature set into non-time features and time features, the non-time features generally reflect the static properties of the battery, such as the model of the battery, the initial capacity and the like, and the time features embody the dynamic changes of the battery in the charging process, such as the change of the charging current with time and the like, the division and fusion of the two types of features can comprehensively utilize the static and dynamic information of the battery, more comprehensively describe the performance state of the battery, provide more abundant input information for the subsequent prediction model, and further improve the accuracy of the battery capacity prediction.
[0131] The application separates the initial fusion features into trend item features and seasonal item features, the trend item features reflect the long-term trend that the battery capacity gradually decreases with time, and the seasonal item features embody the fluctuation law of the battery performance in certain periods, modeling the trend item features and the seasonal item features respectively can more accurately capture the two change laws, avoid the mutual interference of the trend and the seasonal factors, make the understanding of the battery performance degradation of the model more in-depth and accurate. Meanwhile, the trend-seasonal fusion features are generated by modeling and fusing the trend item and the seasonal item, which can comprehensively consider the long-term change and short-term fluctuation of the battery performance, so that the prediction model can better adapt to the performance changes of different time scales and reduce the prediction error caused by ignoring the seasonal factors or the trend changes.
[0132] The local features of the trend-seasonal fusion features are extracted, which can capture the local change mode of the battery performance at different time points, and the multi-head attention mechanism can calculate the dependency relationship of each time step to other time steps, automatically learn the importance weight between different positions in the time sequence, this mechanism makes the model better handle the long-distance dependency problem in the time sequence data, pay attention to the historical information important to the current prediction, enhances the modeling capability of the model for the battery performance degradation time sequence, and thus improves the accuracy of the battery capacity prediction. Meanwhile, the attention fusion features are subjected to pooling processing, which can reduce the dimension of the features and reduce the calculation amount; the pooled features are input into the full connection layer for prediction, which can quickly obtain the prediction capacity sequence of the future continuous time steps under the premise of ensuring the prediction accuracy, thereby improving the practicability and reliability of the whole prediction method.
[0133] In order to better introduce the technical scheme of the application, the embodiment is described through the following several parts.
[0134] I. Label capacity calculation
[0135] In the embodiment, the reliable labeled capacity is obtained by improving the ampere integral formula and combining the monthly statistical average method. The labeled capacity is calculated through the following steps:
[0136] S101: Calculate the labeled capacity based on the state of charge of the battery charging data by improving the type current integral method;
[0137] The formula is expressed as:
[0138]
[0139] In the formula, C represents the labeled capacity; Δt represents a fixed sampling interval; I(t) represents a charging current (a negative value represents a charging process); t1 and t2 represent charging start and end times; SOC t1 , SOC t2 represent the state of charge at the start and end of charging; it should be particularly pointed out that the accuracy of the formula directly depends on the estimation accuracy of the state of charge (SOC), and a sufficient SOC variation interval is required to reduce the integral error.
[0140] S102: The labeled capacity of each battery charging data is subjected to mean or median processing by using a sliding window statistical method, and abnormality caused by SOC error or data noise is excluded;
[0141] The formula is expressed as:
[0142]
[0143] In the formula, Ck represents the labeled capacity of the kth battery charging data; C i represents the kth labeled capacity in the statistical window; N is the total amount of data in the statistical window; and k represents a statistical period number.
[0144] As shown in Figure 2 , the capacity calculated by the current integration method has a large number of abnormal values, and a large number of points in the same month fluctuate. Statistics show that the system can collect more than 90 capacity data points per month. For the vehicle operation and maintenance scene, it is not necessary to pay attention to the battery capacity of each charging process, which is also very difficult to obtain accurately. Therefore, it is reasonable to statistically process the capacity change in units of months, and the capacity curve after monthly statistical processing presents a clear degradation trend, and the difference between the mean value and the median value is less than 0.8%, verifying the symmetric characteristic of the capacity distribution.
[0145] The sliding window statistical method is adopted to process the calculated capacity value by mean or median, which can effectively filter out abnormal points caused by state of charge (SOC) estimation error or data noise, and improve the data reliability.
[0146] II. Feature selection
[0147] In this embodiment, feature selection is realized by the following steps:
[0148] S201: The feature sequence and the labeled capacity sequence of the battery charging data in one period (one month) are subjected to multi-dimensional statistical analysis, and the basic statistical quantities including the arithmetic mean, the cumulative total amount and the standard deviation are calculated;
[0149] S202: Calculate the Pearson correlation coefficient p of the feature and the labeling capacity based on the ratio of the covariance and the standard deviation of the feature sequence and the labeling capacity sequence xi ;
[0150] The formula is:
[0151]
[0152] In the formula: p xi represents the Pearson correlation coefficient of the ith feature x i and the labeling capacity; c represents the labeling capacity; and are the average values of the feature sequence and the labeling capacity sequence, respectively;
[0153] S203: Calculate the gray correlation coefficient r of the feature and the labeling capacity by constructing the correlation function of the feature sequence and the labeling capacity sequence i ;
[0154] The formula is:
[0155]
[0156] In the formula: r i represents the gray correlation coefficient of the ith feature x i and the labeling capacity; n represents the number of features, and i (k) is the gray correlation coefficient of the ith feature x i at time k; p is the discrimination coefficient, usually set to 0.5; x i (k) represents the ith feature at time k; c i (k) is the ith labeling capacity at time k;
[0157] S204: Select the feature with a Pearson correlation coefficient greater than the first coefficient (0.5) with the labeling capacity, or a gray correlation coefficient greater than the second coefficient (0.7) with the labeling capacity as the optimal feature. The Pearson correlation coefficient and the gray correlation coefficient of the feature can be further sorted from large to small, and finally a feature set is constructed.
[0158] In combination with Figure 3 It can be seen that the PCC and GRG average values of 20 vehicles are evaluated in order to obtain accurate correlation analysis results. In addition to the charging data, the time represented by the month is also used as a potential feature, but its correlation with the battery capacity is relatively weak. It can be seen that the features highly correlated with the battery capacity determined by the two methods are almost identical.
[0159] The application comprehensively uses Pearson correlation coefficient (PCC) and grey correlation degree (GRG) to screen out key features (such as temperature difference, voltage fluctuation, etc.) strongly related to capacity degradation from multi-dimensional charging data, and constructs a feature set with low redundancy and high interpretability.
[0160] III. Capacity prediction model
[0161] In combination Figure 4 As shown in the figure, the capacity prediction model includes an input embedding module, a sequence decomposition module, a double encoder module, a multi-scale convolution attention module and an output prediction module.
[0162] 1. Input embedding module
[0163] In combination Figure 5 As shown in the figure, the input embedding module works cooperatively with feature projection and time coding to realize effective fusion of multi-source time sequence features. The module first divides the input into non-time features and timestamp features, and performs linear projection and Time2Vec coding respectively.
[0164] The steps of the working of the input embedding module are as follows:
[0165] S3011: Divide the optimal feature set into non-time features X feature and time features X time ;
[0166] S3012: Linearly project the non-time features to obtain the first feature F embed ;
[0167] The formula is:
[0168] F embed = W f X feature + b f ;
[0169] In the formula: W represents a learnable projection matrix that maps D-1 dimensional features to H dimensional space; b represents a bias vector that enhances the feature expression capability; X feature represents non-time features;
[0170] S3013: Time2Vec encode the time features to obtain the second feature T embed ;
[0171] The formula is:
[0172]
[0173] In the formula: is a linear term parameter that captures the monotonic trend of time; is a period term parameter, encoding the time periodicity; X time represents the time characteristics;
[0174] S3014: Fuse the first feature F embed and the second feature T embed to obtain an initial fusion feature X embd ;
[0175] The formula is:
[0176] X embd = F embed + T embed .
[0177] 2. Sequence decomposition module
[0178] As shown in Figure 4 , the sequence decomposition module explicitly separates the time series feature into a trend term and a seasonal term through a moving average decomposition method based on the idea of Autoformer to enhance the model's independent modeling ability for long-term trends and short-term fluctuations. Unlike the original method, this design optimizes it into a general feature processing component that directly acts on the hidden features of the encoder stage, thus achieving more fine-grained time series pattern decoupling at the feature level.
[0179] The steps of the working of the sequence decomposition module are as follows:
[0180] S3021: One-dimensional average pooling (AvgPool1D) is used to smooth the initial fusion feature X embd , and symmetric padding (Padding) is used to retain the sequence length to extract the trend term feature X trend representing long-term gradual changes;
[0181] The formula is:
[0182] X trend = AvgPool1D(Padding(X embd ));
[0183] In the formula: AvgPool1D represents one-dimensional average pooling; Padding represents symmetric padding operation;
[0184] S3022: Subtract the trend term feature X embd from the initial fusion feature X trend to retain the seasonal term feature X season reflecting periodic fluctuations and short-term details;
[0185] The formula is:
[0186] X season = X embd -Xtrend .
[0187] where, B denotes batch size, L denotes sequence length, and L denotes hidden dimension.
[0188] 3. Double encoder module
[0189] In combination Figure 4 As shown in FIG. 3B, the double encoding module adopts a trend-seasonal double-branch encoder architecture, which independently models the trend and seasonal items output by the time series decomposition module to avoid mutual interference of the two types of signals in the feature space.
[0190] The steps of the operation of the double encoding module are as follows:
[0191] S3031: Model the trend item feature X trend through a Transformer encoder to obtain the feature Z trend .
[0192] The formula is:
[0193]
[0194] S3032: Model the seasonal item feature X season through a Transformer encoder to obtain the feature Z season .
[0195] The formula is:
[0196]
[0197] S3033: Fuse the feature Z trend and the feature Z season to obtain the trend-seasonal fusion feature Z enc .
[0198] The formula is:
[0199] Z enc = Z trend + Z season .
[0200] where,
[0201] 4. Multi-scale convolution attention module
[0202] In combination Figure 6 As shown in FIG. 4B, the multi-scale convolution module adopts a collaborative architecture of multi-scale convolution and multi-head attention, which realizes joint modeling of local time series patterns and global dependency relationships.
[0203] The steps of the operation of the multi-scale convolution attention module are as follows:
[0204] S3041: One-dimensional convolution (Conv1D) parallel extraction of trend-seasonal fusion features Z using the core size set S = {1, 3, 5} enc Local features Z k ;
[0205] The formula is expressed as:
[0206] Z k = Conv1D k (Z enc ), k ∈ S;
[0207] In the formula, Conv1D k represents one-dimensional convolution;
[0208] S3042: Local features Z k and trend-seasonal fusion features Z enc are fused to obtain multi-scale fusion features Z multi ;
[0209] The formula is expressed as:
[0210] Z multi = Avg(Z k +Z enc );
[0211] Where,
[0212] S3043: Multi-scale fusion features Z multi are input into the multi-head attention mechanism to calculate the dependency relationship of each time step to other time steps: self-attention performs weighted summation on each element (time step) of the input sequence to capture long-distance dependencies in the sequence; then the attention output is added to the original feature for normalization (helpful to avoid gradient disappearance problem during training, and accelerate convergence), to obtain attention fusion features Z attn ;
[0213] The formula is expressed as:
[0214] Z attn = LayerNorm(MultiHeadAttention(Z multi )+Z multi ).
[0215] 5, output prediction module
[0216] In combination Figure 4As shown, a mixed pooling strategy (average pooling and max pooling) is adopted in the feature aggregation part to further compress the global feature and capture important patterns in the sequence. The pooled features are then concatenated and output by a fully connected layer to obtain the prediction results of the future time steps.
[0217] The steps of the working of the output prediction module are as follows:
[0218] S3051: respectively perform average pooling and max pooling on the attention fusion features Z attn , and concatenate to obtain the global feature h;
[0219] The formula is as follows:
[0220] h = Concat (AvgPool (Z attn ), MaxPool (Z attn ));
[0221] In the formula: h represents the global feature, Concat represents the concatenation operation;
[0222] S3052: input the global feature h into the fully connected layer for prediction to obtain the predicted capacity sequence
[0223] The formula is as follows:
[0224]
[0225] In the formula: represents the predicted capacity sequence, W o , b o represents the learnable bias parameter in the fully connected layer.
[0226] Four, loss function
[0227] In this embodiment, the loss function when training the capacity prediction model is as follows:
[0228]
[0229] In the formula: H represents the number of data points of the predicted capacity sequence and the labeled capacity sequence; y h and represent the hth labeled capacity and predicted capacity in the labeled capacity sequence Y and the predicted capacity sequence , h∈{1,…,H}.
[0230] Five, experimental description
[0231] In order to better illustrate the advantages of the technical scheme of the present application, the following experiments are disclosed in this embodiment.
[0232] The experiment uses the published battery dataset to test the intelligent diagnosis and prediction method of the automobile key component performance degradation based on the multi-scale attention mechanism model proposed in the application.
[0233] 1. Dataset
[0234] The dataset contains 20 groups of charging operation data of commercial electric vehicles of the same model, and the sample time span is 24 months (the actual collection period is about 29 months). The research object uses a unified specification of battery system, and the experimental samples are sequentially numbered as vehicle_1 to vehicle_20. Data collection relies on on-board charging equipment, which obtains battery state parameters in the charging process through the controller area network (CAN) bus in real time. The system continuously records the charging characteristic parameters at a sampling interval of about 8 seconds. The battery partial charging dataset is shown in Table 1.
[0235] Table 1. Partial charging dataset of automobile battery
[0236]
[0237] 2. Model parameter setting
[0238] By data cleaning, battery capacity calculation, and feature correlation calculation on the original charging data, the best feature set can be constructed as the input of the capacity prediction model as shown in Figure 4 . The parameters of the capacity prediction model are shown in Table 2.
[0239] Table 2. Parameter setting of capacity prediction model
[0240]
[0241] 3. Experimental results
[0242] 1) Battery capacity prediction verification
[0243] In order to verify the effectiveness of the feature selection process, the experiment compares the prediction results based on different feature sets. Four different feature sets are used as the input of the model. F3 is the optimal feature set determined by the feature selection process proposed in the application, F1 and F2 are composed of features with |p|>0.5 or r>0.7, and F4 is a set of features with some related features removed. F3 has 2 more related features compared with F1 and F2. The results of the first test based on different feature sets are shown in Figure 7 , where the prediction results of two vehicles are represented by vehicle_3 and vehicle_6 respectively.
[0244] 2) Ablation experiment
[0245] To verify the effectiveness of the core components of the model, the ablation experiment was conducted, and the results are shown in Table 3. Removing the time encoding module caused the MAE / MSE to increase by 50.7% and 98.5%, respectively, highlighting the key role of Time2Vec in jointly modeling trends and cycles. Removing the sequence decomposition module performed the worst (MAE = 0.6396), with an error increase of 74.5% over the baseline, demonstrating that explicitly separating trend and seasonal components is the cornerstone of the model. The removal of the multi-scale attention module increased the MAE to 0.4108, indicating that the synergy of local convolution and global attention can effectively capture multi-granularity temporal patterns. The simplification of the dual-encoder module and the hybrid pooling introduced 22.3% and 18.1% MAE loss, respectively, verifying the necessity of independent component encoding and feature aggregation strategies. Experiments show that each module significantly improves prediction accuracy through hierarchical collaboration (feature decomposition → independent encoding → multi-scale fusion).
[0246] Table 3 Ablation experiment results
[0247]
[0248]
[0249] 3) Comparison of different prediction models
[0250] The experiment verifies the performance of the model from the aspects of prediction accuracy and long-term stability through quantitative indicators and visual analysis. Table 4 shows the prediction accuracy comparison results of different models under different backtracking window lengths L, where TideFormer is the proposed capacity prediction model. The early data of the battery for the first 6 months, 12 months, 18 months, and 24 months are considered known, and ten-fold cross-validation is still used to evaluate the prediction accuracy.
[0251] Combining Table 4 and Figure 8As shown, the MAE and MSE of TideFormer are significantly better than those of the comparative models under different backtracking window lengths (L = 6, 12, 18, 24), especially in the long window (L = 24) task (MAE = 0.451, MSE = 0.411), which further reduces the error of the suboptimal model (such as Transformer, MAE = 0.460). With the increase of L, the prediction accuracy of TideFormer steadily increases, and the MSE decreases from 0.660 to 0.4111, indicating that its sequence decomposition and multi-scale attention mechanism can effectively model long-term dependencies. In contrast, the traditional time series models (such as LSTM, TCN) and the recent research models (such as TimesNet, PatchTST) significantly increase the error in this task (such as TimesNet, MSE = 1.104 when L = 24, which is 2.68 times that of TideFormer), highlighting the advantages of the architecture of TideFormer in trend-season decoupling and hierarchical feature fusion.
[0252] Table 4 Prediction results under different backtracking window lengths
[0253]
[0254] To further verify the robustness of TideFormer to the length of the backtracking window, the capacity prediction curves under different backtracking window lengths (L = 6, 12, 18, 24) are plotted separately as shown in Figure 9 As shown, it can be found that all four cases can accurately predict the future capacity trajectory. When the amount of prior data decreases, there is no significant decrease in prediction accuracy, which confirms that the proposed method has strong robustness.
[0255] Embodiment Two
[0256] The embodiment discloses a computer device.
[0257] The computer device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor. The processor implements the steps in each of the above-mentioned intelligent diagnosis and prediction methods for performance degradation of various automobile key components when executing the computer program. Alternatively, the processor implements the functions of each module / unit in each of the above-mentioned system embodiments when executing the computer program.
[0258] For example, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the device operation and maintenance intelligent question and answer system.
[0259] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server, and the like. The computer device can include, but is not limited to, a processor, a memory.
[0260] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the technical solutions. Those of ordinary skill in the art should understand that the technical solutions of the present application are modified or replaced equivalently without departing from the purpose and scope of the technical solutions, which should be covered in the scope of claims of the present application.
Claims
1. An intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile, characterized in that, Comprise: S1: obtaining the standard capacity corresponding to each group of battery charging data in the historical charging data set; S2: extracting the multi-dimensional statistical features of each group of battery charging data, and selecting the optimal feature set based on the combination of the Pearson correlation coefficient and the gray correlation degree of the labeled capacity; S3: inputting the optimal feature set into the trained capacity prediction model to output the predicted capacity sequence of the future continuous time steps; The processing steps of the capacity prediction model include: S301: dividing the optimal feature set into non-time features and time features, and fusing to generate initial fusion features; S302: separating the initial fusion features into trend item features and seasonal item features; S303: modeling the trend item features and seasonal item features respectively, and fusing to generate trend-seasonal fusion features; S304: extracting local features of the trend-seasonal fusion features, and calculating the dependency relationship of each time step to other time steps through a multi-head attention mechanism to obtain attention fusion features; S305: after the attention fusion features are processed by pooling, input them into a fully connected layer for prediction to obtain the predicted capacity sequence of the future continuous time steps.
2. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 1, characterized in that: In step S1, the following steps are included: S101: calculating the labeled capacity based on the state of charge of the battery charging data by an improved current integration method; The formula is: wherein: C represents the labeled capacity; Δt represents the fixed sampling interval; I(t) represents the charging current; t1, t2 represent the charging start and end times; SOC t1 , SOC t2 represents the state of charge at the charging start and end. S102: using a sliding window statistical method to process the mean or median of the labeled capacity of each battery charging data to exclude abnormalities caused by SOC errors or data noise; The formula is: In the formula, denotes the labeled capacity of the kth battery charging data; C i denotes the kth labeled capacity within the statistical window; N is the total amount of data within the statistical window; k represents the statistical period number.
3. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 1, characterized in that: In step S2, the following steps are included: S201: performing multi-dimensional statistical analysis on the feature sequence and the labeled capacity sequence of the battery charging data within a period, and calculating the basic statistics including arithmetic mean, cumulative total and standard deviation; S202: Calculate the Pearson correlation coefficient p of the feature and the labeling capacity based on the ratio of the covariance and the standard deviation of the feature sequence and the labeling capacity sequence xi ; The formula is: where: p xi represents the ith feature x i with the labeled capacity Pearson correlation coefficient; c represents the labeled capacity; and are the mean values of the feature sequence and the labeled capacity sequence, respectively; S203: Calculate the grey correlation degree coefficient r of the feature and the labeling capacity by constructing the correlation function of the feature sequence and the labeling capacity sequence i ; The formula is: wherein: r i represents the i-th feature x i is the grey correlation coefficient of the labeled capacity; n represents the number of features, and ξ i (k) is the i-th feature x i is the grey correlation coefficient at the k-th moment; p is a discrimination coefficient; x i (k) represents the i-th feature at the k-th moment; c i (k) is the i-th labeled capacity at the k-th moment; S204: selecting the features with a Pearson correlation coefficient greater than the first coefficient or a gray correlation degree coefficient greater than the second coefficient as the optimal features.
4. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 1, characterized in that: In step S301, the following steps are included: S3011: divide the optimal feature set into non-time features X feature with time features X time ; S3012: Linearly project the non-time feature to obtain a first feature F embed ; The formula is: F embed = W f X feature + b f ; In the formula: represents a learnable projection matrix; bias vector; X feature represents a non-temporal feature; S3013: Time2Vec encodes the time feature to obtain a second feature T embed ; The formula is: wherein: is a linear term parameter; is a periodic term parameter; X time denotes a time characteristic; S3014: fuse the first feature F embed and the second feature T embed to obtain an initial fused feature X embd ; The formula is: X embd = F embed + T embed .
5. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 1, characterized in that: In step S302, the following steps are included: S3021: One-dimensional average pooling is performed on the initial fusion feature X embd Smooth processing is performed, the sequence length is retained by symmetric padding, and a trend item feature X representing long-term gradual change is extracted trend ; The formula is: X trend = AvgPool1D(Padding(X embd )) In the formula, AvgPool1D represents one-dimensional average pooling, and Padding represents a symmetric padding operation; S3022: subtract the trend term feature X embd from the initial fused feature X trend , leaving the seasonal term feature X season that reflects periodic fluctuations and short-term details; The formula is: X season = X embd - X trend .
6. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 1, characterized in that: In step S303, the following steps are included: S3031: encode the trend item feature X by a Transformer encoder trend Modeling is performed to obtain the feature Z trend ; The formula is: S3032: encode the seasonal term feature X by a Transformer encoder season Modeling is performed to obtain the feature Z season ; The formula is: S3033: merge the feature Z trend and the feature Z season to obtain a trend-seasonal merged feature Z enc ; The formula is: Z enc = Z trend + Z season .
7. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 1, characterized in that: In step S304, the following steps are included: S3041: extract the trend-seasonal fused feature Z in parallel using one-dimensional convolution with the kernel size set S = {1, 3, 5} enc local feature Z k ; The formula is: Z k = Conv1D k (Z enc ), k e S; where: Conv1D k represents a one-dimensional convolution; S3042: take the local feature Z k with trend-seasonal fusion feature Z enc fusion, get multi-scale fusion feature Z multi ; The formula is: Z multi = Avg(Z k + Z enc ); S3043: fusing the multi-scale features Z multi The input multi-head attention mechanism calculates the dependency relationship of each time step to other time steps: the self-attention performs weighted summation on each time step of the input sequence to capture the long-distance dependency relationship of the sequence; and the attention output is added to the original feature to obtain the attention fusion feature Z attn ; The formula is: Z attn = LayerNorm(MultiHeadAttention(Z multi ) + Z multi ).
8. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 1, characterized in that: In step S305, the following steps are included: S3051: average pooling and maximum pooling are respectively performed on the attention fusion features Z attn and the global features h are obtained by splicing. The formula is: h = Concat(AvgPool(Z attn ), MaxPool(Z attn )); In the formula, h represents a global feature, and Concat represents a concatenation operation; S3052: input the global feature h into the fully connected layer for prediction to obtain the predicted capacity sequence of the future continuous time steps The formula is represented as: In the formula: represents the predicted capacity sequence; W o , b o represents the learnable bias parameter in the fully connected layer.
9. The intelligent diagnosis large model and prediction method for performance degradation of key components of an automobile according to claim 8, characterized in that: In step S3, the loss function when training the capacity prediction model is as follows: where H denotes the number of data points of the predicted and annotated capacity sequences; y h and denotes the h-th annotated and predicted capacity in the annotated capacity sequence Y and the predicted capacity sequence Y, h ∈ {1,..., H}.
10. A computer device, comprising: Comprise: One or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, the automobile key component performance degradation intelligent diagnosis large model and prediction method according to any one of claims 1-9 is implemented.