QL-RRVM-based super capacitor residual service life prediction method and auxiliary module
By integrating the feature extraction module of reinforcement learning in the correlation vector machine model, the QL-RRVM model is built, and the accuracy of the prediction of the residual service life of supercapacitors under the influence of data variability and outliers is solved, achieving more efficient and robust prediction effects.
Patent Information
- Application Number
- CN202411895091.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-21
- Publication Date
- 2025-05-06
AI Technical Summary
The existing data-driven supercapacitor residual service life prediction methods decrease prediction accuracy when facing a large number of data variability and outliers.
A framework based on robust correlation vector machine (Robust RVM) combined with reinforcement learning is adopted to build a QL-RRVM model. By integrating reinforcement learning-based feature extraction modules in front of improved correlation vector machines, the robustness and prediction accuracy of the model are improved.
In the presence of outliers, the QL-RRVM model can maintain good predictability, improve the real-time and efficiency of the system, and enhance the interpretability of the prediction results.
Smart Images

Figure CN119939531A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of measuring electrical variables and magnetic variables, and in particular to a method for predicting the remaining useful life of a supercapacitor based on QL-RRVM and an auxiliary module. Background Art
[0002] As a new type of energy storage device, supercapacitors are widely used in electric vehicles, energy storage systems and other fields due to their efficient charging and discharging characteristics and long life. However, supercapacitors will inevitably age during use, affecting their energy storage and output capabilities, and thus affecting the stability and safety of the system. Therefore, accurately predicting the remaining useful life (RUL) of supercapacitors is crucial for the reliable operation of the system.
[0003] At present, the remaining service life prediction methods mainly include model-based methods and data-driven methods. Model-based methods predict the remaining service life of supercapacitors by establishing a physical model, mechanism model or random process model to describe the supercapacitor and combining historical data. This type of method requires a large amount of capacitance parameter data and a deep understanding of the physical aging mechanism of supercapacitors. Although the prediction accuracy is high, the model establishment process is complex and costly, and it is difficult to adapt to the complex aging patterns in different working environments. In contrast, the data-driven method is simpler and does not require a clear understanding of the physical mechanism of supercapacitors, and can better adapt to complex and dynamic usage environments.
[0004] Existing data-driven methods include neural networks and relevance vector machines (RVMs), which aim to analyze historical and monitoring data, reveal the potential relationships between the data, and thus predict the degradation trend of the remaining useful life and conduct subsequent analysis. RVMs are sparse models based on a Bayesian framework, and can also perform well in the case of insufficient data. In the Bayesian framework, the weight of each input is controlled by a set of parameters. RVMs are widely used in regression problems due to their wide sparsity and good predictive performance.
[0005] However, although the RVM excels in high accuracy and sparsity, its predictive ability falters under the influence of large amounts of data variability and outliers, resulting in decreased predictive accuracy. Summary of the invention
[0006] In view of the problems existing in the prior art, the present invention provides a supercapacitor remaining service life prediction method and auxiliary module based on QL-RRVM in a framework combining a robust relevance vector machine (Robust RVM) with reinforcement learning, so that good predictability can be maintained even in the presence of outliers.
[0007] The technical solution adopted by the present invention is a method for predicting the remaining useful life of a supercapacitor based on QL-RRVM, constructing an improved relevance vector machine RRVM, and integrating a feature extraction module based on reinforcement learning before the improved relevance vector machine RRVM to obtain a QL-RRVM model for predicting degradation trends according to time series;
[0008] The raw data of the supercapacitor is collected, and the state of the supercapacitor at the current moment is obtained based on any action. After inputting into the QL-RRVM model, the decision is guided based on the predicted target value and expected return, and the remaining service life of the supercapacitor is dynamically estimated using the predicted degradation trend.
[0009] Preferably, based on the Bayesian framework of the relevance vector machine, an improved relevance vector machine RRVM is constructed with parameters following t distribution.
[0010] Preferably, constructing the improved relevance vector machine RRVM comprises the following steps:
[0011] S1.1 define the prediction function y(x);
[0012] S1.2 Given a time series dataset The predicted target value satisfies t i =y(x i )+ε, where x i =(x i ,x i+1 ,…,x i+L-1 ), N represents the number of target groups, L represents the step size; let ε obey the mean of zero and the variance of σ 2 , introduce accuracy β and satisfy β -1 =σ 2 , we get the weights and β -1 The predicted value of t i The probability density function of
[0013] S1.3 defines the weight matrix W as a Gaussian distribution with zero mean and different variances, and sets the precision matrix α corresponding to the weight matrix;
[0014] S1.4 Estimate parameters W, β and α by maximizing the probability density function, and keep the posterior probability distribution of parameter W as Gaussian distribution; let parameters α and β obey gamma distribution, and predict the target value to obey t distribution, and obtain the corresponding probability density function, and the posterior probability density function of weight W and random error ε;
[0015] S1.5 obtains the optimal estimate of α and β. At the simulation time point i, the α and β that best match the current data are trained and marked as and Get the predicted distribution of the calculated data and the predicted value at time i+1.
[0016] Preferably, the reinforcement learning-based feature extraction module estimates the expected reward under a set state using a Q function and optimizes the estimated value using temporal difference.
[0017] Preferably, S t is the environmental state, Q(S t ,AC t ) indicates that the agent (such as the decision module) chooses action AC under the remaining service life L t The maximum cumulative reward obtained after selecting and executing the action is the sum of the immediate reward obtained by the agent after selecting and executing the action and the value obtained by following the optimal strategy.
[0018] Preferably, Q(S t ,AC t )=RE(S t ,AC t )+γmax ac Q(S t+1 ,AC t+1 ), where γ∈[0,1].
[0019] Preferably, Q(S) is modified by a learning rate x t ,AC t ),satisfy,
[0020] Q(S t ,AC t )=(1-χ)Q(S t ,AC t )+χ(RE(S t ,AC t )+γmaxQ(S t+1 ,AC t+1 )).
[0021] Preferably, a cumulative distribution function with truncated conditions is introduced into the life span L of the supercapacitor and t l The expected value of the remaining service life of the supercapacitor is obtained by integrating; the estimated values of the optimal parameters of α and β are sought by maximizing the posterior probability; when new data is collected, the probability density function of the degradation trajectory is obtained, and then the expected value of the degradation trajectory is obtained, and the service life is estimated using the FHT method.
[0022] Preferably, the upper and lower confidence limits of the remaining service life of the supercapacitor are obtained based on the Neyman confidence interval.
[0023] An auxiliary module adopting the supercapacitor remaining service life prediction method based on QL-RRVM is provided, wherein the auxiliary module is arranged in cooperation with a supercapacitor management terminal to perform prediction and decision-making using the supercapacitor remaining service life prediction method based on QL-RRVM; so that the auxiliary module can meet the high-precision prediction requirements without affecting the normal operation of the supercapacitor.
[0024] The present invention relates to a method and an auxiliary module for predicting the remaining service life of a supercapacitor based on QL-RRVM, which constructs an improved relevance vector machine RRVM, and integrates a feature extraction module based on reinforcement learning in front of the improved relevance vector machine RRVM to obtain a QL-RRVM model for predicting degradation trends according to time series; the original data of the supercapacitor is collected, and the state of the supercapacitor at the current moment is obtained based on any action, and after inputting the data into the QL-RRVM model, the decision is guided based on the predicted target value and the expected return, and the remaining service life of the supercapacitor is dynamically estimated using the predicted degradation trend; the method is executed by using the auxiliary module in cooperation with a supercapacitor management terminal.
[0025] The beneficial effects of the present invention are:
[0026] (1) The model training and prediction process is efficient, simple, easy to implement, and applicable to a variety of degradation systems, with broad application prospects;
[0027] (2) By predicting the degradation trend, a complete degradation curve is obtained, which improves the interpretability of the prediction results;
[0028] (3) Compared with the traditional RVM algorithm, it can greatly improve the real-time performance and efficiency of the system while ensuring the prediction accuracy;
[0029] (4) Suppress the loss of effective information while achieving more efficient prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flow chart of the method of the present invention;
[0031] Figure 2 is a flow chart of a method in a specific implementation of the present invention;
[0032] Figure 3 The figure is an application flow chart of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific implementation described herein is only used to explain the present invention and is not intended to limit the present invention.
[0034] The present invention relates to a method for predicting the remaining useful life of a supercapacitor based on QL-RRVM, constructing an improved relevance vector machine RRVM, and integrating a feature extraction module based on reinforcement learning in front of the improved relevance vector machine RRVM to obtain a QL-RRVM model for predicting degradation trends according to time series;
[0035] The raw data of the supercapacitor is collected, and the state of the supercapacitor at the current moment is obtained based on any action. After inputting into the QL-RRVM model, the decision is guided based on the predicted target value and expected return, and the remaining service life of the supercapacitor is dynamically estimated using the predicted degradation trend.
[0036] In the present invention, RRVM is agreed to be a robust relevance vector machine, and QL is Q-Learning;
[0037] On the basis of completing RRVM modeling and parameter estimation, the degradation trend of the system is predicted, and the degradation trend characteristics are extracted and guided by QL. In the process, the predicted degradation trend is used to dynamically estimate the remaining service life of the supercapacitor.
[0038] In the present invention, the original data includes but is not limited to current, temperature, number of charging times, etc. The more types and characteristics of the original signals collected, the more accurate the final prediction result will be.
[0039] In the actual application process, two RRVM models are involved. One is trained based on the degradation trend characteristic data to predict the unknown degradation trend and obtain the complete degradation curve. The other RRVM model is trained and predicted based on the trend data to achieve an accurate and explainable remaining service life prediction process.
[0040] The method of the present invention comprises the following steps:
[0041] (1) Complete the RRVM model establishment and parameter estimation to predict the degradation trend of the system;
[0042] (2) The process of introducing the QL learning algorithm to extract the characteristics of degradation trends;
[0043] (3) Use the predicted degradation trend to estimate the remaining life of the supercapacitor.
[0044] The following is a detailed description of the three steps.
[0045] (1) Complete RRVM model establishment and parameter estimation to predict system degradation trends
[0046] Based on the Bayesian framework of relevance vector machine, an improved relevance vector machine RRVM is constructed with parameters following t distribution.
[0047] Building the improved relevance vector machine RRVM includes the following steps:
[0048] S1.1 define the prediction function y(x);
[0049]
[0050] in, is a design matrix, k(x i ,x k ) is the Gaussian kernel function, used to represent x i and x k The correlation between them, k = 1,…,N; is the weight matrix.
[0051] S1.2 Given a dataset about time series in, The predicted target value meets
[0052] t i =y(x i )+ε (2)
[0053] N represents the number of target groups, and L represents the step size;
[0054] Let ε have a mean of zero and a variance of σ 2 , introduce accuracy β and satisfy β -1 =σ 2 , we get the weights and β -1 The predicted value of t i The probability density function of
[0055]
[0056] Considering the N groups of data mentioned above, it can be expressed as:
[0057]
[0058] In the formula, T, X, and W are respectively i 、x i 、w i The matrix form of .
[0059] S1.3 defines the weight matrix W as a Gaussian distribution with zero mean and different variances, so that it satisfies the sparse prediction of the relevant vector machine, expressed as
[0060]
[0061] In the formula, the precision matrix α is set corresponding to the weight matrix, And α iExpressed as the precision corresponding to W, i∈[1,N+1].
[0062] S1.4 Estimate the parameters W, β and α by maximizing the probability density function,
[0063] p(W,α,β|t i )=p(W|t i ,α,β)p(α,β|t i ) (6)
[0064] The posterior probability distribution of parameter W remains Gaussian, then
[0065] p(W|t i ,α,β)=N(W|Λ,Σ) (7)
[0066] Where Λ=βΣΦ T T is the mean, Σ=(A+βΦ T Φ) -1 is the variance,
[0067] diag represents a diagonal matrix.
[0068] Considering that the assumed weights and random errors of the traditional relevance vector machine obey Gaussian distribution, it is sensitive to outliers and abnormal data points, resulting in poor robustness of the system. Therefore, here we let the parameters α and β obey gamma distribution, and the predicted target value obeys t distribution (more robust than Gaussian distribution), and get the corresponding probability density function,
[0069]
[0070] Among them, a and c are shape parameters, b and d are scale parameters; Γ(x|y,z)=Γ(y) -1 z y x y-1 e -zx is the gamma function,
[0071] Introducing the above expression into the probability density function, the posterior probability density function of the weight W and the random error ε is,
[0072]
[0073]
[0074] Among them, the parameters λ=a / b, λ′=c / d and ν=2a, v′=2c correspond to the accuracy and degree of freedom of the weight W and error ε respectively, and all obey t distribution.
[0075] S1.5 obtains the optimal estimate of α and β. At the simulation time point i, the α and β that best match the current data are trained and marked as and Get the predicted distribution of the calculated data and the predicted value at time i+1.
[0076] In fact, the optimal estimate of the current assumption parameters α and β has been achieved. At the simulation time point i, the α and β that best match the current data are trained and marked as and Get the predicted distribution of the calculated data,
[0077] p(t i |t i-1 )=∫p(t i |W,α,β)p(W,α,β|t i-1 )dWdαdβ (11)
[0078]
[0079] Substituting the probability density function and simplifying it, we get:
[0080]
[0081] Among them, y i =W T Φ(x i-1 )and Mean y i Considered data x i-1 The predicted value of , so the predicted value at time i+1 is,
[0082]
[0083] Among them, y i+1 =W T Φ(x i )and
[0084] At this point, the traditional support vector machine is improved to the RRVM framework, and the time series data will be predicted sequentially.
[0085] (2) The process of introducing the QL learning algorithm to extract degradation trend characteristics
[0086] In the present invention, the traditional feature extraction process is replaced by the QL algorithm, and the intelligent agent selects data processing actions to improve the accuracy and efficiency of prediction. QL optimizes feature selection by learning the relationship between actions and rewards to reduce the risk of information loss. Specifically, the selection of each action is based on a Q function, which estimates the expected return that can be obtained by selecting an action under a specific state. By repeatedly updating the value of the Q function, the optimal decision under different states is calculated to improve data processing efficiency. The estimated value is gradually optimized using time series difference. By calculating the cumulative distribution, the algorithm can better evaluate the potential benefits of each action and ultimately achieve the optimal choice of data processing.
[0087] In the decision-making process, the agent perceives the environment and performs actions from a series of possible choices. Assume that at time t, the current state of the environment is S t , the agent selects and executes action AC t , causing the state to change from S t To S t+1 , while providing the agent with a reward RE(S t ,AC t ). By repeating this process continuously, until the end of the training and learning process.
[0088] S t is the environmental state, Q(S t ,AC t ) indicates that the agent chooses action AC under the remaining service life L t The maximum cumulative reward obtained after the action is the sum of the immediate reward obtained by the agent after selecting and executing the action and the value obtained by following the optimal strategy; L = inf{t l :x(t k +t l )≡x k+l ∈B|X 1:k}.
[0089] Q(S t ,AC t )=RE(S t ,AC t )+γmax ac Q(S t+1 ,AC t+1 ) (15)
[0090] Among them, γ is the discount factor. During the training and learning process of the agent, it is always iteratively trained with the maximum Q value corresponding to the state, satisfying γ∈[0,1].
[0091] In multiple trainings, the Q value is continuously updated. However, in order to ensure that the QL algorithm converges at the appropriate time, an appropriate learning rate χ needs to be introduced in formula (15) to modify Q(S t ,AC t ),satisfy,
[0092] Q(S t ,AC t )=(1-χ)Q(S t ,AC t )+χ(RE(S t ,AC t )+γmaxQ(S t+1 ,AC t+1 )) (16)
[0093] In traditional machine learning methods, the feature extraction process is usually relatively simple, which easily leads to the loss of effective information and makes it difficult to extract features with high variability. To address this problem, the present invention achieves efficient and accurate feature extraction through intelligent decision-making. The reinforcement learning-based method is integrated into the front end of the RRVM model to provide the model with more accurate feature information, laying a solid foundation for the subsequent accurate prediction process.
[0094] (3) Using the predicted degradation trend to estimate the remaining life of the supercapacitor
[0095] In the present invention, in the process of predicting the remaining service life, a forward prediction method is adopted to dynamically adjust the prediction results according to the degradation trend in the time series, so as to further improve the accuracy and interpretability of the prediction; the moment when the predicted degradation exceeds the given threshold for the first time is defined as the failure moment, and the time from the current moment to the failure moment is the remaining service life at the current moment. In the known history and the observed data at the current moment {x 1 ,x 2 ,…,x k}, t k The remaining service life L at the moment can be defined based on the first hitting time as
[0096] L = inf{t l :x(t k +t l )≡x k+l ∈B|X 1:k} (17)
[0097] Among them, inf{·} represents the infimum of the set; t l Indicates that from t k The time from the moment to the expiration moment, X 1:k Indicates that from t 1 tok B represents a known bounded set whose elements are fault or failure thresholds. The specific values are usually determined by expert experience. k+l Indicates t k +t l Here, the degradation trend of the system is predicted based on the observed data set, and the obtained degradation trend is used to predict the RUL of the supercapacitor, so x k+l Indicates an intermediate process;
[0098] The cumulative distribution function of the supercapacitor life L (always greater than or equal to 0) is introduced with the truncated condition and t l integral,
[0099]
[0100]
[0101] Among them, △g k+l Indicates g k+l About time t k+l The derivative of , φ(·) represents the probability density function of a random variable that obeys a standard normal distribution;
[0102] Get the expected value of the remaining service life of the supercapacitor,
[0103]
[0104] The optimal parameters max(p(α,β|t i ), then the probability density function of the predicted value with respect to the hyperparameters α and β is,
[0105] p(t i |α,β)
[0106]
[0107] After taking the logarithm, we get the following likelihood function:
[0108] LF=lnp(t i |lnα,lnβ)+lnp(lnα)+lnp(lnβ) (22)
[0109] If we ignore the parts that are not related to α and β, the likelihood function LF is transformed into the following form:
[0110]
[0111] Find the partial derivatives of α and β respectively and set them to zero.
[0112]
[0113]
[0114] Through the above derivation, we get the estimated expressions of hyperparameters α and β:
[0115]
[0116] where γ i =1-α i Σ ii ,i=1,...,N+1,Σ ii represents the corresponding elements of the i-th row and i-th column of the matrix Σ;
[0117] When new data is collected, the probability density function of the degradation trajectory can be obtained by referring to formula (21), and then the expected value of the degradation trajectory can be obtained, and the service life can be estimated using the FHT method.
[0118] Considering the randomness of the remaining service life, the upper and lower confidence limits of the remaining service life of the supercapacitor are obtained based on the Neyman confidence interval, which specifically includes the following steps:
[0119] According to the expected value of the remaining useful life derived from the above formula (20), the relationship between the variance and the mean is expressed as The standard deviation of the remaining useful life σ can be obtained L ;
[0120] Construct a pivot quantity using the mean and standard deviation of the remaining useful life distribution Among them, μ L is the true value of the remaining useful life, and M represents the number of batches;
[0121] Given a confidence interval 1-α, use the formula That is, the confidence interval is z α is the critical value of the standard normal distribution.
[0122] Applying the above method to the specific prediction of the remaining service life of supercapacitors includes the following steps:
[0123] S1 collects the raw data of the supercapacitor, such as current, temperature, charging times, etc., and collects as many types and characteristics of the raw signals as possible;
[0124] The S2 agent selects an action from the action set according to the random strategy based on the obtained raw data, which is convenient for the subsequent Q value update and feature extraction;
[0125] The S3 agent processes the original signal at the current moment according to the randomly selected action, and obtains the current state St Import into the QL-RRVM framework;
[0126] S4 uses RRVM for training. First, according to the current state S t The matrix A, Λ and Σ are calculated by using the Gaussian kernel function, and then the optimal parameters at the current moment are determined according to formula (26): and Finally based on Then calculate the predicted value y i+1 =W T Φ(x i );
[0127] S5 combines the true value corresponding to the current moment with the predicted value corresponding to the moment to calculate the absolute value of the prediction error. In the specific calculation, the error between the training set and the test set is considered at the same time. AC j,i represents the jth action randomly selected at time i, and 1≤j≤6, is the predicted value of the i-th point in the training set, is the corresponding true value, which is also applicable to the test set;
[0128] S6 saves the current Q value and passes it to the agent, which then compares whether it reaches the threshold of the relevant parameters or the number of iterations. The iterative process repeats the above steps until all actions are traversed or the threshold is reached;
[0129] S7 summarizes the Q values obtained after executing all actions into the Q table, and processes the original data accordingly based on the data in the table. By using the Neyman confidence interval, the upper and lower limits of the remaining service life can be obtained, and the prediction results can be obtained.
[0130] The present invention also relates to an auxiliary module that adopts the supercapacitor remaining service life prediction method based on QL-RRVM. The auxiliary module is arranged in conjunction with a supercapacitor management terminal to perform prediction and decision-making using the supercapacitor remaining service life prediction method based on QL-RRVM; so that it can meet high-precision prediction requirements without affecting the normal operation of the supercapacitor.
[0131] Specifically, raw data, including but not limited to voltage, current, charging times, etc., are collected from the sensors and monitoring modules of the supercapacitor management terminal to form a data set. An external auxiliary module is used to specifically perform state prediction and decision-making tasks. The data set formed by the sensor and inspection module is used to directly obtain the real-time data of the supercapacitor through the system interface, and this data information is input into the model for analysis and processing.
[0132] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0134] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0135] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0136] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0137] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A method for predicting the remaining useful life of a supercapacitor based on QL-RRVM, characterized in that: An improved relevance vector machine RRVM is constructed, and a feature extraction module based on reinforcement learning is integrated before the improved relevance vector machine RRVM to obtain a QL-RRVM model for predicting degradation trends according to time series; The raw data of the supercapacitor is collected, and the state of the supercapacitor at the current moment is obtained based on any action. After inputting into the QL-RRVM model, the decision is guided based on the predicted target value and expected return, and the remaining service life of the supercapacitor is dynamically estimated using the predicted degradation trend.
2. The method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 1, characterized in that: Based on the Bayesian framework of relevance vector machine, an improved relevance vector machine RRVM is constructed with parameters following t distribution.
3. A method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 1 or 2, characterized in that: Building the improved relevance vector machine RRVM includes the following steps: S1.1 define the prediction function y(x); S1.2 Given a dataset about time series The predicted target value satisfies t i =y(x i )+ε, where x i =(x i ,x i+1 ,…,x i+L-1 ), N represents the number of target groups, L represents the step size; let ε obey the mean of zero and the variance of σ 2 , introduce accuracy β and satisfy β -1 =σ 2 , we get the weights and β -1 The predicted value of t i The probability density function of S1.3 defines the weight matrix W as a Gaussian distribution with zero mean and different variances, and sets the precision matrix α corresponding to the weight matrix; S1.4 Estimate parameters W, β and α by maximizing the probability density function, and keep the posterior probability distribution of parameter W as Gaussian distribution; let parameters α and β obey gamma distribution, and predict the target value to obey t distribution, and obtain the corresponding probability density function, and the posterior probability density function of weight W and random error ε; S1.5 obtains the optimal estimate of α and β. At the simulation time point i, the α and β that best match the current data are trained and marked as and Get the predicted distribution of the calculated data and the predicted value at time i+1.
4. The method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 1, characterized in that: The reinforcement learning-based feature extraction module estimates the expected reward under a set state using a Q function and optimizes the estimated value using temporal difference.
5. The method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 4, characterized in that: S t is the environmental state, Q(S t ,AC t ) indicates that the agent chooses action AC under the remaining service life L t The maximum cumulative reward obtained after selecting and executing the action is the sum of the immediate reward obtained by the agent after selecting and executing the action and the value obtained by following the optimal strategy.
6. The method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 5, characterized in that: Q(S t , AC t ) = RE(S t , AC t ) + γmax ac Q(S t+1 , AC t+1 ), where γ ∈ [0, 1].
7. The method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 6, characterized in that: Modify Q(S) with learning rate χ t ,AC t ), satisfying, Q(S t ,AC t )=(1-χ)Q(S t ,AC t )+χ(RE(S t ,AC t )+γmaxQ(S t+1 ,AC t+1 )).
8. The method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 3, characterized in that: The cumulative distribution function of the supercapacitor life L is introduced with the truncated condition and t l The expected value of the remaining service life of the supercapacitor is obtained by integrating; the estimated values of the optimal parameters of α and β are sought by maximizing the posterior probability; when new data is collected, the probability density function of the degradation trajectory is obtained, and then the expected value of the degradation trajectory is obtained, and the service life is estimated using the FHT method.
9. The method for predicting the remaining useful life of a supercapacitor based on QL-RRVM according to claim 8, characterized in that: The upper and lower confidence limits of the remaining service life of the supercapacitor are obtained based on the Neyman confidence interval.
10. An auxiliary module that uses the QL-RRVM-based supercapacitor remaining useful life prediction method as described in any one of claims 1 to 9, wherein the auxiliary module is configured in conjunction with a supercapacitor management terminal to perform prediction and decision-making using the QL-RRVM-based supercapacitor remaining useful life prediction method.