An Online Prediction Method for Power Battery Manufacturing Capacity Based on Reinforcement Learning
Through a combined prediction model based on reinforcement learning, the weight and hidden layers of the recurrent neural network and long-term memory network are optimized, and the problem of inaccurate prediction of lithium-ion battery manufacturing capacity in the prior art is solved, and a higher accuracy and reliability prediction of battery manufacturing capacity is achieved.
Patent Information
- Application Number
- CN202210098257.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-01-19
AI Technical Summary
The existing prediction models cannot accurately predict the production capacity in the manufacturing process of lithium-ion batteries, resulting in improper manufacturing cycle arrangement and affecting corporate efficiency.
Using reinforcement learning-based methods, recurrent neural networks and long-term memory network models are constructed, and by optimizing the weight and hidden layers of a single prediction model, a combined prediction model is constructed to achieve online accurate prediction of battery manufacturing capabilities.
It improves the accuracy and reliability of battery manufacturing capacity prediction, helps enterprises reasonably arrange manufacturing cycles and improve economic benefits.
Smart Images

Figure CN114418234B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an online prediction method for power battery manufacturing capacity based on reinforcement learning, and belongs to the field of power battery manufacturing prediction. Background Art
[0002] In recent years, as an upgraded product of nickel-metal hydride and lead-acid batteries, lithium-ion batteries have the characteristics of high energy density, high rate, high safety, etc., and have become the focus of current technology research and industrialization. At the same time, battery manufacturing technology has developed from workshop production to automation and today's intelligentization, and the industrial scale has been continuously expanding. China has become the world's largest production and consumption place for lithium-ion batteries.
[0003] The manufacturing process of batteries is a highly complex system, and the battery manufacturing cycle is uncertain. For example, the production process of lithium batteries mainly includes mixing, coating, rolling, slitting, stacking, welding, injection, formation, aging, assembly, etc.; the time required for each process of each type of battery is uncertain. Then, for enterprises, if they can predict the time required for the battery manufacturing process, it can help enterprises reasonably arrange the manufacturing cycle and maximize economic benefits. Therefore, studying the battery manufacturing capacity prediction method has important practical significance for improving battery manufacturing efficiency and enhancing enterprise revenue.
[0004] Most of the existing prediction models are for objects such as stocks and short-term traffic flow. The value of these objects at a certain future moment usually has a certain relationship with the current value, so the existing prediction models can accurately predict them; however, these existing prediction models cannot accurately predict the battery manufacturing capacity, mainly because there is no linear relationship between the production capacity at the current moment and the production capacity at a certain future moment during the battery production process, and the production time required for the same process of different types of batteries is different, and there is no reference between each process. Therefore, there is a large deviation when using the existing prediction models to analyze and predict the battery manufacturing capacity. Summary of the Invention
[0005] In order to achieve accurate prediction of battery manufacturing capacity to help enterprises reasonably arrange the manufacturing cycle and improve economic benefits, the present invention provides an online prediction method for power battery manufacturing capacity based on reinforcement learning. When performing online prediction of battery manufacturing capacity, the method first defines a new combined prediction form, then uses reinforcement learning to obtain the optimal hidden layer of the recurrent neural network model and the long short-term memory network model, and then uses reinforcement learning to obtain the optimal weight of the combined prediction model. Finally, a combined prediction model for power battery manufacturing capacity based on reinforcement learning is obtained, realizing accurate online prediction of battery manufacturing capacity.
[0006] An online prediction method for the manufacturing capacity of power batteries based on reinforcement learning, the method comprising:
[0007] Using m single prediction models to respectively predict the manufacturing capacity of power batteries within a future time period of length N, and obtaining the primary prediction results of each single prediction model;
[0008] Based on the primary prediction results of each single prediction model, optimizing each single prediction model based on reinforcement learning;
[0009] Using the optimized single prediction models to respectively re-predict the manufacturing capacity of power batteries within a future time period of length N, and obtaining the secondary prediction results of each single prediction model;
[0010] Based on the secondary prediction results of each single prediction model, using reinforcement learning to divide the time period of length N to determine the optimal weight of each single prediction model, thereby determining the expression of the combined prediction model:
[0011]
[0012]
[0013] wherein, w i is the weight of the i-th single prediction model, N i is the demarcation point of the i-th single prediction model, f ij represents the j-th prediction vector value of the i-th single prediction model, represents rounding down the numerical value a;
[0014] w i satisfies:
[0015]
[0016] Obtaining the predicted value of the manufacturing capacity of the power battery according to the expression of the combined prediction model determined by the optimal weights of each single prediction model.
[0017] Optionally, the single prediction models include a recurrent neural network model and a long short-term memory network model, m = 2; the method comprising:
[0018] Step 1: Defining the production volume obtained per unit time in the battery manufacturing process to represent the manufacturing capacity of the power battery;
[0019] Step 2: Define the number of iterations \(k1 = 1\) for the recurrent neural network model, \(k2 = 1\) for the long short-term memory network model, \(k3 = 1\) for the weight learning iteration, and the corresponding iteration steps \(N1\), \(N2\), \(N3\). The Q1 table of the recurrent neural network model representing state-action, the Q2 table of the long short-term memory network model, and the Q3 table of weight learning are all 0; the optimal initial value \(l\) of the hidden layer of the recurrent neural network model 1,0 and the optimal initial value \(l\) of the hidden layer of the long short-term memory network model 2,0 , the action matrix \(A1\) of the recurrent neural network model, the action matrix \(A2\) of the long short-term memory network model, and the action matrix \(A3\) of weight learning, the weights \(w1\), \(w2\) of the combined prediction model, and calculate the first prediction results of the recurrent neural network model and the long short-term memory network model;
[0020] Step 3: Use reinforcement learning to construct the hidden layer learning environment, establish the loss function \(L1\) of the recurrent neural network and the loss function \(L2\) of the long short-term memory network, the reward and punishment function \(R1\) of the recurrent neural network and the reward and punishment function \(R2\) of the long short-term memory network, and select action behaviors using the greedy algorithm according to the current state, Q1 table, and Q2 table;
[0021] Step 4: Calculate the loss functions \(L1\) and \(L2\), the reward and punishment functions \(R1\) and \(R2\) according to the first prediction results in Step 2, and update the Q1 table and Q2 table;
[0022] Step 5: Let \(k1 = k1 + 1\) and \(k2 = k2 + 1\), return to Step 3 until \(k1 = N1\) and output the optimal hidden layer \(l1\) of the recurrent neural network; when \(k2 = N2\), output the optimal hidden layer \(l2\) of the long short-term memory network, and jump to Step 6 after outputting the two hidden layers;
[0023] Step 6: Substitute the optimal hidden layer numbers \(l1\) and \(l2\), and recalculate the second prediction results of the recurrent neural network and the long short-term memory network;
[0024] Step 7: Construct the weight learning environment, set the state matrix \(S1\) and the action matrix \(A3\) of the battery manufacturing capacity combined prediction model, and establish the weight learning loss function \(L3\) and the reward and punishment function \(R3\), and select action behaviors using the greedy algorithm according to the current state and Q3 table;
[0025] Step 8: Calculate the loss function \(L3\) and the reward and punishment function \(R3\) according to the second prediction results in Step 6, and update the Q3 table;
[0026] Step 9: Let \(k3 = k3 + 1\), return to Step 7 until \(k3 = N3\) and output the recurrent neural network weight \(w1\) and the long short-term memory network weight \(w2\), and jump to Step 10;
[0027] Step Ten: According to the output result of Step Nine, construct a combined prediction model for power battery manufacturing capacity based on reinforcement learning, input the historical data of battery manufacturing capacity into the constructed combined prediction model for power battery manufacturing capacity based on reinforcement learning, and output the predicted value of battery manufacturing capacity.
[0028] Optionally, in the primary prediction results of the recurrent neural network model and the long short-term memory network model calculated in Step Two, the output of the recurrent neural network model at time t is:
[0029]
[0030] where x t is the input of the system at time t, that is, the historical data of power battery manufacturing capacity, s t and o t are the outputs of the hidden layer and the output layer at time t. The output of the output layer is the predicted value of power battery manufacturing capacity. U is the weight of the hidden layer, V is the weight of the output layer, and W represents the weight of the previous value of the hidden layer as the current input; g and f are activation functions;
[0031] The output of the long short-term memory network model at time t is:
[0032]
[0033] where h t represents the short-term historical information at time t, C t represents the long-term historical information at time t, is the candidate long-term historical information at time t, x t represents the input sample at time t, σ is the Sigmoid activation function, tanh is the hyperbolic tangent activation function, W f and b f are the weight matrix and bias vector of the forget gate respectively, f t is the output of the long short-term memory network forget gate at time t, W i and b i are the weight matrix and bias vector of the input gate respectively, i t is the output of the long short-term memory network input gate at time t, W o and b o are the weight matrix and bias vector of the output gate respectively, O t is the output of the long short-term memory network output gate at time t, that is, the predicted value of power battery manufacturing capacity;
[0034] Substitute the optimal initial value l 1,0 of the hidden layer of the recurrent neural network and the optimal initial value l 2,0 of the hidden layer of the long short-term memory network to calculate the primary prediction results y1 and y2.
[0035] Optionally, step three includes:
[0036] Define variables l1 and l2 to represent the number of hidden layers of the recurrent neural network model and the long short-term memory network model respectively; define action matrices A1 and A2:
[0037] A1 = A2 = [Δv1 -Δv1] (7)
[0038] where Δv1 is the amplitude of actions A1 and A2; construct the Q1 table of the recurrent neural network model and the Q2 table of the long short-term memory network model with the target hidden layer variable as the row vector and the action state as the column vector, and set the loss functions L1 and L2 and the reward and punishment functions R1 and R2 as follows:
[0039]
[0040]
[0041]
[0042]
[0043] where y 1,t and y 2,t represent the true values of the battery manufacturing capabilities of the recurrent neural network model and the long short-term memory network model at time t respectively, and represent the predicted values of the battery manufacturing capabilities of the recurrent neural network model and the long short-term memory network model at time t respectively, L 1,t and L 2,t are the loss values of the recurrent neural network model and the long short-term memory network model at time t respectively, L 1,t+1 and L 1,t+1 are the loss values of the recurrent neural network model and the long short-term memory network model at time t+1 respectively, and N is the output sample length;
[0044] The selection mechanism of A1 is
[0045]
[0046] where represents the action corresponding to the maximum Q value in the Q1 table, A R represents randomly selecting an action from the action matrix A1, δ1 is a random number between [0,1], and ε1 is the greed rate of action A1;
[0047] The selection mechanism of the action matrix A2 is:
[0048]
[0049] Among them, represents the action corresponding to the maximum Q value in the Q2 table, A R2 represents randomly selecting an action from the action matrix A2, δ2 is a random number between [0, 1], and ε2 is the greed rate of the action A2.
[0050] Optionally, the fourth step includes:
[0051] The formula for updating the Q1 table is as follows:
[0052] Q 1,t+1 (l 1,t , A 1,t ) = Q 1,t (l 1,t , A 1,t ) + α1[R1(l 1,t , A 1,t ) + γ1maxQ 1,t (l 1,t+1 , A 1,t+1 ) - Q 1,t (l 1,t , A 1,t )] (14)
[0053] Among them, Q 1,t (l 1,t , A 1,t ) represents the Q 1,t table constructed by using l1 and A1 at time t as the row and column respectively, α1 is the update learning rate of the Q1 table, and γ1 is the discount factor of the Q1 table;
[0054] The formula for updating the Q2 table is as follows:
[0055] Q 2,t+1 (l 2,t , A 2,t ) = Q 2,t (l 2,t , A 2,t ) + α2[R2(l 2,t , A 2,t ) + γ2maxQ 2,t (l 2,t+1 , A 2,t+1 ) - Q 2,t (l 1,t , A 2,t )] (15)
[0056] Among them, Q 2,t (l 2,t , A 2,t ) represents the Q 2,tThe table, α2 is the update learning rate of Q2 table, and γ2 is the discount factor of Q2 table; Define the number of training iterations as N1 and N2 respectively, and train to obtain the optimal hidden layers l1 and l2.
[0057] Optionally, step seven includes:
[0058] The state matrix S1 and the action matrix A3 are as follows:
[0059] S1 = [w1 -w1] (16)
[0060] A3 = [Δv2 -Δv2] (17)
[0061] Where, w1 is the weight of the recurrent neural network model, w2 is the weight of the long short-term memory network model; △v2 is the amplitude of action A3; At the same time, establish a combined prediction model Q3 table with rows representing the target state and columns representing the action state, and set the loss function L3 and the reward and punishment function R3:
[0062]
[0063]
[0064] Where, y 3,t represents the true value of the battery manufacturing capacity of the combined prediction model at time t, respectively represent the predicted values of the battery manufacturing capacity of the combined prediction model at time t, L 3,t is the loss value of the combined prediction model at time t, L 3,t+1 is the loss value of the combined prediction model at time t+1, and N is the output sample length;
[0065] Set the A3 selection mechanism as:
[0066]
[0067] Where, represents the action corresponding to the maximum Q value in the Q3 table, A r represents randomly selecting an action from the action matrix A3, δ3 is a random number between [0,1], and ε3 is the greed rate of action A3.
[0068] Optionally, step eight includes:
[0069] Calculate the loss function L3 and the reward and punishment function R3 according to the secondary prediction result in step six and equations (18) and (19);
[0070] Update the Q3 table according to the following formula:
[0071]
[0072] Where, Q3,t (S 1,t , A 3,t ) represents the Q table constructed by using S1 and A3 at time t as the row and column respectively. α3 is the update learning rate of the Q3 table, and γ3 is the discount factor of the Q3 table. 3,t Table, α3 is the update learning rate of the Q3 table, and γ3 is the discount factor of the Q3 table.
[0073] Optionally, the ninth step includes:
[0074] Define the number of training iterations as N3 respectively; let k3 = k3 + 1, return to the seventh step until k3 = N3, output the weights w1 of the recurrent neural network and the weights w2 of the long short-term memory network, and jump to the tenth step.
[0075] Optionally, the tenth step includes:
[0076] After obtaining the weights w1 of the recurrent neural network and the weights w2 of the long short-term memory network, according to the battery manufacturing capacity combination prediction forms given by formulas (2) and (3), construct a prediction model for the battery manufacturing capacity based on reinforcement learning, and output the predicted value of the battery manufacturing capacity.
[0077] The beneficial effects of the present invention are:
[0078] By proposing a new combination prediction model form to reasonably segment and combine the predicted values of different prediction methods. Secondly, compared with traditional neural networks that often set the number of hidden layers using empirical data and it is difficult to achieve the best adaptation to the battery manufacturing capacity prediction model, this application uses reinforcement learning to construct a hidden layer learning environment for the recurrent neural network and the long short-term memory network model, solves the optimal number of hidden layers of the network model, reduces the prediction deviation of the hidden layer; furthermore, constructs a weight learning environment for the combination model, obtains the optimal weights after iterative training, and finally constructs a prediction model for the battery manufacturing capacity combination, further improving the prediction accuracy and reliability for the prediction of the battery manufacturing capacity. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0080] Figure 1 is a flowchart of an online prediction method for the battery manufacturing capacity based on reinforcement learning disclosed in an embodiment of the present invention.
[0081] Figure 2 is a structural diagram of an online prediction method for the battery manufacturing capacity based on reinforcement learning disclosed in an embodiment of the present invention.
[0082] Figure 3 It is a graph of the online prediction results of the power battery manufacturing capacity using the method of the present application and four existing methods disclosed in an embodiment of the present invention.
[0083] Figure 4 It is a graph of the online prediction error results of the power battery manufacturing capacity using the method of the present application and four existing methods disclosed in an embodiment of the present invention. Detailed implementation mode
[0084] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0085] Embodiment 1:
[0086] This embodiment provides an online prediction method for the power battery manufacturing capacity based on reinforcement learning. The method uses m single prediction models to respectively predict the power battery manufacturing capacity within a future time period of length N, and obtains the primary prediction results of each single prediction model.
[0087] Based on the primary prediction results of each single prediction model, each single prediction model is optimized based on reinforcement learning.
[0088] The optimized single prediction models are used to respectively predict the power battery manufacturing capacity within a future time period of length N again, and obtain the secondary prediction results of each single prediction model.
[0089] Based on the secondary prediction results of each single prediction model, reinforcement learning is used to divide the time period of length N to determine the optimal weight of each single prediction model, so as to determine the expression of the combined prediction model:
[0090]
[0091]
[0092] where w i is the weight of the i-th single prediction model, N i is the demarcation point of the i-th single prediction model, f ij represents the j-th prediction vector value of the i-th single prediction model, represents rounding down the numerical value a;
[0093] w i satisfies:
[0094]
[0095] The predicted value of the power battery manufacturing capacity is obtained according to the expression of the combined prediction model determined by the optimal weights of each single prediction model.
[0096] See Figure 1 , the method includes:
[0097] Step 1: Define the battery manufacturing capacity and the form of the new combined prediction model;
[0098] Step 2: Initialize the parameters of the recurrent neural network model, long short-term memory network model, and reinforcement learning model.
[0099] Define the number of iterations of the recurrent neural network \(k1 = 1\), the number of iterations of the long short-term memory network \(k2 = 1\), the number of iterations of weight learning \(k3 = 1\), and the corresponding iteration steps \(N1\), \(N2\), \(N3\). The recurrent neural networks \(Q1\), long short-term memory networks \(Q2\), and weight learning \(Q3\) representing state behaviors are all 0; the optimal initial value \(l\) of the hidden layer of the recurrent neural network 1,0 and the optimal initial value \(l\) of the hidden layer of the long short-term memory network 2,0 , the action matrix \(A1\) of the recurrent neural network, the action matrix \(A2\) of the long short-term memory network, and the action matrix \(A3\) of weight learning, the weights \(w1\), \(w2\) of the combined model, and calculate the first prediction results of the recurrent neural network and the long short-term memory network;
[0100] Step 3: Use reinforcement learning to construct the hidden layer learning environment, establish the loss function \(L1\) of the recurrent neural network and the loss function \(L2\) of the long short-term memory network, the reward and punishment function \(R1\) of the recurrent neural network and the reward and punishment function \(R2\) of the long short-term memory network, and select action behaviors using the greedy algorithm according to the current state, \(Q1\) table, and \(Q2\) table;
[0101] Step 4: Calculate the loss functions \(L1\) and \(L2\), the reward and punishment functions \(R1\) and \(R2\) according to the first prediction results in Step 2, and update the \(Q1\) table and \(Q2\) table;
[0102] Step 5: Let \(k1 = k1 + 1\) and \(k2 = k2 + 1\), return to Step 3 until \(k1 = N1\) and output the optimal hidden layer \(l1\) of the recurrent neural network; when \(k2 = N2\), output the optimal hidden layer \(l2\) of the long short-term memory network, and jump to Step 6 after outputting the two hidden layers;
[0103] Step 6: Substitute the optimal hidden layer numbers \(l1\) and \(l2\) and recalculate the second prediction results of the recurrent neural network and the long short-term memory network;
[0104] Step 7: Construct the weight learning environment, set the state matrix \(S1\) and action matrix \(A3\) of the battery manufacturing capacity prediction model, and establish the weight learning loss function \(L3\) and reward and punishment function \(R3\). Select action behaviors using the greedy algorithm according to the current state and \(Q3\) table;
[0105] Step 8: Calculate the loss function L3 and the reward and punishment function R3 based on the secondary prediction result in Step 6, and update the Q3 table;
[0106] Step 9: Let k3 = k3 + 1, and return to Step 7 until k3 = N3, then output the weights w1 of the recurrent neural network and the weights w2 of the long short-term memory network, and jump to Step 10;
[0107] Step 10: Construct a combined prediction model for power battery manufacturing capacity based on the output result in Step 9, and output the predicted value of battery manufacturing capacity.
[0108] Embodiment 2
[0109] This embodiment provides an online prediction method for power battery manufacturing capacity based on reinforcement learning, and the method includes:
[0110] Step 1: Define battery manufacturing capacity and a new form of combined prediction model:
[0111] Define battery manufacturing capacity and a new form of combined prediction model as follows:
[0112] The prediction target in this article is battery manufacturing capacity, and the production volume (CycleTime, CT) obtained per unit time is used as the manufacturing capacity evaluation index:
[0113] CT = P T / P N (1)
[0114] In the formula, P T represents the factory manufacturing time, and P N represents the number of workpieces manufactured by the factory.
[0115] If there are m single prediction models in the set battery manufacturing capacity prediction combination model and the length of the prediction result vector is N, then the form of the combined prediction model given in this application is expressed as:
[0116]
[0117]
[0118] In the formula, w i is the weight of the i-th prediction method, N i is the demarcation point of the i-th prediction method, f ij represents the j-th prediction vector value of the i-th prediction method, represents rounding down the numerical value a. The demarcation point between prediction methods is determined by multiplying the weight by the length of the prediction result vector and then rounding down, so as to obtain a new form of combined prediction representation, which can reasonably segment and combine the predicted values of different prediction methods.
[0119] Combined prediction weight w i Satisfy the following constraints
[0120]
[0121] Note: If the weight w appears i = 1, it means that the obtained combined model is the same as one of the single models; if the weight w appears i = 0, then the rationality of the i-th single model needs to be reconsidered.
[0122] Step 2: Initialize the parameters of the recurrent neural network model, long short-term memory network model, and reinforcement learning model. Define the number of iterations of the recurrent neural network k1 = 1, the number of iterations of the long short-term memory network k2 = 1, the number of iterations of weight learning k3 = 1, and the corresponding iteration steps N1, N2, N3. The recurrent neural networks Q1, long short-term memory networks Q2, and weight learning Q3 representing state actions are all 0. The optimal initial value l of the hidden layer of the recurrent neural network 1,0 and the optimal initial value l of the hidden layer of the long short-term memory network 2,0 , the action matrix A1 of the recurrent neural network, the action matrix A2 of the long short-term memory network, and the action matrix A3 of weight learning, the combined model weights w1, w2, and calculate the first prediction results of the recurrent neural network and the long short-term memory network;
[0123] The parameters of the recurrent neural network model are defined in Table 1 below:
[0124] Table 1: Parameters of the recurrent neural network model
[0125]
[0126] The parameters of the long short-term memory network model are defined in Table 2 below
[0127] Table 2: Parameters of the long short-term memory network model
[0128]
[0129] The parameters of the reinforcement learning model are defined in Table 3 below
[0130] Table 3: Parameters of the reinforcement learning model
[0131]
[0132]
[0133] The output of the recurrent neural network model at time t is:
[0134]
[0135] Among them, xt is the input of the system at time t, s t and o t are the outputs of the hidden layer and the output layer at time t, U is the weight of the hidden layer, V is the weight of the output layer, and W represents the weight of the previous value of the hidden layer as the input of this time. g and f are activation functions
[0136] The output of the long short-term memory network model at time t is
[0137]
[0138] where h t represents the short-term historical information at time t, C t represents the long-term historical information at time t, is the candidate long-term historical information at time t, x t represents the input sample at time t, σ is the Sigmoid activation function, tanh is the hyperbolic tangent activation function, W f and b f are the weight matrix and the bias vector of the forget gate respectively, f t is the output of the long short-term memory network forget gate at time t, W i and b i are the weight matrix and the bias vector of the input gate respectively, i t is the output of the long short-term memory network input gate at time t, W o and b o are the weight matrix and the bias vector of the output gate respectively, O t is the output of the long short-term memory network output gate at time t.
[0139] Substitute the optimal initial value l of the hidden layer of the recurrent neural network 1,0 and the optimal initial value l of the hidden layer of the long short-term memory network 2,0 Calculate to obtain the first prediction results y1 and y2.
[0140] Step 3: Use reinforcement learning to construct the hidden layer learning environment, establish the loss function L1 of the recurrent neural network and the loss function L2 of the long short-term memory network, the reward and punishment function R1 of the recurrent neural network and the reward and punishment function R2 of the long short-term memory network. According to the current state, Q1 table and Q2 table, use the greedy algorithm to select action behaviors:
[0141] First, define variables l1 and l2 to represent the number of layers of the hidden layer of the recurrent neural network model and the number of layers of the hidden layer of the long short-term memory network model respectively. Secondly, define the action matrices A1 and A2
[0142] A1 = A2 = [Δv1 -Δv1] (7)
[0143] Among them, △v1 is the amplitude of actions A1 and A2. Using the target hidden layer variable as a row vector and the action state as a column vector, the Q1 table of the recurrent neural network model and the Q2 table of the long short-term memory network model are constructed respectively, and the loss functions L1 and L2 and the reward and punishment functions R1 and R2 are set as follows:
[0144]
[0145]
[0146]
[0147]
[0148] Among them, y 1,t and y 2,t respectively represent the true values of the battery manufacturing capabilities of the recurrent neural network model and the long short-term memory network model at time t, and respectively represent the predicted values of the battery manufacturing capabilities of the recurrent neural network model and the long short-term memory network model at time t, L 1,t and L 2,t are respectively the loss values of the recurrent neural network model and the long short-term memory network model at time t, L 1,t+1 and L 1,t+1 are respectively the loss values of the recurrent neural network model and the long short-term memory network model at time t + 1, and N is the length of the output sample.
[0149] This application sets the A1 selection mechanism as
[0150]
[0151] Among them, represents the action corresponding to the maximum Q value in the Q1 table, A R represents randomly selecting an action from the action matrix A1, δ1 is a random number between [0, 1], and ε1 is the greed rate of action A1.
[0152] The selection mechanism of the action matrix A2 is:
[0153]
[0154] Among them, represents the action corresponding to the maximum Q value in the Q2 table, A R2 represents randomly selecting an action from the action matrix A2, δ2 is a random number between [0, 1], and ε2 is the greed rate of action A2;
[0155] Step 4: Calculate the loss functions L1 and L2, the reward and punishment functions R1 and R2 based on the first prediction result in Step 2, and update the Q1 table and Q2 table:
[0156] The formula for updating the Q1 table is as follows
[0157] Q 1,t+1 (l 1,t ,A 1,t ) = Q 1,t (l 1,t ,A 1,t ) + α1[R1(l 1,t ,A 1,t ) + γ1maxQ 1,t (l 1,t+1 ,A 1,t+1 ) - Q 1,t (l 1,t ,A 1,t )] (14)
[0158] Among them, Q 1,t (l 1,t ,A 1,t ) represents the Q 1,t table constructed by using l1 and A1 at time t as the row and column respectively. α1 is the update learning rate of the Q1 table, and γ1 is the discount factor of the Q1 table;
[0159] The formula for updating the Q2 table is as follows:
[0160] Q 2,t+1 (l 2,t ,A 2,t ) = Q 2,t (l 2,t ,A 2,t ) + α2[R2(l 2,t ,A 2,t ) + γ2maxQ 2,t (l 2,t+1 ,A 2,t+1 ) - Q 2,t (l 1,t ,A 2,t )] (15)
[0161] Among them, Q 2,t (l 2,t ,A 2,t ) represents the Q 2,t table constructed by using l2 and A2 at time t as the row and column respectively. α2 is the update learning rate of the Q2 table, and γ2 is the discount factor of the Q2 table; Define the number of training iterations as N1 and N2 respectively, and train to obtain the optimal hidden layers l1 and l2.
[0162] Step 5: Let \(k1 = k1 + 1\) and \(k2 = k2 + 1\), and return to Step 3 until \(k1 = N1\), then output the optimal hidden layer \(l1\) of the recurrent neural network; when \(k2 = N2\), output the optimal hidden layer \(l2\) of the long short-term memory network. After outputting the two hidden layers, jump to Step 6:
[0163] When \(k1 = N1\), the optimal hidden layer \(l1\) of the recurrent neural network output is 16; when \(k2 = N2\), the optimal hidden layer \(l2\) of the long short-term memory network output is 152.
[0164] Step 6: Substitute the optimal hidden layer numbers \(l1\) and \(l2\) to recalculate the secondary prediction results of the recurrent neural network and the long short-term memory network: After substituting the optimal hidden layer numbers \(l1\) and \(l2\) into the two models respectively, calculate the secondary prediction results of the recurrent neural network and the long short-term memory network according to Equations (5) and (6). and
[0165] Step 7: Construct a weight learning environment, set the state matrix \(S1\) and action matrix \(A3\) of the battery manufacturing capacity prediction model, and establish a loss function \(L3\) and a reward and punishment function \(R3\). According to the current state and the \(Q3\) table, use the greedy algorithm to select action behaviors: The state matrix \(S1\) and action matrix \(A3\) are as follows
[0166] \(S1 = [w1 -w1] (16)\)
[0167] \(A3 = [Δv2 -Δv2] (17)\)
[0168] where \(w1\) is the weight of the recurrent neural network model, and \(w2\) is the weight of the long short-term memory network model. \(△v2\) is the amplitude of action \(A3\). At the same time, establish a combined prediction model \(Q3\) table with the target state represented by rows and the action state represented by columns, and set the loss function \(L3\) and the reward and punishment function \(R3\):
[0169]
[0170]
[0171] where \(y\) 3,t represents the true value of the battery manufacturing capacity of the combined prediction model at time \(t\), respectively represent the predicted values of the battery manufacturing capacity of the combined prediction model at time \(t\), \(L\) 3,t is the loss value of the combined prediction model at time \(t\), \(L\) 3,t+1 is the loss value of the combined prediction model at time \(t + 1\), and \(N\) is the output sample length.
[0172] This paper sets the A3 selection mechanism as:
[0173]
[0174] Among them, represents the action corresponding to the maximum Q value in the Q3 table, A r represents randomly selecting an action in the action matrix A3, δ3 is a random number between [0, 1], and ε3 is the greed rate of action A3.
[0175] Step Eight: Calculate the loss function L3 and the reward and punishment function R3 according to the secondary prediction result in Step Six and Equations (18) and (19), and update the Q3 table;
[0176] The formula for updating the Q3 table is as follows
[0177]
[0178] Among them, Q 3,t (S 1,t , A 3,t ) represents the Q 3,t table constructed by using S1 at time t and A3 as the row and column respectively. α3 is the update learning rate of the Q3 table, and γ3 is the discount factor of the Q3 table. Define the number of training iterations as N3 respectively, and train to obtain the optimal combined weights w1 and w2.
[0179] Step Nine: Let k3 = k3 + 1, return to Step Seven, and until k3 = N3, output the weights w1 of the recurrent neural network and the weights w2 of the long short-term memory network, and jump to Step Ten;
[0180] When k3 = N3, output the weights w1 of the recurrent neural network and the weights w2 of the long short-term memory network, and jump to Step Ten.
[0181] Step Ten: According to the output result of Step Nine, construct a combined prediction model for the power battery manufacturing capacity based on reinforcement learning, and output the predicted value of the battery manufacturing capacity.
[0182] After obtaining the weights w1 of the recurrent neural network and the weights w2 of the long short-term memory network, according to the combined prediction form of the battery manufacturing capacity proposed in Equations (2) and (3), construct a combined prediction model for the power battery manufacturing capacity based on reinforcement learning, and output the predicted value of the battery manufacturing capacity.
[0183] To evaluate the prediction performance of the method of this application (subsequently abbreviated as the RL-RNN-LSTM method), in this embodiment, by comparing with the estimation results of four existing methods, to judge the advantages and disadvantages of this method. The four existing methods are respectively the method using the recurrent neural network (subsequently abbreviated as the RNN method), the method using the long short-term memory network (subsequently abbreviated as the LSTM method), the method using the random weight combination model (subsequently abbreviated as the RRL-RNN-LSTM method), and the general form combination model method (subsequently abbreviated as the RL-R-LSTM method).
[0184] For the introduction of the recurrent neural network method, please refer to: "Zhou Heng, Wu Zhongyuan, Zhang Xin, et al. Shear wave prediction method based on LSTM recurrent neural network [J]. Fault-Block Oil & Gas Field, 2021, 28(06): 829-834."
[0185] For the introduction of the long short-term memory network method, please refer to: "Wang Qiuwen, Chen Yanru, Liu Yuanchun. Short-term passenger flow prediction of urban rail transit based on convolutional long short-term memory neural network [J]. Control and Decision, 2021, 36(11): 2760-2770."
[0186] For the introduction of the random weight combination model method, please refer to: "Lu Tianzhu, Qian Xiaochao, He Shu, et al. A time series prediction method based on deep learning [J]. Control and Decision, 2021, 36(3): 645-652."
[0187] For the introduction of the general form combination model method, please refer to: "Kong Wei, Liu Yun, Li Hui, et al. Review of pedestrian trajectory prediction methods based on deep learning [J]. Control and Decision, 2021, 36(12): 2841-2850."
[0188] The prediction error comparison of battery manufacturing capacity under different prediction methods is shown in Table 4 below
[0189] Table 4: Comparison of prediction errors of battery manufacturing capacity under different prediction methods
[0190]
[0191] To verify the accuracy and effectiveness of an online prediction method for battery manufacturing capacity proposed in this application, the following simulation experiments are carried out using the method of this application and existing RNN, LSTM, RRL-RNN-LSTM, and RL-R-LSTM methods. For the actual production process situation Figure 3 、 Figure 4 The changes in manufacturing capacity and error conditions of each method are respectively shown. During the experiment, the production data of the 22nd soft-pack battery workshop of Zhejiang Tianneng Battery Co., Ltd. is used for engineering verification. Taking the number of soft-pack batteries produced by each machine in the 22nd soft-pack battery workshop working 8 hours as the prediction output, the output of a single workshop machine for 200 days is collected as the data sample, and 75% of the previous data is selected as the training set sample, and 25% of the later data is used as the test set sample.
[0192] By Figure 3It can be seen that the circular line represents the true value of battery manufacturing capacity, the diamond dotted line represents the prediction result of the RNN method, the pentagram dotted line represents the prediction result of the LSTM method, the dot solid line represents the RL-RNN-LSTM method proposed in this application, the rectangular line represents the RRL-RNN-LSTM method, and the × line represents the RL-R-LSTM method. All these methods can roughly predict the overall change trend of manufacturing capacity.
[0193] It can be seen from Figure 4 that when the time is in the range of {0, 25}, the prediction result of the battery manufacturing capacity of the RNN method at the initial moment is closer to the true value than that of the LSTM method. However, it begins to gradually deviate from the true value around the 26th day. Although the compared RRL-RNN-LSTM method adds combined weights on the basis of a single model, due to the random setting of the weights, the prediction result is not stable and it is difficult to guarantee the prediction accuracy, indicating that not any combined prediction model with randomly set weights can improve the prediction accuracy of battery manufacturing capacity.
[0194] After combining reinforcement learning to find the optimal weights, it can be seen that the prediction accuracy has been significantly improved. Moreover, compared with the combined form of the RL-R-LSTM method, the prediction effect of the RL-RNN-LSTM method of this application is better, verifying the better effectiveness and accuracy of the RL-RNN-LSTM method proposed in this application. Table 4 shows the comparison of the prediction errors of the battery manufacturing capacity of each prediction method. It can be seen from Table 4 that overall, the prediction effect of LSTM is a little better than that of RNN in the single method. However, in terms of the MAD and RMSE error indicators, the prediction error of RRL-RNN-LSTM is actually worse than that of the LSTM method, which also reflects the importance of solving the weights in the combined method. In the comparison of the three error indicators, the prediction accuracy of the battery manufacturing capacity of the RL-R-LSTM method has increased by 34.8%, 16.7%, and 19.2% respectively compared with the RNN method, and has increased by 53.3%, 15.8%, and 24.9% respectively compared with the LSTM model. The prediction accuracy of the RL-RNN-LSTM method proposed in this application has increased by 5.8%, 20.4%, and 25.7% respectively compared with the RL-R-LSTM method. It can be seen that in the battery manufacturing process, using the RL-RNN-LSTM method proposed in this paper can further improve the prediction accuracy on the basis of the prediction effect of the RL-R-LSTM method. It shows that the online prediction method of battery manufacturing capacity proposed by the present invention has the characteristics of high prediction accuracy and strong reliability.
[0195] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as a CD or a hard disk, etc.
[0196] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An online prediction method for the manufacturing capacity of power batteries based on reinforcement learning, characterized in that, The method includes: Using m single prediction models to respectively predict the power battery manufacturing capacity within a future time period of length N, and obtaining the primary prediction results of each single prediction model; Based on the primary prediction results of each single prediction model, optimizing each single prediction model based on reinforcement learning; Using the optimized single prediction models to respectively re-predict the power battery manufacturing capacity within a future time period of length N, and obtaining the secondary prediction results of each single prediction model; Based on the secondary prediction results of each single prediction model, using reinforcement learning to segment the time period of length N to determine the optimal weights of each single prediction model, thereby determining the expression of the combined prediction model: Among them, w i is the weight of the i-th single prediction model, N i is the cut-off point of the i-th single prediction model, f ij represents the j-th predicted vector value of the i-th single prediction model, represents rounding down the numerical value a; w i Satisfy: Obtaining the predicted value of the power battery manufacturing capacity according to the expression of the combined prediction model determined by the optimal weights of each single prediction model; The single prediction models include a recurrent neural network model and a long short-term memory network model, and m = 2; the method includes: Step 1: Define that the production volume obtained per unit time in the battery manufacturing process represents the power battery manufacturing capacity; Step 2: Define the number of iterations of the recurrent neural network model \(k1 = 1\), the number of iterations of the long short-term memory network model \(k2 = 1\), the number of iterations of the weight learning \(k3 = 1\), and the corresponding iteration step sizes \(N1\), \(N2\), \(N3\). The Q1 table of the recurrent neural network model representing state-action in reinforcement learning, the Q2 table of the long short-term memory network model, and the Q3 table of weight learning are all 0. The optimal initial value \(l\) of the hidden layer of the recurrent neural network model 1,0 and the optimal initial value \(l\) of the hidden layer of the long short-term memory network model 2,0 , the action matrix \(A1\) of the recurrent neural network model, the action matrix \(A2\) of the long short-term memory network model, and the action matrix \(A3\) of weight learning, the weights \(w1\), \(w2\) of the combined prediction model, and calculate the first prediction results of the recurrent neural network model and the long short-term memory network model; Step 3: Use reinforcement learning to construct a hidden layer learning environment, establish a recurrent neural network loss function L1 and a long short-term memory network loss function L2, a recurrent neural network reward and punishment function R1 and a long short-term memory network reward and punishment function R2, and select action behaviors using the greedy algorithm according to the current state, Q1 table, and Q2 table; Step 4: Calculate the loss functions L1 and L2, the reward and punishment functions R1 and R2 according to the primary prediction results in Step 2, and update the Q1 table and Q2 table; Step 5: Let k1 = k1 + 1 and k2 = k2 + 1, return to Step 3 until k1 = N1 to output the optimal hidden layer l1 of the recurrent neural network; when k2 = N2, output the optimal hidden layer l2 of the long short-term memory network, and after outputting the two hidden layers, jump to Step 6; Step 6: Substitute the optimal hidden layer numbers l1 and l2, and recalculate to obtain the secondary prediction results of the recurrent neural network and the long short-term memory network; Step 7: Construct a weight learning environment, set the state matrix S1 and action matrix A3 of the power battery manufacturing capacity combined prediction model, and establish a weight learning loss function L3 and a reward and punishment function R3, and select action behaviors using the greedy algorithm according to the current state and Q3 table; Step 8: Calculate the loss function L3 and the reward and punishment function R3 according to the secondary prediction results in Step 6, and update the Q3 table; Step 9: Let k3 = k3 + 1, return to Step 7 until k3 = N3 to output the recurrent neural network weight w1 and the long short-term memory network weight w2, and jump to Step 10; Step 10: According to the output results in Step 9, construct a power battery manufacturing capacity combined prediction model based on reinforcement learning, input the historical data of the power battery manufacturing capacity into the constructed power battery manufacturing capacity combined prediction model based on reinforcement learning, and output the predicted value of the power battery manufacturing capacity.
2. The method according to claim 1, wherein In the primary prediction results of the recurrent neural network model and the long short-term memory network model calculated in Step 2, the output of the recurrent neural network model at time t is: where x t is the input of the system at time t, i.e., the historical data of the power battery manufacturing capacity, s t and o t are the outputs of the hidden layer and the output layer at time t. The output of the output layer is the predicted value of the power battery manufacturing capacity. U is the weight of the hidden layer, V is the weight of the output layer, and W represents the weight of the previous value of the hidden layer as the current input; g and f are activation functions; The output of the long short-term memory network model at time t is: Among them, h t represents the short-term historical information at time t, C t represents the long-term historical information at time t, is the candidate long-term historical information at time t, x t represents the input sample at time t, σ is the Sigmoid activation function, tanh is the hyperbolic tangent activation function, W f and b f are the weight matrix and bias vector of the forgetting gate respectively, f t is the output of the long short-term memory network forgetting gate at time t, W i and b i are the weight matrix and bias vector of the input gate respectively, i t is the output of the long short-term memory network input gate at time t, W o and b o are the weight matrix and bias vector of the output gate respectively, O t is the output of the long short-term memory network output gate at time t, that is, the predicted value of the power battery manufacturing capacity; Substitute the optimal initial value l of the hidden layer of the recurrent neural network 1,0 and the optimal initial value l of the hidden layer of the long short-term memory network 2,0 Calculate the first prediction results y1 and y2.
3. The method according to claim 2, characterized in that, Step 3 includes: Define variables l1 and l2 to represent the number of hidden layers of the recurrent neural network model and the number of hidden layers of the long short-term memory network model respectively; define action matrices A1 and A2: A1 = A2 = [Δv1 -Δv1] (7) where Δv1 is the amplitude of actions A1 and A2; construct the Q1 table of the recurrent neural network model and the Q2 table of the long short-term memory network model with the target hidden layer variables as row vectors and the action states as column vectors respectively, and set the loss functions L1 and L2 and the reward and punishment functions R1 and R2 as follows: Among them, y 1,t and y 2,t respectively represent the true values of the battery manufacturing capabilities of the recurrent neural network model and the long short-term memory network model at time t, and respectively represent the predicted values of the battery manufacturing capabilities of the recurrent neural network model and the long short-term memory network model at time t, L 1,t and L 2,t are respectively the loss values of the recurrent neural network model and the long short-term memory network model at time t, L 1,t+1 and L 1,t+1 are respectively the loss values of the recurrent neural network model and the long short-term memory network model at time t + 1, and N is the length of the output sample; The selection mechanism of A1 is Among them, represents the action corresponding to the maximum Q value in the Q1 table, A R1 represents randomly selecting an action from the action matrix A1, δ1 is a random number between [0, 1], and ε1 is the greed rate of the action A1; The selection mechanism of the action matrix A2 is: Among them, represents the action corresponding to the maximum Q value in the Q2 table, A R2 represents randomly selecting an action from the action matrix A2, δ2 is a random number between [0, 1], and ε2 is the greed rate of the action A2.
4. The method according to claim 3, characterized in that, The said step four includes: The formula for updating the Q1 table is as follows: Q 1,t+1 (l 1,t ,A 1,t ) = Q 1,t (l 1,t ,A 1,t ) + α1[R1(l 1,t ,A 1,t ) + γ1 maxQ 1,t (l 1,t+1 ,A 1,t+1 ) - Q 1,t (l 1,t ,A 1,t )] (14) Among them, Q 1,t (l 1,t , A 1,t ) represents the Q table constructed by using l1 and A1 at time t as the row and column respectively. 1,t The update learning rate of the Q1 table is α1, and the discount factor of the Q1 table is γ1; The formula for updating the Q2 table is as follows: Q 2,t+1 (l 2,t ,A 2,t ) = Q 2,t (l 2,t ,A 2,t ) + α2[R2(l 2,t ,A 2,t ) + γ2maxQ 2,t (l 2,t+1 ,A 2,t+1 ) - Q 2,t (l 1,t ,A 2,t )] (15) Among them, Q 2,t (l 2,t , A 2,t ) represents the Q table constructed by using l2 and A2 at time t as the row and column respectively. 2,t The update learning rate of the Q2 table is α2, and the discount factor of the Q2 table is γ2; the training iteration numbers are defined as N1 and N2 respectively, and the optimal hidden layers l1 and l2 are obtained through training.
5. The method according to claim 4, wherein The said step seven includes: The state matrix S1 and the action matrix A3 are respectively as follows: S1 = [w1 -w1] (16) A3 = [Δv2 -Δv2] (17) where w1 is the weight of the recurrent neural network model, w2 is the weight of the long short-term memory network model; Δv2 is the amplitude of the action A3; at the same time, establish a combined prediction model Q3 table with rows representing the target states and columns representing the action states, and set the loss function L3 and the reward and punishment function R3: Among them, y 3,t represents the true value of the battery manufacturing capacity of the combined prediction model at time t, respectively represent the predicted values of the battery manufacturing capacity of the combined prediction model at time t, L 3,t is the loss value of the combined prediction model at time t, L 3,t+1 is the loss value of the combined prediction model at time t + 1, and N is the length of the output sample; Set the selection mechanism of A3 as: Among them, A Q3max represents the action corresponding to the maximum Q value in the Q3 table. A r represents randomly selecting an action from the action matrix A3. δ3 is a random number between [0, 1], and ε3 is the greed rate of the action A3.
6. The method according to claim 5, characterized in that, The said step eight includes: Calculate the loss function L3 and the reward and punishment function R3 according to the secondary prediction result of step six and formulas (18) and (19); Update the Q3 table according to the following formula: Among them, Q 3,t (S 1,t , A 3,t ) represents the Q 3,t table constructed by using S1 and A3 at time t as the row and column respectively. α3 is the update learning rate of the Q3 table, and γ3 is the discount factor of the Q3 table.
7. The method according to claim 6, characterized in that, The said step nine includes: Define the training iteration times as N3 respectively; let k3 = k3 + 1, return to step seven until k3 = N3, output the recurrent neural network weight w1 and the long short-term memory network weight w2, and jump to step ten.
8. The method according to claim 7, wherein The said step ten includes: After obtaining the recurrent neural network weight w1 and the long short-term memory network weight w2, construct a combined prediction model of power battery manufacturing capacity based on reinforcement learning according to the combined prediction form of battery manufacturing capacity given by formulas (2) and (3), and output the predicted value of battery manufacturing capacity.
Citation Information
Patent Citations
Combined prediction method based on adaptive neural network
CN112001740A
Load prediction model training method and device, storage medium and equipment
CN112200373A