Approximate dynamic programming control method and system based on long short-term memory network

Through the approximate dynamic programming control method based on long and short-term memory networks, the problems of low control reliability and efficiency of traditional nonlinear systems are solved, and efficient and reliable control strategy generation and parameter optimization are realized, which is suitable for real-time optimization of complex systems.

CN120386206AInactive Publication Date: 2025-07-29NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510574352.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional optimal control scheme for nonlinear systems at continuous time has the problems of poor control reliability, low efficiency and insufficient scalability, especially when facing modern energy systems and smart grids, it is difficult to deal with high-dimensional, strong coupling characteristics and nonlinear response.

Method used

The approximate dynamic programming control method based on long and short-term memory network is adopted, and the long and short-term memory network is constructed, the initial parameter setting is performed, the initial state is uniformly sampled, the training data set is established, and the optimal value function and control strategy are obtained through iterative training. The gate mechanism and timing processing capabilities of LSTM are used to achieve efficient modeling and control of nonlinear systems.

Benefits of technology

It significantly improves the control reliability, efficiency and scalability of nonlinear systems, can handle multi-source heterogeneous data and real-time dynamic interactions in complex systems, reduces computing complexity, and achieves millisecond-level optimization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386206A_ABST
    Figure CN120386206A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of optimal control, and provides an approximate dynamic programming control method and system based on a long short-term memory network, and the method comprises the steps: constructing the long short-term memory network, and carrying out the initialization setting of network parameters and algorithm iteration parameters of the long short-term memory network, and obtaining initialization parameters; uniformly sampling a plurality of initial states from a state space of the controlled object, and determining trajectory data and a terminal state corresponding to each initial state according to the initialization parameters; determining a cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establishing a training data set according to all the initial states and the respective corresponding cost target values; and carrying out iterative training on the long-short-term memory network according to the training data set until a set iteration termination condition is met, and obtaining an optimal value function and an optimal control strategy. According to the scheme provided by the invention, the control reliability, the control efficiency and the expansibility of the optimal control link of the nonlinear system in continuous time are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimal control, and particularly to an approximate dynamic programming control method and system based on a long short-term memory network. Background Art

[0002] In modern control theory, the optimal control problem aims to find the global optimal strategy of the system performance index under dynamic constraints. Since the foundation of the Pontryagin maximum principle and the Bellman dynamic programming theory, the application scenarios have expanded from classical fields such as spacecraft orbit dynamics and industrial process control to frontier areas such as new energy systems and smart grids, promoting the development of complex system control technologies.

[0003] Facing the characteristics of high-dimensionality and strong coupling in modern energy systems, traditional control strategies are ineffective in complex tasks such as dynamic scheduling of smart grids and coordinated optimization of wind-solar-storage. Such systems need to process the dynamic interaction of multi-source heterogeneous data in real time and accurately compensate for random disturbances and non-linear responses at the millisecond scale. Especially when the penetration rate of renewable energy increases, the non-convex characteristics and uncertainties of the system dynamic equations pose a fundamental challenge to classical optimal control methods based on linearization approximation, which is rooted in the difficulty of solving the HJB (Hamilton-Jacobi-Bellman) partial differential equation. To break through the problem of solving the HJB equation, control strategies such as fuzzy logic control, neural network control, and adaptive dynamic programming have emerged. However, the static characteristics of the rule base in fuzzy logic control are difficult to adapt to unmodeled dynamics such as sudden load mutations or equipment failures. The network architecture design in neural network control highly depends on prior knowledge, and the completeness of the training data set directly affects the control robustness.

[0004] When the adaptive dynamic programming scheme faces a high-dimensional state space (such as a smart grid with distributed energy) or strong non-linear dynamics (such as the aerodynamic-mechanical coupling system of a wind turbine), there are problems such as a sharp increase in the number of iterations leading to super-linear computational complexity, and it is easy to fall into the local minimum trap in non-convex optimization. Moreover, in actual engineering deployment, its approximator design (such as activation function selection, network depth configuration), iteration step size tuning, and Lyapunov stability guarantee mechanism lack systematic design criteria to balance the convergence speed and numerical stability.

[0005] It is not difficult to find that traditional optimal control schemes for non-linear systems in continuous time have technical problems such as poor control reliability, low efficiency, and insufficient scalability. Summary of the Invention

[0006] The present invention provides an approximate dynamic programming control method and system based on a long short-term memory network, so as to solve the defects of poor control reliability, low efficiency and insufficient scalability in the traditional optimal control scheme for nonlinear systems in continuous time.

[0007] On the one hand, the present invention provides an approximate dynamic programming control method based on a long short-term memory network, including: Construct a long short-term memory network, and initialize the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; Uniformly sample a plurality of initial states from the state space of the controlled object, and determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters; Determine the cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establish a training data set according to all the initial states and their corresponding cost target values; Iteratively train the long short-term memory network according to the training data set until the set iteration termination condition is satisfied, and obtain the optimal value function and the optimal control strategy.

[0008] According to the approximate dynamic programming control method based on a long short-term memory network provided by the present invention, determining the trajectory data and terminal state corresponding to each initial state according to the initialization parameters includes: Determine the control strategy corresponding to each initial state according to the initialization parameters; Solve the preset closed-loop system equation according to the control strategy corresponding to each initial state to obtain the trajectory data and the terminal state.

[0009] According to the approximate dynamic programming control method based on a long short-term memory network provided by the present invention, the expression of the control strategy is: ; Wherein, represents the control strategy obtained based on the state i at the th iteration, R represents the control weight matrix, represents the transpose of the control gain matrix corresponding to the state , represents the weight parameter of the long short-term memory network for the gradient of the state .

[0010] According to the approximate dynamic programming control method based on a long short-term memory network provided by the present invention, the closed-loop system equation is: ; Wherein, Denote the state variable x with respect to time derivative of denote the nonlinear drift matrix corresponding to the moment denote the control gain matrix corresponding to the moment denote the i control strategy corresponding to the moment under the th iteration denote the initial moment corresponding to each initial state T denote the duration

[0011] According to the approximate dynamic programming control method based on long short-term memory network provided by the present invention, based on the trajectory data and the terminal state, determine the cost target value corresponding to each initial state, including: Based on the trajectory data, calculate the trajectory cost value corresponding to each initial state through numerical integration; Predict the value function at the terminal state to obtain the terminal cost value; Sum the trajectory cost value and the terminal cost value to obtain the cost target value corresponding to each initial state.

[0012] According to the approximate dynamic programming control method based on long short-term memory network provided by the present invention, the expression of the trajectory cost value is: ; wherein, denote the trajectory cost value corresponding to starting from the initial state and adopting the control strategy , denote the state at the moment in the trajectory data denote the initial moment j denote the count of discrete time steps denote the time interval denote the value of the control strategy at the moment Q denote the state weight matrix R denote the control weight matrix denote the sum of data on the discrete time steps from j 0 to N -1 for a total of N discrete time steps

[0013] According to the approximate dynamic programming control method based on the long short-term memory network provided by the present invention, the long short-term memory network is iteratively trained according to the training data set until a set iteration termination condition is satisfied, and an optimal value function and an optimal control strategy are obtained, including: In the current iteration round, any initial state in the training data set is input into the long short-term memory network to obtain a cost prediction value; According to the cost prediction value and the cost target value corresponding to any initial state, the gradient of the loss function with respect to the weight parameters of the long short-term memory network is calculated in the current round; According to the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round, the weight parameters of the long short-term memory network are updated to obtain updated weight parameters; On a preset independent verification set, according to the weight parameters before and after the update, the difference value of the value function before and after the update is calculated; If the difference value of the value function before and after the update is less than a set difference threshold, and / or the current number of iterations reaches the maximum number of iterations, it is determined that the set iteration termination condition is satisfied, and an optimal value function and an optimal control strategy are obtained.

[0014] According to the approximate dynamic programming control method based on the long short-term memory network provided by the present invention, the expression corresponding to the gradient of the loss function with respect to the weight parameters of the long short-term memory network is: ; Wherein, represents the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the i-th iteration, M represents the number of samples in the training data set, represents the cost prediction value predicted based on the weight parameters of the long short-term memory network in the i-th iteration for the state , represents the cost target value corresponding to the time represents the partial derivative of the long short-term memory network with respect to the weight parameters for the state .

[0015] According to the approximate dynamic programming control method based on the long short-term memory network provided by the present invention, according to the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round, the weight parameters of the long short-term memory network are updated to obtain updated weight parameters, including: Input the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round into a parameter optimizer to obtain a parameter update amount; Multiply the parameter update amount by the set learning rate to obtain the final update amount; Subtract the final update amount from the weight parameters of the long short-term memory network before update to obtain the updated weight parameters.

[0016] On the other hand, the present invention also provides an approximate dynamic programming control system based on a long short-term memory network, including: An initialization module, configured to construct a long short-term memory network, and perform initialization settings on the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; A sampling module, configured to uniformly sample a plurality of initial states from the state space of the controlled object, and determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters; A building module, configured to determine the cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establish a training data set according to all the initial states and their corresponding cost target values; An iteration module, configured to iteratively train the long short-term memory network according to the training data set until a set iteration termination condition is satisfied, to obtain an optimal value function and an optimal control strategy.

[0017] The approximate dynamic programming control method and system based on a long short-term memory network provided by the present invention construct a long short-term memory network, and perform initialization settings on the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; uniformly sample a plurality of initial states from the state space of the controlled object, and determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters; determine the cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establish a training data set according to all the initial states and their corresponding cost target values; iteratively train the long short-term memory network according to the training data set until a set iteration termination condition is satisfied, to obtain an optimal value function and an optimal control strategy. Due to the introduction of the long short-term memory network, the nonlinear approximation ability and the time series modeling performance of the adaptive dynamic programming link are significantly improved, a complete closed loop from value function approximation, control strategy generation to parameter optimization is realized, and at the same time, the calculation amount is greatly reduced, and the control reliability, control efficiency and scalability of the optimal control link of the nonlinear system in continuous time are improved. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of the approximate dynamic programming control method based on the long short-term memory network provided by an embodiment of the present invention; Figure 2 It is a schematic structural diagram of the approximate dynamic programming control system based on the long short-term memory network provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] This embodiment relates to the field of optimal control and can be specifically applied to the optimal control scenario of a non-linear system in continuous time. For example, in the autonomous driving scenario, based on a non-linear vehicle dynamics model in continuous time, the control strategy of the power system is optimized to minimize fuel consumption or power consumption on the premise of ensuring driving safety and comfort.

[0022] Considering that the state changes of non-linear dynamic systems often have long-term dependencies in time. However, due to structural limitations, traditional feedforward neural networks often have difficulty capturing the correlations between distant states. To address this defect, LSTM (Long Short Term Memory) effectively manages memory cells through its unique gating mechanism (i.e., input gate, forget gate, output gate), and can not only selectively retain key historical information, but also more accurately model the long-term dependencies in the state sequence.

[0023] Considering that the adaptive dynamic programming link often needs to process the integration or prediction of state trajectories (such as the Bellman equation in value function iteration), the time series processing ability of LSTM naturally adapts to such requirements. It can directly model the state transition process in continuous time or discrete time, thus avoiding the cumbersome manual discretization or model simplification steps in traditional solutions. Specifically, in a continuous-time system, by dividing the state trajectory into time windows and inputting them into LSTM, the system can automatically learn the time series characteristics of state evolution, which significantly reduces the computational burden of explicit integration of dynamic equations in traditional solutions.

[0024] Considering that traditional recurrent neural networks are prone to the problem of gradient vanishing or explosion during training, which makes it difficult for deep time steps to converge. However, with the collaborative design of gating units and cell state, LSTM enables stable propagation of gradients in the time dimension. This advantage is particularly suitable for scenarios that require multiple iterative updates in adaptive dynamic programming (such as the policy evaluation - improvement loop), ensuring the robustness of the training process.

[0025] More importantly, in the face of high - dimensional or implicit features hidden in non - linear state equations (such as the phase - space reconstruction of chaotic systems), LSTM can automatically extract features related to optimal control from the original state data, thus reducing the dependence on artificially designed basis functions (such as polynomials and radial basis functions). Taking the control of complex fluids as an example, its state space contains multi - scale features of turbulence, and LSTM can automatically identify key dynamic patterns through a hierarchical gating mechanism, significantly improving the generalization performance of control strategies.

[0026] In the face of the problems of partially unobservable states or sensor noise commonly existing in practical systems, the memory function of LSTM shows unique value: by integrating historical observation data, its hidden state can effectively estimate missing information, thus enhancing the robustness of the adaptive dynamic programming link in non - ideal environments. Furthermore, the differentiability property of LSTM allows it to be jointly trained with the policy network to construct an Actor - Critic framework, realizing the synchronous optimization of the value function and the policy. This end - to - end learning method directly calculates the policy gradient through backpropagation, significantly improving the overall efficiency of the adaptive dynamic programming algorithm. For example, in an energy management system, jointly training LSTM Critic (evaluating the value function) and policy Actor (generating control instructions) can avoid the delay of alternating iterations in traditional adaptive dynamic programming links and achieve millisecond - level real - time optimization.

[0027] Finally, in response to the challenges of high - dimensional non - linear systems (such as multi - agent cooperative control), LSTM can significantly reduce the computational complexity through parameter sharing and sequence compression mechanisms. Compared with methods such as support vector regression, its adaptability in large - scale problems is more prominent, effectively solving the engineering implementation problem caused by dimensional explosion in traditional methods.

[0028] Based on the above advantages of the long short - term memory network in the adaptive dynamic programming link, this embodiment proposes an approximate dynamic programming control method based on the long short - term memory network for solving the optimal control problem of non - linear systems in continuous time. When processing this kind of time - series data, LSTM can better model the dynamic characteristics of the system, especially the estimation of the value function over a long time span.

[0029] Next, the specific feasible theory on which the approximate dynamic programming control method based on the long short - term memory network proposed in this embodiment is based will be described.

[0030] Traditional value iteration may require discretization of time steps, while LSTM can naturally handle continuous-time data or longer sequence inputs. Taking a class of nonlinear dynamic systems with time-invariant affine structure as an example, its state-space description can be expressed as: (1); where, represents the derivative of the system state variable x with respect to time t; represents the state variable of the system at time t, such as the position or velocity of a mechanical system, , that is, the state variable is an n-dimensional vector; represents the nonlinear drift matrix in the state , represents the control gain matrix in the state . In this embodiment, both the nonlinear drift matrix and the control gain matrix are Lipschitz continuous on , ensuring the uniqueness and stability of the system trajectory and satisfying controllability, that is, there exists at least one controllable control strategy such that the closed-loop system is asymptotically stable within ; represents the control input at time t, that is, the externally applied control signal, such as motor torque, , that is to say, the control input is an m-dimensional vector.

[0031] Subsequently, this embodiment designs a control strategy such that starting from any initial state, the system minimizes the performance index in the infinite time domain as: (2); where, represents the performance index value, represents the state variable of the system at the moment . Q represents a positive definite state weight matrix, which is used to punish the degree of state deviation from the equilibrium point, such as tracking error; represents the control input at the moment, R represents a positive definite control weight matrix, which is used to punish the consumption of control energy, such as actuator output; x0 represents the initial state of the system at the initial moment t = 0.

[0032] To deeply explore the time-series modeling advantages of LSTM, this embodiment realizes the extraction of high-dimensional spatio-temporal features of nonlinear dynamic systems through its gated memory mechanism, and can at least solve the following technical problems: First, overcome the interpretability bottleneck of the black-box model and establish a convergence guarantee system based on Lyapunov stability constraints; Second, design an online incremental learning architecture with hardware awareness to improve the parameter update efficiency in a dynamic environment; Third, construct a network training paradigm guided by control theory.

[0033] At the same time, to solve the optimal control problem, the stability of the system needs to be satisfied: the closed-loop system satisfies asymptotic stability, that is , and the optimality of the control strategy is satisfied: the control strategy globally minimizes the performance index.

[0034] According to the dynamic programming theory, the optimal value function needs to satisfy the HJB equation, and the control input needs to minimize the Hamiltonian, that is, it satisfies: (3); Among them, x represents the state variable of the system, which can characterize the state information of the system at a certain moment; u represents the control input of the system at a certain moment, and satisfies this value range; Q represents the state positive definite weight matrix; R represents the control positive definite weight matrix; represents the optimal value function the gradient of the state variable x; represents the nonlinear drift matrix in state x; represents the control gain matrix in state x.

[0035] Further, an explicit optimal control strategy can be obtained by solving the HJB equation: (4); Among them, represents the optimal control strategy, that is, the control input that can make the system reach the optimal in state x.

[0036] Since the HJB equation is a partial differential equation and it is difficult to solve the nonlinear system analytically, therefore, in this embodiment, the optimal value function and the optimal control strategy need to be approximated iteratively.

[0037] Design a function approximator for the value iteration algorithm of the continuous-time nonlinear system to estimate the scalar value function. Since LSTM is good at processing sequence data and can model long-term dependencies through the gating mechanism, therefore, in this embodiment, the LSTM network is used to process the time series data. However, in practical applications, the value function only explicitly depends on the current state, so the state at each time point is regarded as an independent sample.

[0038] At the same time, the defined input quantity is the n-dimensional state vector at the current moment, and the output quantity is the estimated value of the scalar value function, that is, the cost prediction value. Since the value function only depends on the current state and LSTM is suitable for sequence modeling. Therefore, in this embodiment, this contradiction is solved through structural design, and the network structure is designed to approximate the value function as: (5); Among them, denotes a scalar value function, and \(x\) represents the state vector at the current moment. denotes the weight parameters of the long short-term memory network in the \(i\)-th iteration.

[0039] The hidden state of the LSTM captures the long-term dependencies of the state sequence through time unfolding, and its dynamic equations are: (6); (7); (8); (9); (10); (11); where denotes the forget gate; denotes the input gate; denotes the output gate; denotes the candidate memory; denotes the memory cell state; denotes the hidden state at time step \(t\), and the initial hidden state ; denotes the Sigmoid function; denotes element-wise multiplication.

[0040] The final value function is mapped from the hidden state through a fully connected layer: (12); where denotes the final hidden state, and both denote the output layer weight parameters.

[0041] The optimal value function The gradient of the state variable \(x\) is calculated through automatic differentiation as follows: (13); At the same time, the rolling optimization cost is reconstructed as follows: (14); where denotes the trajectory cost of the finite time domain; denotes the terminal cost predicted by the LSTM.

[0042] Subsequently, \(M\) initial states are uniformly sampled from the state space , for each initial state, applying the control strategy to generate a trajectory through the Runge-Kutta method can obtain trajectory data and the terminal state, and then obtain the cost objective value. Furthermore, using the time-expanded LSTM, the weight parameters are optimized through backpropagation and gradient descent. Subsequently, weight optimization is performed, and the weight parameters of the LSTM are updated by the gradient descent method until the iteration termination condition is satisfied, and then the optimal value function and the optimal control strategy can be obtained.

[0043] The following combines Figures 1 to 3 to describe the detailed solution of the approximate dynamic programming control method based on the long short-term memory network provided by the embodiments of the present invention.

[0044] Figure 1 is the flow schematic diagram of the approximate dynamic programming control method based on the long short-term memory network provided by the embodiments of the present invention.

[0045] As Figure 1 shown, for the approximate dynamic programming control method based on the long short-term memory network provided by the embodiments of the present invention, the execution subject can be a computer or a server with data processing capabilities. The above method mainly includes the following steps: Step 110: Construct a long short-term memory network, and initialize the network parameters and algorithm iteration parameters of the long short-term memory network to obtain the initialization parameters.

[0046] In this embodiment, in the initialization link of the network parameters, the dimension of the input layer of the long short-term memory network is set to the system state dimension n, and the dimension of the hidden layer d n is set to 2n to 4n, and the output layer is a scalar value function estimation. Set the initial weight of the long short-term memory network to , the iteration number i = 0, and the initial control strategy is , and the initial value function is defined as follows: (15); where x represents the state vector at the beginning of the iteration, represents the initial weight.

[0047] In the algorithm iteration parameter setting link, the learning rate, regularization system, maximum iteration number, convergence threshold, and sampling parameters including the number of sampling points, trajectory integration duration, and discrete time step are mainly set.

[0048] Step 120: Uniformly sample multiple initial states from the state space of the controlled object, and determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters.

[0049] In this embodiment, the trajectory data is mainly used to describe the set of states corresponding to each step after a set number of steps are taken according to the set discrete time step starting from the moment corresponding to the current initial state. The terminal state is mainly used to describe the final state of the controlled object after a set duration starting from the moment corresponding to the current initial state.

[0050] Step 130: Determine the cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establish a training data set based on all the initial states and their corresponding cost target values.

[0051] In this embodiment, each initial state and the corresponding cost target value are mainly used as a sample array, and the data set composed of multiple sample arrays is used as the training data set.

[0052] Step 140: Iteratively train the long short-term memory network according to the training data set until the set iteration termination condition is met, and obtain the optimal value function and the optimal control strategy.

[0053] The solution provided in this embodiment uses the long short-term memory network for adaptive dynamic programming, realizes a complete closed-loop from value function approximation, control strategy generation to parameter optimization. By introducing the long short-term memory network, the nonlinear approximation ability and the temporal modeling performance are significantly improved. At the same time, through automatic differentiation and gradient descent optimization, the computational complexity is greatly reduced, providing an efficient and scalable solution for the real-time optimal control of complex dynamic systems.

[0054] In one embodiment, according to the initialization parameters, determine the trajectory data and the terminal state corresponding to each initial state, specifically including: First, determine the control strategy corresponding to each initial state according to the initialization parameters.

[0055] In one embodiment, the expression of the control strategy is: (16); Where represents the control strategy obtained based on the state at the i-th iteration, R represents the control weight matrix, represents the transpose of the control gain matrix corresponding to the state , represents the weight parameter of the long short-term memory network for the gradient of the state

[0056] Then, solve the preset closed-loop system equation according to the control strategy corresponding to each initial state to obtain the trajectory data and the terminal state .

[0057] In this embodiment, the closed-loop system equation is: (17); Wherein, represents the derivative of the state variable x with respect to time , represents the nonlinear drift matrix corresponding to the time represents the control gain matrix corresponding to the time represents the control strategy corresponding to the time at the i-th iteration, , represents the initial time corresponding to each initial state, and T represents the duration.

[0058] In one embodiment, according to the trajectory data and the terminal state, the cost target value corresponding to each initial state is determined, specifically including: On the one hand, according to the trajectory data, the trajectory cost value corresponding to each initial state is calculated by numerical integration.

[0059] In this embodiment, the expression of the trajectory cost value is: (18); Wherein, represents the trajectory cost value corresponding to starting from the initial state and adopting the control strategy , represents the state at the time in the trajectory data, represents the initial time, j represents the count of the discrete time step, represents the time interval, represents the value of the control strategy at the time, Q represents the state weight matrix, and R represents the control weight matrix, represents the sum of the data for j from 0 to N - 1 over these N discrete time steps.

[0060] On the other hand, the value function at the terminal state is predicted to obtain the terminal cost value.

[0061] In this embodiment, the current LSTM can be used to predict the value function of the terminal state to obtain the terminal cost value, specifically as follows: (19); Wherein, represents the value function at the terminal state , that is, the terminal cost value, Indicates the output of the long short-term memory network under the weight parameters and the terminal state The output quantity under the condition.

[0062] Finally, the trajectory cost value and the terminal cost value are summed to obtain the cost target value corresponding to each initial state.

[0063] In this embodiment, the cost target value can be expressed as follows: (20); Wherein, Indicates the cost target value, Indicates starting from the initial state The trajectory cost value corresponding to the control strategy When used, Indicates the terminal state The terminal cost value under the condition.

[0064] In one embodiment, the long short-term memory network is iteratively trained according to the training data set until the set iteration termination condition is met, and the optimal value function and the optimal control strategy are obtained, specifically including: In the first step, at the current iteration round, any initial state in the training data set is input into the long short-term memory network to obtain a cost prediction value.

[0065] In this embodiment, the cost prediction value can be expressed as follows: (21); Wherein, Indicates the state The corresponding cost prediction value, Indicates the output of the long short-term memory network under the weight parameters And the state The output quantity under the condition.

[0066] In the second step, according to the cost prediction value and the cost target value corresponding to any initial state, the gradient of the loss function with respect to the weight parameters of the long short-term memory network is calculated in the current round.

[0067] In this embodiment, the mean squared error loss function is first defined and the L2 regularization term is added. The expression of the loss function is as follows: (22); Wherein, Indicates the loss function value of the long short-term memory network under the weight parameters Under the condition, Indicates the L2 regularization term, M represents the number of samples in the training data set, that is, the number of initial states, Indicates based on the weight parameters of the long short-term memory network in the i-th iteration For the state The predicted cost prediction value represents The cost target value corresponding to the moment.

[0068] Specifically, the expression corresponding to the gradient of the loss function with respect to the weight parameters of the long short-term memory network is: (23); where represents the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the i-th iteration, M represents the number of samples in the training dataset, represents the predicted cost prediction value for the state predicted based on the weight parameters of the long short-term memory network in the i-th iteration represents The cost target value corresponding to the moment, represents the partial derivative of the long short-term memory network with respect to the weight parameters for the state .

[0069] In the third step, according to the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round, update the weight parameters of the long short-term memory network to obtain the updated weight parameters.

[0070] In this embodiment, according to the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round, update the weight parameters of the long short-term memory network to obtain the updated weight parameters, which specifically includes: First, input the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round into the parameter optimizer to obtain the parameter update amount.

[0071] Then, multiply the parameter update amount by the set learning rate to obtain the final update amount.

[0072] Finally, subtract the final update amount from the weight parameters of the long short-term memory network before the update to obtain the updated weight parameters.

[0073] It should be noted that in this embodiment, the Adam optimizer is used to update the weight parameters, and the updated weight parameters can be expressed as follows: (24); where represents the updated weight parameters, represents the weight parameters before the update, represents the learning rate, represents the parameter update amount output by the Adam optimizer.

[0074] In the fourth step, on a preset independent validation set, according to the weight parameters before and after the update, the value function difference values before and after the update are calculated.

[0075] In this embodiment, the value function difference values before and after the update can be calculated as follows: (25); Wherein, represents the value function difference value before and after the update, represents the value function after the update, represents the value function before the update, represents the independent validation set.

[0076] In the fifth step, if the value function difference value before and after the update is less than the set difference threshold, and / or the current iteration number reaches the maximum iteration number, it is determined that the set iteration termination condition is satisfied, and the optimal value function and the optimal control strategy .

[0077] It can be understood that if the iteration termination condition is not satisfied at this time, return to the iteration policy evaluation and control strategy generation link, continue to determine the trajectory data and the terminal state, so as to perform the next round of iterative training until the iteration termination condition is satisfied, and then the training ends.

[0078] Based on the same general inventive concept, the present invention also protects an approximate dynamic programming control system based on a long short-term memory network. The approximate dynamic programming control system based on a long short-term memory network provided by the present invention will be described below. The approximate dynamic programming control system based on a long short-term memory network described below can be mutually corresponding and referred to the approximate dynamic programming control method described above.

[0079] As Figure 2 shown, the approximate dynamic programming control system based on a long short-term memory network provided by the embodiment of the present invention specifically includes: An initialization module 210, configured to construct a long short-term memory network, and perform initialization settings on the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters.

[0080] A sampling module 220, configured to uniformly sample a plurality of initial states from the state space of the controlled object, and determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters.

[0081] A building module 230, configured to determine the cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establish a training data set according to all the initial states and their respective corresponding cost target values.

[0082] An iterative module 240 is configured to iteratively train a long short-term memory network based on a training data set until a set iterative termination condition is met, so as to obtain an optimal value function and an optimal control strategy.

[0083] Regarding the system in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0084] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention.

[0085] As Figure 3 shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340. Among them, the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communication bus 340. The processor 310 may call logic instructions in the memory 330 to execute an approximate dynamic programming control method based on a long short-term memory network, including: constructing a long short-term memory network, and initializing network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; uniformly sampling a plurality of initial states from the state space of the controlled object, and determining trajectory data and terminal states corresponding to each initial state according to the initialization parameters; determining a cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establishing a training data set according to all the initial states and their corresponding cost target values; iteratively training the long short-term memory network based on the training data set until a set iterative termination condition is met, so as to obtain an optimal value function and an optimal control strategy.

[0086] In addition, when the logic instructions in the above-mentioned memory 330 are implemented in the form of software function units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0087] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute an approximate dynamic programming control method based on a long short-term memory network, including: constructing a long short-term memory network, and initializing the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; uniformly sampling a plurality of initial states from the state space of the controlled object, and determining the trajectory data and terminal state corresponding to each initial state according to the initialization parameters; determining the cost target value corresponding to each initial state according to the trajectory data and terminal state, and establishing a training data set according to all the initial states and their corresponding cost target values; iteratively training the long short-term memory network according to the training data set until a set iteration termination condition is satisfied, to obtain an optimal value function and an optimal control strategy.

[0088] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an approximate dynamic programming control method based on a long short-term memory network, including: constructing a long short-term memory network, and initializing the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; uniformly sampling a plurality of initial states from the state space of the controlled object, and determining the trajectory data and terminal state corresponding to each initial state according to the initialization parameters; determining the cost target value corresponding to each initial state according to the trajectory data and terminal state, and establishing a training data set according to all the initial states and their corresponding cost target values; iteratively training the long short-term memory network according to the training data set until a set iteration termination condition is satisfied, to obtain an optimal value function and an optimal control strategy.

[0089] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An approximate dynamic programming control method based on a long short-term memory network, characterized in that Including: Construct a long short-term memory network, and initialize the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; Uniformly sample multiple initial states from the state space of the controlled object, and determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters; Determine the cost objective value corresponding to each initial state according to the trajectory data and the terminal state, and establish a training data set according to all initial states and their corresponding cost objective values; Iteratively train the long short-term memory network according to the training data set until the set iteration termination condition is met, and obtain the optimal value function and optimal control strategy.

2. The approximate dynamic programming control method based on a long short-term memory network according to claim 1, wherein Determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters, including: Determine the control strategy corresponding to each initial state according to the initialization parameters; Solve the preset closed-loop system equation according to the control strategy corresponding to each initial state to obtain the trajectory data and terminal state.

3. The approximate dynamic programming control method based on a long short-term memory network according to claim 2, wherein The expression of the control strategy is: ; Among them, represents the control strategy obtained based on the state i at the -th iteration, R represents the control weight matrix, represents the transpose of the control gain matrix corresponding to the state , represents the weight parameter of the long short-term memory network for the gradient of the state .

4. The approximate dynamic programming control method based on a long short-term memory network according to claim 2, characterized in that The closed-loop system equation is: ; Among them, represents the state variable x derivative with respect to time , represents the non - linear drift matrix corresponding to the moment represents the control gain matrix corresponding to the moment represents the i th iteration control strategy corresponding to the moment , represents the initial moment corresponding to each initial state, T represents the duration.

5. The approximate dynamic programming control method based on a long short-term memory network according to claim 1, characterized in that Determine the cost objective value corresponding to each initial state according to the trajectory data and the terminal state, including: Calculate the trajectory cost value corresponding to each initial state by numerical integration according to the trajectory data; Predict the value function at the terminal state to obtain the terminal cost value; Sum the trajectory cost value and the terminal cost value to obtain the cost objective value corresponding to each initial state.

6. The approximate dynamic programming control method based on long short-term memory network according to claim 5, characterized in that The expression of the trajectory cost value is: ; Among them, represents the trajectory cost value corresponding to starting from the initial state and adopting the control strategy . represents the state at the moment in the trajectory data, represents the initial moment, j represents the count of discrete time steps, represents the time interval, represents the value of the control strategy at the Q moment, R represents the state weight matrix, represents the control weight matrix, j represents the sum of data from 0 to N -1 for these N discrete time steps.

7. The approximate dynamic programming control method based on a long short-term memory network according to claim 4, wherein Iteratively train the long short-term memory network according to the training data set until the set iteration termination condition is met, and obtain the optimal value function and optimal control strategy, including: In the current iteration round, input any initial state in the training data set into the long short-term memory network to obtain a cost prediction value; Calculate the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round according to the cost prediction value and the cost objective value corresponding to any initial state; Update the weight parameters of the long short-term memory network according to the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round to obtain the updated weight parameters; Calculate the difference value of the value function before and after the update on a preset independent validation set according to the weight parameters before and after the update; If the difference value of the value function before and after the update is less than the set difference threshold, and / or the current iteration number reaches the maximum iteration number, it is determined that the set iteration termination condition is met, and the optimal value function and optimal control strategy are obtained.

8. The approximate dynamic programming control method based on a long short-term memory network according to claim 7, wherein The corresponding expression of the gradient of the loss function with respect to the weight parameters of the long short-term memory network is: ; Among them, represents the loss function in the i-th iteration with respect to the weight parameters of the long short-term memory network, M represents the number of samples in the training dataset, represents the cost prediction value predicted for the state based on the weight parameters of the long short-term memory network in the i-th iteration , represents the cost target value corresponding to the time represents the partial derivative of the long short-term memory network with respect to the weight parameter for the state .​ 9. The approximate dynamic programming control method based on a long short-term memory network according to claim 7, wherein Update the weight parameters of the long short-term memory network according to the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round to obtain the updated weight parameters, including: Input the gradient of the loss function with respect to the weight parameters of the long short-term memory network in the current round into the parameter optimizer to obtain a parameter update amount; Multiply the parameter update amount by the set learning rate to obtain the final update amount; Subtract the weight parameters of the long short-term memory network before update from the final update amount to obtain the updated weight parameters.

10. An approximate dynamic programming control system based on a long short-term memory network, characterized in that, It includes: An initialization module, configured to construct a long short-term memory network, and perform initialization settings on the network parameters and algorithm iteration parameters of the long short-term memory network to obtain initialization parameters; A sampling module, configured to uniformly sample a plurality of initial states from the state space of the controlled object, and determine the trajectory data and terminal state corresponding to each initial state according to the initialization parameters; A building module, configured to determine the cost target value corresponding to each initial state according to the trajectory data and the terminal state, and establish a training data set according to all the initial states and their corresponding cost target values; An iteration module, configured to perform iterative training on the long short-term memory network according to the training data set until a set iteration termination condition is met, to obtain an optimal value function and an optimal control strategy.

Citation Information

Patent Citations

  • Multi-agent following control method and system based on adaptive dynamic programming

    CN115755615A

  • Database query optimization method and system based on self-attention and reinforcement learning

    CN117827889A

  • Learning type predictive control method for pilotless automobile based on long and short-term memory network

    CN118025223A

Cited By

  • Northern greenhouse tropical fruit microclimate adaptive control system based on reinforcement learning

    CN121325625A