Multi-stage Q-learning optimal control method and device based on system decoupling and medium
Through the multi-stage Q-learning method of system decoupling, the control strategy is optimized, and the optimal control problem of high-dimensional discrete time systems is solved, efficient and real-time control decisions are achieved, and it is suitable for fields such as automated production and distributed energy management.
Patent Information
- Application Number
- CN202510826565.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The prior art solves the optimal control problem in high-dimensional discrete time systems with low efficiency and high difficulty, and traditional methods are prone to falling into local optimality.
A multi-stage Q-learning optimal control method based on system decoupling is adopted, multiple value functions are optimized through a non-cooperative game framework, feedback control strategies are designed, and the approximate function is iterated by the least squares method to realize system decoupling and policy updates.
Effectively reduce computing complexity, improve real-time and robustness of control decisions, and reduce computing resource consumption. It is suitable for complex control systems such as automated production, intelligent scheduling and distributed energy management.
Smart Images

Figure CN120507989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of optimal control technology, and in particular to a multi-stage Q-learning optimal control method, device and storage medium based on system decoupling. Background Art
[0002] Optimal control is a crucial component of modern control theory and plays a key role in areas such as drone control, traffic planning, and industrial process automation. The core concept of optimal control is to optimize the performance of a system by developing an optimal control strategy. Control strategies derived from this concept can reduce control costs and improve control system efficiency.
[0003] Adaptive dynamic programming (ADP) is an optimal control method integrated with the reinforcement learning (RL) theoretical framework. Compared to traditional methods, it can solve more complex optimal control problems. For example, the ADP-based Q-learning algorithm can solve the algebraic Riccati equations in optimal control problems online without requiring knowledge of the control system. Summary of the Invention
[0004] The present invention proposes a multi-stage Q-learning optimal control method, device and storage medium based on system decoupling, which aims to solve the current problem of low efficiency and high difficulty in solving a type of optimal control problem based on high-dimensional discrete-time systems (DT).
[0005] To achieve the above object, the present invention adopts the following technical solutions: Step S1 specifically includes: Assume that there exists the following linear DT system:
[0006] Where, for Maintain system status, for Dimensional control strategy, is an unknown disturbance, , , ; For the optimal control problem of the above system, assume that both the control strategy and the disturbance strategy are feedback stable; under the condition that the optimal disturbance is obtained, that is, the value function given below is maximized, design a set of feedback control strategies , which minimizes the value function:
[0007] in, , for the system in The utility function at time , , Respectively The weights of the value function on the state, control strategy and disturbance strategy, Represents the system Gain, and Respectively The control and perturbation strategies under the value function are written in vector form as follows:
[0008]
[0009] The control strategy and disturbance strategy designed for this set of value functions are expressed as:
[0010]
[0011] because In the form of linear quadratic form, then the value function is further expressed as:
[0012] here is the solution of the following game Riccati equation GARE:
[0013] Let the optimal strategy , ,in and are the control and disturbance gains respectively; if the matrix It can be solved, then The control and disturbance gains of the value function are expressed as:
[0014]
[0015] Define a set of Q functions as: in, is a matrix that stores Q function information, and:
[0016] At this time, the Q function and the value function satisfy the following relationship:
[0017] therefore, and satisfy:
[0018] No. The control and disturbance gains of the value function are transformed into:
[0019] .
[0020] S2. Decouple the control system. The specific implementation steps are as follows: The conditions and methods for system decoupling are as follows: If the control system satisfies: is the identity matrix, , the weight of the value function satisfies , , ,and , , , and The dimensions are the same. Then there must be ,in, . In addition, the control and policy gains satisfy: ,and , , and The dimensions are the same, , , and The dimensions are the same. At this time, and It can be expressed as:
[0021]
[0022] The proof of the decoupling conditions and methods is as follows: Define the matrix ,in , , , Then, we can further define The inverse matrix of ,in: , , , .in, is a matrix Shure patch.
[0023] Easy to find, , , , are all diagonal block matrices, and their dimensions are the same as the matrix are the same. Therefore, , , and can be written as: , , , , where their dimensions are consistent.
[0024] therefore, can be rewritten as:
[0025] make ,in, and The dimensions are the same, and The dimensions are the same. is the identity matrix, then the Riccati equation can be rewritten as:
[0026] in:
[0027]
[0028] Then, the Riccati equation can be further written as:
[0029] Further deducing the above formula, we can get:
[0030]
[0031]
[0032] According to the above expression, we can see that if and only if and are all 0 matrices, that is When and The matrix is always 0, that is, Therefore, there must be in .
[0033] Based on the above derivation process, the strategy gain matrix can be further written as:
[0034]
[0035] It is easy to see that and is also a diagonal block matrix, so the gain matrix must also be a diagonal block matrix, that is, , ,in:
[0036]
[0037]
[0038] The proof is complete.
[0039] If the weights of the control system and the value function meet the decoupling conditions, the system can be decoupled into two subsystems. Furthermore, if the subsystems and their corresponding weights still meet the decoupling conditions, the subsystems can continue to be decoupled until all subsystems no longer meet the decoupling conditions.
[0040] Furthermore, after decoupling the control system, the control system The input and perturbation strategy of the value function satisfy the following expression:
[0041]
[0042] in ( ) indicates the The status of each subsystem, and Indicates the Control and disturbance strategies for each subsystem, and The corresponding gain.
[0043] Initialize system number .
[0044] S3, Initialization value function number: .
[0045] S4. Initialize the gain number of the control strategy: .
[0046] S5. Set and initialize the structural parameters. The specific steps are as follows: The present invention uses the least squares method to iteratively approximate the value function, thereby obtaining the optimal control strategy. First, initialize the Q function , here It is non-optimal. Next, the iterative equation of the Q function is given as follows: therefore, and The following relationship exists:
[0047] Further:
[0048] Based on the above derivation process of iterative Q function, we can know that as long as we determine
[0049] The value of , then we can get Therefore, choose a suitable iterative algorithm and let When , then there will be , so that , .
[0050] The parameter structure of the Q function is as follows:
[0051] in:
[0052]
[0053] here, Corresponding matrix elements.
[0054] Furthermore, the two action parameter structures are designed as follows:
[0055] Initialize iteration number .
[0056] S6, strategy evaluation and strategy update. The specific implementation is as follows: In order to obtain , define the following objective function:
[0057] If the optimal strategy is After iterations, it must satisfy:
[0058] make:
[0059] in, and , so we have:
[0060] According to the form of the above formula, it can be solved by the least squares method In order to make the least squares problem solvable, detection noise needs to be introduced, so the parameter structure of the strategy becomes:
[0061]
[0062] in, , further, the objective function is rewritten as: here, .
[0063] Furthermore, It can be updated by the following least squares solution:
[0064] So, for each subsystem, there is:
[0065] Furthermore, Can be achieved through In addition, for each subsystem, the detection noise needs to be selected , , and appropriate parameters , .
[0066] In a given Under the condition of , the updated strategy gain is as follows:
[0067]
[0068] Furthermore, the updated policy is as follows:
[0069]
[0070] S7, convergence check. The specific implementation is as follows: when The optimal strategy is obtained when . However, it is impractical to perform an infinite number of calculations. Therefore, we can check The convergence of is used to stop the policy iteration: like , output the optimal strategy:
[0071] make , And transfer to step 8, otherwise let And return to step 6.
[0072] S8, if , go to step 5; otherwise, let , go to step 9.
[0073] S9, if , go to step 4; otherwise, end.
[0074] Furthermore, a computer-readable storage device stores a computer program, wherein when the computer program is executed, a multi-stage Q-learning optimal control method based on system decoupling is implemented.
[0075] In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0076] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0077] Compared with the prior art, the present invention has the following technical effects: This method assigns different weights to different control objectives and optimizes multiple value functions using a non-cooperative game framework, enabling the control strategy to dynamically balance multiple constraints. This approach not only avoids the local optimum often encountered by traditional single-value-function approaches, but also allows for flexible strategy adjustment to better meet the overall system optimization requirements. Compared to traditional deep Q-learning (DQN) methods, this approach, through non-cooperative game modeling and dynamic decoupled computation, effectively reduces computational complexity, improves the real-time and robustness of control decisions, and significantly reduces computational resource consumption while ensuring optimal control strategies. Because this method relies on data-driven decision-making, it can continuously optimize control strategies through reinforcement learning even in the absence of precise mathematical models. This approach has broad application potential in decentralized control, multi-objective optimization, and high-real-time control. Experimental results demonstrate that this method's computational efficiency is significantly higher than that of traditional Q-learning algorithms. This method can be applied to a variety of complex control systems, such as automated production, intelligent scheduling, distributed energy management, and traffic optimization, providing efficient, intelligent, and scalable control solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 This is a flow chart of a multi-stage Q-learning optimal control method based on system decoupling of the present invention; Figure 2 This is a UAV intersection conflict scenario in an embodiment of the present invention; Figure 3 In the embodiment of the present invention Changes within 200 iterations; Figure 4 In the embodiment of the present invention Changes within 200 iterations; Figure 5 In the embodiment of the present invention Changes within 200 iterations; Figure 6 In the embodiment of the present invention The convergence of Figure 7 In the embodiment of the present invention The convergence of Figure 8 The changes in the length of the UAV queue under the two control strategies in the embodiment of the present invention are shown; Figure 9 The changes in the travel time allocated to the UAV under the two control strategies in the embodiment of the present invention are shown; Figure 10 This shows the change in the number of drones in each direction in the embodiment of the present invention. DETAILED DESCRIPTION
[0079] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.
[0080] like Figure 1 As shown, this embodiment discloses a multi-stage Q-learning optimal control method based on system decoupling, comprising the following steps: S1. According to the general linear discrete time (DT) system, construct the cost function and Q function.
[0081] Assume that there exists the following linear DT system:
[0082] For the optimal control problem of the above system, it is assumed that both the control strategy and the disturbance strategy are feedback stable. Then, the goal is to design a set of feedback control strategies under the condition that the optimal disturbance is obtained, that is, to maximize the value function given below , which minimizes the value function:
[0083] and Written in vector form:
[0084]
[0085] Therefore, the control strategy and disturbance strategy designed for this set of value functions can be expressed as:
[0086]
[0087] because In the form of linear quadratic form, then the value function can be further expressed as:
[0088] is the solution of the following Game Riccati Equation (GARE):
[0089] Let the optimal strategy , , if the matrix It can be solved, then The control and disturbance gains of a value function can be expressed as:
[0090]
[0091] Define a set of Q functions as: in:
[0092] At this time, the Q function and the value function satisfy the following relationship:
[0093] therefore, and satisfy:
[0094] So, the The control and disturbance gains of the value function can be transformed into:
[0095]
[0096] S2. Decouple the control system and initialize the system number The specific steps are as follows: If the weights of the control system and the value function meet the decoupling conditions, the system can be decoupled into two subsystems. Furthermore, if the subsystems and their corresponding weights still meet the decoupling conditions, the subsystems can continue to be decoupled until all subsystems do not meet the decoupling conditions.
[0097] Furthermore, after decoupling the control system, the control system The input and perturbation strategy of the value function satisfy the following expression:
[0098] S3, Initialization value function number: .
[0099] S4. Initialize the gain number of the control strategy: .
[0100] S5. Set and initialize structural parameters and iteration number , select Detection Noise , , and appropriate parameters , The specific steps are as follows: The present invention uses the least squares method to iteratively approximate the value function, thereby obtaining the optimal control strategy. First, initialize the Q function , here It is non-optimal. Next, the iterative equation of the Q function is given as follows: therefore, and The following relationship exists:
[0101] Further:
[0102] Based on the above derivation process of iterative Q function, we can know that as long as we determine
[0103] The value of , then we can get Therefore, choose a suitable iterative algorithm and let When , then there will be , so that , .
[0104] The parameter structure of the Q function is as follows:
[0105] in:
[0106]
[0107] Furthermore, the two action parameter structures are designed as follows: ,
[0108] S6, strategy evaluation and strategy update. The specific implementation is as follows: In order to obtain , define the following objective function:
[0109] If the optimal strategy is After iterations, it must satisfy:
[0110] make:
[0111] So we have:
[0112] According to the form of the above formula, it can be solved by the least squares method After the introduction of detection noise, the parameter structure of the strategy becomes:
[0113]
[0114] Furthermore, the objective function is rewritten as: here, .
[0115] Further, It can be updated by the following least squares solution:
[0116] So, for each subsystem, there is:
[0117] Further, Can be achieved through Get, given Under the condition of , the updated strategy gain is as follows:
[0118]
[0119] Furthermore, the updated policy is as follows:
[0120]
[0121] S7, if , output the optimal strategy:
[0122] make , , And transfer to step 8, otherwise, go to step 6.
[0123] S8, if , go to step 5; otherwise, let , go to step 9.
[0124] S9, if , go to step 4; otherwise, end.
[0125] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc. Various media that can store program codes.
[0126] The following simulation example is given. The traffic control problem at drone intersections has significant research value in the future low-altitude economy field. This paper will verify the effectiveness of the multi-stage Q-learning optimal control method based on system decoupling through the traffic control model of a three-directional drone intersection. Consider Figure 2 The traffic flow model for the drone intersection conflict scenario shown is as follows:
[0127] in, , , It is a group of drones coming from all directions at a time The equivalent queue length is There are drones coming from all directions at the same time. The duration of the allocated pass instruction, It's time The number of newly added drones in each direction. Assume that the sampling step size is equal to a traffic cycle, and the traffic cycle .
[0128] In this problem, it is assumed that the drones entering from both sides of the main channel have a higher priority in the allocation of travel time than the drones entering from the secondary channel. Assume that the travel time allocated to the drones entering from both sides of the main channel is and , then we have:
[0129] Therefore, a set of value functions is given as: in,
[0130] According to the decoupling conditions and methods, the control system can be decomposed into two independent subsystems, namely:
[0131] in, , , , , , .
[0132] Since the transit period is constant, the strategy of the second subsystem can be obtained by simply solving the strategy of the first subsystem, that is:
[0133] in addition, , and The value does not affect and The value of , so it can be left undefined. Further, the weight is selected , , , , , .
[0134] Then, the reference solutions of the Riccati equation are: ,
[0135] Furthermore, there are:
[0136]
[0137] The final policy gain under different value functions is: Therefore, the optimal strategy gain is:
[0138] Next, the multi-stage Q-learning optimal control method based on system decoupling is applied to solve the optimal strategy, such as Figure 3 and Figure 4 As shown, and It converges after 200 iterations, as Figure 5 As shown, It also converges after 200 iterations. Figure 6 As shown, It converges to 0 after 200 iterations, as Figure 7 As shown, After 200 iterations, it converges to 0. The above simulation results show that the algorithm can converge to the vicinity of the reference solution, proving the effectiveness of this method.
[0139] In addition, the deep Q network (DQN) is used to solve the optimal strategy. The optimal strategy gain is:
[0140] Assumptions:
[0141] The optimal control strategies solved by the above two methods are applied to the control system respectively. The queue length, the allocated travel time, and the number growth of drones are shown as follows: Figure 8 , Figure 9 and Figure 10 The results in the figure show that the UAV queue lengths of both methods can converge at the 180th time step.
[0142] Furthermore, the simulation results of the two methods after 10,000 iterations are compared as shown in Table 1: Table 1 Comparison of simulation results
[0143] The results show that under the same number of iterations and the same convergence effect, the computational efficiency and value function performance of the proposed method are better than those of the DQN algorithm, which reflects the superiority of the multi-stage Q-learning optimal control method based on decoupling. In another aspect, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0144] On the other hand, the present invention further discloses a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0145] In another embodiment provided by the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the multi-stage Q-learning optimal control methods based on system decoupling in the above embodiments.
[0146] It is understandable that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts of the above methods.
[0147] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0148] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0149] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.
[0150] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A multi-stage Q-learning optimal control method based on system decoupling, characterized in that: The following steps are included: S1. Construct the cost function and Q function according to the general linear discrete-time system; S2. Decouple the system and initialize the subsystem number after decoupling; S3. Determine the multi-stage value function form based on the constructed cost function and Q function form, and initialize the stage number of the value function; S4, initializing the gain number of the subsystem control strategy of the current stage; S5. Set the structural parameters of the Q function and control strategy of the current stage, and initialize the iteration number of the control strategy; S6. Evaluate and update the control strategy; S7, check the convergence of the control strategy. If the control strategy converges, update the stage number of the value function and the gain number of the corresponding control strategy, and go to step S8; otherwise, update the iteration number of the strategy gain and go to step S6; S8. If the control strategies for each stage of the current subsystem have not yet been obtained, go to step S5; otherwise, go to step S9; S9. If the control strategies of all stages, i.e., all subsystems, have been obtained, the process ends; otherwise, the subsystem numbers are updated and the process goes to step S4.
2. The multi-stage Q-learning optimal control method based on system decoupling according to claim 1, characterized in that: Step S1 specifically includes: Assume that there exists the following linear DT system: Where, for Maintain system status, for Dimensional control strategy, is an unknown disturbance, , , ; For the optimal control problem of the above system, assume that both the control strategy and the disturbance strategy are feedback stable; under the condition that the optimal disturbance is obtained, that is, the value function given below is maximized, design a set of feedback control strategies , which minimizes the value function: in, , for the system in The utility function at time , , Respectively The weights of the value function on the state, control strategy and disturbance strategy, Represents the system Gain, and Respectively The control and perturbation strategies under the value function are written in vector form as follows: The control strategy and disturbance strategy designed for this set of value functions are expressed as: because In the form of linear quadratic form, then the value function is further expressed as: here is the solution of the following game Riccati equation GARE: Let the optimal strategy , ,in and are the control and disturbance gains respectively; if the matrix It can be solved, then The control and disturbance gains of the value function are expressed as: Define a set of Q functions as: in, is a matrix that stores Q function information, and: At this time, the Q function and the value function satisfy the following relationship: therefore, and satisfy: No. The control and disturbance gains of the value function are transformed into: 。 3. The multi-stage Q-learning optimal control method based on system decoupling according to claim 2, characterized in that: Step S2 specifically includes: If the weights of the control system and the value function meet the decoupling conditions, the system can be decoupled into two subsystems. Furthermore, if the subsystems and their corresponding weights still meet the decoupling conditions, the subsystems continue to decouple until all subsystems do not meet the decoupling conditions. The specific decoupling conditions and methods are as follows: If the control system satisfies: is the identity matrix, , the weight of the value function satisfies , , ,and , , , and The dimensions are the same; then there must be ,in, ; In addition, the control and policy gains satisfy: ,and , , and The dimensions are the same, , , and The dimensions are the same; in this case, and Expressed as: After decoupling the control system through the above method, the control system The input and perturbation strategy of the value function satisfy the following expression: , in , Indicates the The status of each subsystem, and Indicates the Control and disturbance strategies for each subsystem, and is the corresponding gain.
4. The multi-stage Q-learning optimal control method based on system decoupling according to claim 3, characterized in that: Step S5 specifically includes: The optimal control strategy is obtained by iteratively approximating the value function through the least square method; First initialize the Q function , here It is non-optimal. Next, the iterative equation of the Q function is given as follows: therefore, and The following relationship exists: but: The parameter structure of the Q function is as follows: in: here, Corresponding matrix Elements of The two action parameter structures are designed as follows: , Initialize iteration number .
5. The multi-stage Q-learning optimal control method based on system decoupling according to claim 4, characterized in that: Step S6 includes, In order to obtain , define the following objective function: If the optimal strategy is After iterations, it must satisfy: make: in, and , so we have: According to the form of the above formula, it can be solved by the least squares method ; Introducing detection noise, the parameter structure of the strategy becomes: in, ; The objective function is rewritten as: here, ; Updated by the following least squares solution: For each subsystem, there is: pass get; For each subsystem, select the detection noise , , and parameters that meet the requirements , ; In a given Under the condition of , the updated strategy gain is as follows: The updated policy is as follows: 。 6. The multi-stage Q-learning optimal control method based on system decoupling according to claim 5, characterized in that: Step S7 specifically includes: Pass inspection The convergence of is used to stop the policy iteration: like , output the optimal strategy: make , And transfer to step 8, otherwise let And return to step 6.
7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 6.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the computer program is executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-individual optimization control method based on non-strategy Q learning
CN110083063A
Optimal state consistency control method for multi-agent system
CN112445132A