An integral reinforcement learning-based hybrid-order system event-triggered cooperative control strategy

CN122613745APending Publication Date: 2026-08-21GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610921526.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0004]本发明的目的是为了解决在系统动力学未知的条件下,如何实现最优控制与优化系统资源调度,从而克服现有需依赖精确的系统模型适应力差以及通信资源利用率低的问题,本发明提供了一种基于积分强化学习的混合阶协同队列控制策略

Benefits of technology

[0047]本发明的有益效果在于通过积分强化学习实现数据驱动的最优控制,无需精确系统模型,有效应对未知非线性与扰动;结合事件触发机制,仅在必要时进行控制,有效地降低资源消耗,提升系统运行效率与可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122613745A_ABST
    Figure CN122613745A_ABST
Patent Text Reader

Abstract

The application discloses a hybrid order cooperative queue control strategy based on integral reinforcement learning, belongs to the field of intelligent control, and mainly solves the problems of unknown nonlinear dynamics and system resource waste in a hybrid order system through an integral reinforcement learning algorithm and an event triggering mechanism, so as to realize optimal control of the system and saving of system resources. A hybrid order queue system model with an actuator saturation, unknown nonlinear dynamics and external disturbance is established; a queue cooperative tracking error is established; a distributed hybrid order sliding mode surface is established; an integral reinforcement learning mechanism based on a neural network is established; and an event triggering cooperative control strategy of the hybrid order system based on the integral reinforcement learning is established. The application is used for intelligent control of a cooperative queue.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A hybrid-order cooperative queue control strategy based on integral reinforcement learning, belonging to the field of intelligent control, mainly addresses the problems of unknown nonlinear dynamics and wasted system resources in hybrid-order systems through integral reinforcement learning algorithms and event-triggered mechanisms, thereby achieving optimal system control and resource conservation. The invention establishes a hybrid-order queue system model with actuator saturation, unknown nonlinear dynamics, and external disturbances; establishes a queue cooperative tracking error; establishes a distributed hybrid-order sliding surface; establishes a neural network-based integral reinforcement learning mechanism; and establishes an event-triggered cooperative control strategy for the hybrid-order system based on integral reinforcement learning. This invention is used for intelligent control of cooperative queues. Technical Field

[0002] This invention belongs to the field of intelligent control, and mainly relates to a hybrid-order cooperative queue control strategy based on integral reinforcement learning. Background Technology

[0003] In cooperative queuing control systems, achieving synchronization and tracking of the states of multiple agents is crucial for enhancing the overall system's reliability, safety, and energy efficiency. However, existing control methods still face significant challenges in handling unknown nonlinearities in vehicle dynamics, input saturation constraints, and external disturbances. Traditional optimal control methods often rely on precise system mathematical models, requiring the identification of unknown dynamics through fuzzy systems or neural networks. This is not only time-consuming but also introduces additional identification errors, affecting the system's convergence speed and steady-state performance. Reinforcement learning, as a data-driven online learning method, can learn optimal control strategies without relying on precise system models, significantly enhancing the controller's adaptability under uncertainty. Furthermore, traditional sampled-data control systems use time-triggered methods, where system inputs periodically act on the controlled object regardless of whether the control state needs updating. Event-triggered control, on the other hand, drives agents to communicate, compute, and control output only when necessary, based on preset conditions, resulting in more scientific and efficient resource scheduling. By designing a reasonable triggering mechanism, the transmission and control frequency of agents can be optimized, effectively saving control resources and improving system operating efficiency. In practical applications, frequent triggering increases energy consumption and component wear, impacting system performance. In the design of multi-agent control strategies, energy consumption reduction and component lifespan extension should be fully considered to effectively control operating costs. Therefore, there is an urgent need to design a hybrid-order cooperative queue control strategy based on integral reinforcement learning to achieve optimal system control and save system resources. This invention is proposed precisely to address this technical need. Summary of the Invention

[0004] The purpose of this invention is to solve the problem of how to achieve optimal control and optimize system resource scheduling under the condition of unknown system dynamics, thereby overcoming the problems of poor adaptability and low communication resource utilization of existing systems that rely on accurate system models. This invention provides a hybrid-order cooperative queue control strategy based on integral reinforcement learning.

[0005] A hybrid-order cooperative queue control based on integral reinforcement learning is characterized by addressing the problems of unknown nonlinear dynamics and wasted system resources in hybrid-order systems through integral reinforcement learning algorithms and event-triggered mechanisms, thereby achieving optimal system control and resource conservation. The control strategy includes the following steps:

[0006] Step 1: Establish a mixed-order queue system model with actuator saturation, unknown nonlinear dynamics, and external disturbances:

[0007] Second-order queue system model

[0008]

[0009] Three-order queue system model

[0010]

[0011] Where i is the sequence number, s i (t), v i (t) and α i (t) represents the position, velocity, and acceleration of the i-th agent, respectively, and t represents time. Represents a column vector containing position, velocity, and acceleration. It is the uncertainty that the system is subject to, u i (t) is the control input. It is the mass coefficient of the agent, τ i N1 is the system delay constant, N2 represents the set of second-order agents, and N2 represents the set of third-order agents.

[0012] Step 2: Establish queue-based collaborative error tracking:

[0013] Cooperative queue tracking error

[0014]

[0015] Where ξ s,i (t), ξ v,i (t) and ξ a,i (t) represents the tracking errors for position, velocity, and acceleration, respectively. It is the safe distance between adjacent intelligent agents. Let s0(t), v0(t), and a0(f) be the safe distance between agent i and the leader, and let s0(t), v0(t), and a0(f) be the position, velocity, and acceleration of the leader in the system. It is the set of neighboring agents of agent i. It is the set of neighboring agents of agent i with third order, α ij Let b represent the adjacency matrix. i Let represent the connection weight between agent i and the leader, and N represent the set containing all agents.

[0016] Step 3: Establish a distributed hybrid sliding surface:

[0017] Distributed hybrid sliding surface

[0018]

[0019] Where, σ i (t) represents the sliding mode error. and All are sliding mode control coefficients.

[0020] Step 4: Establish an integral reinforcement learning mechanism based on neural networks:

[0021] Integral reinforcement learning algorithm

[0022]

[0023] Among them, Q i (σ i (t) represents the cost function. Indicates from 0 to u i The integral, where τ is the integration variable, A ii Let λ be a positive definite matrix. i >0.

[0024] Neural Network Approximation

[0025]

[0026] in, This represents an estimate of the expected cost function. It is the activation function, W i These are ideal weights. It is an estimate of the ideal weights, μ i It is the approximation error of the neural network.

[0027] Based on the time difference error of event triggering

[0028]

[0029] Where, θ i It is time difference error. T is the time constant. This represents the integral from time t to time t+T. is u i The estimated value, t i,g >0 represents the sampling time.

[0030] Differential error based on event triggering and experience playback techniques

[0031]

[0032] Where, θ i,h It is a differential error based on empirical playback technology. is u i The estimated value, Indicates from time t i,h By time t i,h The integral of +T, t i,h >0 represents the h-th historical data point collected.

[0033] Neural Network Update Law

[0034]

[0035] Where, β i h represents the learning rate. m ζ represents the total number of historical data points. i >0.

[0036] Step 5: Establish an event-triggered collaborative control strategy for a hybrid-order system based on integral reinforcement learning.

[0037] Event triggering error

[0038]

[0039] in, Represents the cost function, σ i,g =σ i (t i,g ).

[0040] Event triggering conditions

[0041]

[0042] Here, inf{} represents the infm. ζ i >0, |||| denotes the norm, o i Values ​​greater than 0 are ideal weights. Ψ () represents the smallest eigenvalue.

[0043] Event-triggered cooperative control strategy for hybrid-order systems based on integral reinforcement learning

[0044]

[0045] in, l ii b is the diagonal element of the Laplace matrix. ii These are the diagonal elements of the traction matrix.

[0046] The beneficial effects of this invention are mainly reflected in:

[0047] The beneficial effects of this invention are that it achieves data-driven optimal control through integral reinforcement learning, without the need for a precise system model, and effectively copes with unknown nonlinearities and disturbances; combined with an event-triggered mechanism, it only performs control when necessary, effectively reducing resource consumption and improving system operating efficiency and reliability. Attached Figure Description

[0048] Figure 1 This is a flowchart illustrating the event-triggered cooperative control strategy for a hybrid-order system based on integral reinforcement learning, as described in Implementation Method 1. Detailed Implementation

[0049] Specific implementation method one: Combining Figure 1 This embodiment describes an event-triggered cooperative control strategy for a hybrid-order system based on integral reinforcement learning. The cooperative queue control strategy includes the following steps:

[0050] Step 1: Establish a mixed-order queue system model with actuator saturation, unknown nonlinear dynamics, and external disturbances:

[0051] Second-order queue system model

[0052]

[0053] Three-order queue system model

[0054]

[0055] Where i is the sequence number, s i (t), v i (t) and a i( t) represent the position, velocity, and acceleration of the i-th agent, respectively, and t represents time. Represents a column vector containing position, velocity, and acceleration. It is the uncertainty that the system is subject to, u i (t) is the control input. It is the mass coefficient of the agent, τ i N1 is the system delay constant, N2 represents the set of second-order agents, and N2 represents the set of third-order agents.

[0056] Step 2: Establish queue-based collaborative error tracking:

[0057] Cooperative queue tracking error

[0058]

[0059] Where ξ s,i (t), ξ v,i (t) and ξ a,i (t) represents the tracking errors for position, velocity, and acceleration, respectively. It is the safe distance between adjacent intelligent agents. Let s0(t), v0(t), and a0(t) be the safe distance between agent i and the leader, and let s0(t), v0(t), and a0(t) be the position, velocity, and acceleration of the leader in the system, respectively. It is the set of neighboring agents of agent i. It is the set of neighboring agents of agent i with third order, α ij Let b represent the adjacency matrix. i Let represent the connection weight between agent i and the leader, and N represent the set containing all agents.

[0060] Step 3: Establish a distributed hybrid sliding surface:

[0061] Distributed hybrid sliding surface

[0062]

[0063] Where, σ i (t) represents the sliding mode error. and All are sliding mode control coefficients.

[0064] Step 4: Establish an integral reinforcement learning mechanism based on neural networks:

[0065] Integral reinforcement learning algorithm

[0066]

[0067] Among them, Q i (σ i (t) represents the cost function. Indicates from 0 to u i The integral, where τ is the integration variable, A ii Let λ be a positive definite matrix. i >0.

[0068] Neural Network Approximation

[0069]

[0070] in, This represents an estimate of the expected cost function. It is the activation function, W i These are ideal weights. It is an estimate of the ideal weights, μ i It is the approximation error of the neural network.

[0071] Based on the time difference error of event triggering

[0072]

[0073] Where, θ i It is time difference error. T is the time constant. This represents the integral from time t to time t+T. is u i The estimated value, t i,g >0 represents the sampling time.

[0074] Differential error based on event triggering and experience playback techniques

[0075]

[0076] Where, θ i,h It is a differential error based on empirical playback technology. is u i The estimated value, Indicates from time t i,h By time t i,h The integral of +T, t i,h >0 represents the h-th historical data point collected.

[0077] Neural Network Update Law

[0078]

[0079] Where, β i h represents the learning rate. m ζ represents the total number of historical data points. i >0.

[0080] Step 5: Establish an event-triggered collaborative control strategy for a hybrid-order system based on integral reinforcement learning.

[0081] Event triggering error

[0082]

[0083] in, Represents the cost function, σ i,g =σ i (t i,g ).

[0084] Event triggering conditions

[0085]

[0086] Here, inf{} represents the infm. ζ i >0, |||| denotes the norm, o i Values ​​greater than 0 are ideal weights. Ψ () represents the smallest eigenvalue.

[0087] Event-triggered cooperative control strategy for hybrid-order systems based on integral reinforcement learning

[0088]

[0089] in, l ii b is the diagonal element of the Laplace matrix. ii These are the diagonal elements of the traction matrix.

[0090] This invention achieves data-driven optimal control through integral reinforcement learning, eliminating the need for a precise system model and effectively addressing unknown nonlinearities and disturbances. Combined with an event-triggered mechanism, control is only performed when necessary, effectively reducing resource consumption and improving system operating efficiency and reliability.

Claims

1. A hybrid-order cooperative queue control based on integral reinforcement learning, characterized by: By employing integral reinforcement learning algorithms and event-triggered mechanisms, this study addresses the issues of unknown nonlinear dynamics and wasted system resources in hybrid-order systems, aiming to achieve optimal system control and resource conservation. The control strategy includes the following steps: Step 1: Establish a model of a hybrid-order queue system with actuator saturation, unknown nonlinear dynamics, and external disturbances; Step 2: Establish a queue for collaborative error tracking; Step 3: Establish a distributed hybrid sliding surface; Step 4: Establish an integral reinforcement learning mechanism based on neural networks; Step 5: Establish an event-triggered collaborative control strategy for a hybrid-order system based on integral reinforcement learning; In step one, Second-order queue system model Three-order queue system model Where i is the sequence number, s i (t), v i (t) and a i (t) represents the position, velocity, and acceleration of the i-th agent, respectively, and t represents time. Represents a column vector containing position, velocity, and acceleration. It is the uncertainty that the system is subject to, u i (t) is the control input. It is the mass coefficient of the agent, τ i N1 is the system delay constant, N2 represents the set of second-order agents, and N2 represents the set of third-order agents. In step two, Cooperative queue tracking error Where ξ s,i (t), ξ v,i (t) and ξ a,i (t) represents the tracking errors for position, velocity, and acceleration, respectively. It is the safe distance between adjacent intelligent agents. Let s0(t), v0(t), and a0(t) be the safe distance between agent i and the leader, and let s0(t), v0(t), and a0(t) be the position, velocity, and acceleration of the leader in the system, respectively. It is the set of neighboring agents of agent i. It is the set of neighboring agents of agent i with third order, a ij Let b represent the adjacency matrix. i Let represent the connection weight between agent i and the leader, and N represent the set containing all agents. In step three Distributed hybrid sliding surface Where, σ i (t) represents the sliding mode error. and All are sliding mode control coefficients. In step four, Integral reinforcement learning algorithm Among them, Q i (σ i (t) represents the cost function. C ii >0, Indicates from 0 to u i The integral, where τ is the integration variable, A ij Let λ be a positive definite matrix. i >0. Neural Network Approximation in, This represents an estimate of the expected cost function. It is the activation function, W i These are ideal weights. It is an estimate of the ideal weights, μ i It is the approximation error of the neural network. Based on the time difference error of event triggering Where, θ i It is time difference error. T is the time constant. This represents the integral from time t to time t+T. is u i The estimated value, t i,g >0 represents the sampling time. Differential error based on event triggering and experience playback techniques Where, θ i,h It is a differential error based on empirical playback technology. is u i The estimated value, Indicates from time t i,h By time t i,h The integral of +T, t i,h >0 represents the h-th historical data point collected. Neural Network Update Law Where, β i h represents the learning rate. m This indicates the total number of historical data points. In step five, Event triggering error in, Represents the cost function, σ i,g =σ i (t i,g ). Event triggering conditions Here, inf{} represents the infm. ζ i >0, ||||| represents the norm, o i Values ​​greater than 0 are ideal weights. ψ ( ) represents the smallest eigenvalue. Event-triggered cooperative control strategy for hybrid-order systems based on integral reinforcement learning in, l ii b is the diagonal element of the Laplace matrix. ii These are the diagonal elements of the traction matrix.