A neural network adaptive formation cooperative optimization control method based on reinforcement learning

CN122815876APending Publication Date: 2026-09-25GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610921593.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本发明旨在解决编队协同控制中由执行器延迟、未知时变控制系数以及未知非线性动态所引发的系统不稳定、跟踪误差增大和振荡加剧等问题

Benefits of technology

[0047]本发明旨在解决编队协同控制中由执行器延迟、未知时变控制系数以及未知非线性动态所引发的系统不稳定、跟踪误差增大和振荡加剧等问题。通过所提出的控制方法,不仅能够保障编队内部各智能体的动态稳定性与整体编队的弦稳定性,还显著提升了系统的鲁棒性与控制精度。在此基础上,还提高道路通行效率、增强行驶安全性,并有效降低能源消耗,为多智能体编队在实际场景中的安全、高效、绿色运行提供有力支撑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122815876A_ABST
    Figure CN122815876A_ABST
Patent Text Reader

Abstract

The application discloses a neural network adaptive formation cooperative optimization control method based on reinforcement learning, and belongs to the field of multi-agent cooperative control. The method aims to effectively estimate and compensate the actuator delay, unknown time-varying control coefficient and unknown nonlinear dynamics commonly existing in a multi-agent formation system, so as to significantly improve the overall stability, robustness and cooperative control performance of the formation. The application comprises the following steps: establishing a third-order nonlinear formation dynamics model which comprehensively considers the actuator delay, unknown time-varying control coefficient and unknown nonlinear dynamics; establishing a Gaussian error function approximation strategy; establishing a formation tracking error; establishing a formation sliding mode error; establishing a reinforcement learning optimization algorithm based on an actuator-judge neural network system; establishing an adaptive parameter update rate and a neural network estimation mechanism aiming at unknown uncertainties in the formation; and establishing a formation adaptive sliding mode control strategy based on reinforcement learning, the Gaussian error function approximation strategy, the adaptive parameter update rate and the neural network estimation mechanism. The application is used for multi-agent cooperative control.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention discloses a neural network-based adaptive formation cooperative optimization control method based on reinforcement learning, belonging to the field of multi-agent cooperative control. This method aims to effectively estimate and compensate for actuator delays, unknown time-varying control coefficients, and unknown nonlinear dynamics commonly found in multi-agent formation systems, thereby significantly improving the overall stability, robustness, and cooperative control performance of the formation. The invention includes: establishing a third-order nonlinear formation dynamics model that comprehensively considers actuator delays, unknown time-varying control coefficients, and unknown nonlinear dynamics; establishing a Gaussian error function approximation strategy; establishing formation tracking error; establishing formation sliding mode error; establishing a reinforcement learning optimization algorithm based on an actuator-evaluator neural network system; establishing an adaptive parameter update rate and neural network estimation mechanism for unknown uncertainties in the formation; and establishing an adaptive sliding mode control strategy for formation based on reinforcement learning, the Gaussian error function approximation strategy, the adaptive parameter update rate, and the neural network estimation mechanism. This invention is used for multi-agent cooperative control. Technical Field

[0002] This invention belongs to the field of multi-agent cooperative control, and mainly relates to a neural network adaptive formation cooperative optimization control method based on reinforcement learning. Background Technology

[0003] my country's logistics and transportation industry has developed rapidly, and the number of cars has continued to rise. Consequently, the road traffic system faces social problems such as environmental pollution, traffic congestion, frequent accidents, and energy crises. Intelligent transportation systems (ITS) offer a key breakthrough in solving these road management challenges, and platooning cooperative control technology, as its core component, has become a key research area. Placing cooperative control refers to designing control strategies to drive a cluster of intelligent agents with autonomous decision-making and neighborhood communication capabilities to achieve consistent tracking of the desired trajectory while maintaining a compact topological spacing. This technology has profound research value for overcoming the bottleneck of traffic system efficiency and building a highly robust and safe system. However, during the long-term operation of intelligent agents, the execution units of these agents generally experience significant input delays due to the coupling effects of multiple factors such as mechanical wear, actuator response lag, and sensor computation delay. This not only leads to degradation of cooperative performance but also directly threatens the stability and operational safety of the platooning system. Furthermore, given the inherent characteristics of power fluctuations, nonlinearity of drive / braking mechanisms, and time-varying loads, the dynamics models of intelligent agents often contain unknown time-varying control coefficients. Coupled with the superposition of external uncertainties such as wind resistance disturbances, sudden operating conditions, and sensor measurement noise, these factors severely restrict the longitudinal robustness and cooperative consistency of multi-agent systems in complex dynamic environments. In summary, to meet the higher standards and requirements of intelligent systems for formation control, existing formation control methods need further improvement and refinement. Summary of the Invention

[0004] This invention aims to address the problems of system instability, increased tracking error, and exacerbated oscillations caused by actuator delays, unknown time-varying control coefficients, and unknown nonlinear dynamics in formation cooperative control. This invention provides a neural network-based adaptive formation cooperative optimization control method based on reinforcement learning.

[0005] A neural network-based adaptive formation cooperative optimization control method based on reinforcement learning aims to effectively estimate and compensate for actuator delays, unknown time-varying control coefficients, and unknown nonlinear dynamics prevalent in formation systems, thereby significantly improving the overall stability, robustness, and cooperative control performance of the formation. The control method includes the following steps:

[0006] Step 1: Establish a formation dynamics model that considers actuator delay, unknown time-varying control coefficients, and unknown nonlinear dynamics.

[0007] Formation dynamics model

[0008]

[0009] Where i is the agent's index, t represents time, and p i (t), v i (t) and a i (t) represent the position, velocity, and acceleration of the i-th agent, respectively, and α i (t) represents the unknown time-varying control coefficient, u i (t-τ i (t) represents a time delay τ i The control input for (t), f i (v i (t), a i (t) represents an unknown nonlinear function. This is due to external interference.

[0010] Step 2: Establish a Gaussian error function approximation strategy:

[0011] Gaussian error function approximation strategy

[0012]

[0013] Where exp is the exponential function, Z is the neural network input vector, and φ k (Z) Gaussian radial basis functions, where n is the number of neurons, and c k With b k These represent the center and width of the k-th hidden layer neural node, respectively.

[0014] f(Z) = θ *T φ(Z)+ε(Z) (3)

[0015] Where f(Z) is an arbitrarily continuous function, θ * Z represents the optimal weight vector, and ε(Z) represents the minimum approximation error.

[0016] Step 3: Establish formation tracking error:

[0017] Formation tracking error

[0018]

[0019] Among them, e t m (t) represents the position tracking error, d i (t)=p i-1 (t)-p i (t)-l i l represents the actual distance between adjacent agents. i Δ is the length of the i-th vehicle. i Represents the stationary distance, h is the front end of the agent, and K i It is a delayed compensation adjustment factor. Let τ represent the integral from time 0 to time f, where τ is the integration variable and β is the integral. i (t) Delay compensation variable, w i It is an auxiliary variable for correcting tracking error, Ψ i (ψ i ) is a smooth, monotonically increasing function, ρ(t) represents the preset performance function, and ψ i It represents the unconstrained transformation error, and ln is the natural logarithm. and These are design parameters.

[0020] Step 4: Establish formation sliding mode error:

[0021] First-order sliding mode error

[0022]

[0023] Among them, s i λ is the sliding mode error, and λ is the design parameter.

[0024] Coupled sliding mode error

[0025]

[0026] Among them, Π i It is the coupling sliding mode error, 0 < δ ≤ 1, and N represents the number of agents.

[0027] Step 5: Establish a reinforcement learning optimization algorithm with an integrated actor-critic architecture:

[0028] Reinforcement learning optimization algorithm with an integrated actor-critic architecture

[0029]

[0030] in, This represents the learning law of the criterion. F ci >0 is the learning rate. These are design parameters, I m It is the identity matrix. yes The estimated value.

[0031]

[0032] in, F represents the learning law of the actuator. ai >0 is the learning rate, and satisfies F ai >F ci .

[0033] Step Six: Establish adaptive parameter estimation mechanisms and neural network estimation mechanisms to address unknown uncertainties in formation:

[0034] Adaptive parameter estimation mechanism

[0035]

[0036] in, tanh represents the hyperbolic tangent function, F ψi >0 and σ ψi >0 is a design parameter, ζ i It is a constant. It is ψ i The estimated value.

[0037] Neural network estimation mechanism

[0038]

[0039] Among them, F αi >0 and σ αi >0 is a design parameter. α i >0 is α i The minimum value, yes The estimated value.

[0040]

[0041] Among them, F fi >0 and σ fi >0 is a design parameter. yes The estimated value.

[0042] Step 7: Establish an adaptive sliding mode control strategy for formation based on reinforcement learning, Gaussian error function approximation, adaptive parameter estimation, and neural network estimation.

[0043] Adaptive sliding mode control strategy for formation based on reinforcement learning, Gaussian error function approximation strategy, adaptive parameter estimation mechanism, and neural network estimation mechanism.

[0044]

[0045] in,

[0046] ξ i >0 and These are design parameters.

[0047] This invention aims to address the problems of system instability, increased tracking error, and exacerbated oscillations caused by actuator delays, unknown time-varying control coefficients, and unknown nonlinear dynamics in platoon cooperative control. The proposed control method not only ensures the dynamic stability of each agent within the platoon and the overall chordal stability of the platoon, but also significantly improves the system's robustness and control accuracy. Furthermore, it enhances road traffic efficiency, improves driving safety, and effectively reduces energy consumption, providing strong support for the safe, efficient, and environmentally friendly operation of multi-agent platoons in real-world scenarios. Attached Figure Description

[0048] Figure 1 This is a flowchart illustrating the control method described in Specific Implementation Method 1. Detailed Implementation

[0049] Specific implementation method one: Combining Figure 1 This embodiment describes a neural network adaptive formation cooperative optimization control method based on reinforcement learning. The control method includes the following steps:

[0050] Step 1: Establish a formation dynamics model that considers actuator delay, unknown time-varying control coefficients, and unknown nonlinear dynamics.

[0051] Formation dynamics model

[0052]

[0053] Where i is the agent's index, t represents time, and p i (t), v i (t) and a i (t) represent the position, velocity, and acceleration of the i-th agent, respectively, and α i(t) represents the unknown time-varying control coefficient, u i (t-τ i (t) represents a time delay τ i The control input for (t), f i (v i (t), a i (t) represents an unknown nonlinear function. This is due to external interference.

[0054] Step 2: Establish a Gaussian error function approximation strategy:

[0055] Gaussian error function approximation strategy

[0056]

[0057] Where exp is the exponential function, Z is the neural network input vector, and φ k (Z) Gaussian radial basis functions, where n is the number of neurons, and c k With b k These represent the center and width of the k-th hidden layer neural node, respectively.

[0058] f(Z) = θ *T φ(Z)+ε(Z) (15)

[0059] Where f(Z) is an arbitrarily continuous function, θ * Z represents the optimal weight vector, and ε(Z) represents the minimum approximation error.

[0060] Step 3: Establish formation tracking error:

[0061] Formation tracking error

[0062]

[0063] Among them, e i m (t) represents the position tracking error, d i (t)=p i-1 (t)-p i (t)-l i l represents the actual distance between adjacent agents. i Δ is the length of the i-th vehicle. i Represents the stationary distance, h is the front end of the agent, and K i It is a delayed compensation adjustment factor. Let τ represent the integral from time 0 to time t, where τ is the integration variable and β is the integral. i (t) Delay compensation variable, w i It is an auxiliary variable for correcting tracking error, Ψ i (ψ i) is a smooth, monotonically increasing function, ρ(t) represents the preset performance function, and ψ i It represents the unconstrained transformation error, and ln is the natural logarithm. and These are design parameters.

[0064] Step 4: Establish formation sliding mode error:

[0065] First-order sliding mode error

[0066]

[0067] Among them, s i λ is the sliding mode error, and λ is the design parameter.

[0068] Coupled sliding mode error

[0069]

[0070] Among them, Π i It is the coupling sliding mode error, 0 < δ ≤ 1, and N represents the number of vehicles.

[0071] Step 5: Establish a reinforcement learning optimization algorithm with an integrated actor-critic architecture:

[0072] Reinforcement learning optimization algorithm with an integrated actor-critic architecture

[0073]

[0074] in, This represents the learning law of the criterion. F ci >0 is the learning rate. These are design parameters, I m It is the identity matrix. yes The estimated value.

[0075]

[0076] in, F represents the learning law of the actuator. ai >0 is the learning rate, and satisfies F ai >F ci .

[0077] Step Six: Establish adaptive parameter estimation mechanisms and neural network estimation mechanisms to address unknown uncertainties in formation:

[0078] Adaptive parameter estimation mechanism

[0079]

[0080] in, tanh represents the hyperbolic tangent function, F ψi >0 and σ ψi >0 is a design parameter, ζ i These are design parameters. It is ψ i The estimated value.

[0081] Neural network estimation mechanism

[0082]

[0083] Among them, F αi >0 and σ αi >0 is a design parameter. α i >0 is α i The minimum value, yes The estimated value.

[0084]

[0085] Among them, F fi >0 and σ fi >0 is a design parameter. yes The estimated value.

[0086] Step 7: Establish an adaptive sliding mode control strategy for formation based on reinforcement learning, Gaussian error function approximation, adaptive parameter estimation, and neural network estimation.

[0087] Adaptive sliding mode control strategy for formation based on reinforcement learning, Gaussian error function approximation strategy, adaptive parameter estimation mechanism, and neural network estimation mechanism.

[0088]

[0089] in,

[0090] ξ i >0 and These are design parameters.

[0091] Effects of this implementation method:

[0092] This invention effectively solves the problems of system instability, increased tracking error, and aggravated oscillation caused by actuator delay, unknown time-varying control coefficients, and unknown nonlinear dynamics in platoon cooperative control. The proposed control method not only ensures the dynamic stability of each agent within the platoon and the chordal stability of the overall platoon, but also significantly improves the system's robustness and control accuracy. Furthermore, it enhances road traffic efficiency, improves driving safety, and effectively reduces energy consumption, providing strong support for the safe, efficient, and environmentally friendly operation of multi-agent platoons in real-world scenarios.

Claims

1. A neural network-based adaptive formation cooperative optimization control method based on reinforcement learning aims to effectively estimate and compensate for actuator delays, unknown time-varying control coefficients, and unknown nonlinear dynamics commonly found in fleet systems, thereby significantly improving the overall stability, robustness, and cooperative control performance of the formation. The control method includes the following steps: Step 1: Establish a formation dynamics model that considers actuator delay, unknown time-varying control coefficients, and unknown nonlinear dynamics; Step 2: Establish a Gaussian error function approximation strategy; Step 3: Establish formation tracking error; Step 4: Establish formation sliding mode error; Step 5: Establish a reinforcement learning optimization algorithm based on the executor-evaluator neural network system; Step 6: Establish an adaptive parameter estimation mechanism and a neural network estimation mechanism to address unknown uncertainties in formation; Step 7: Establish an adaptive sliding mode control strategy for formation based on reinforcement learning, Gaussian error function approximation strategy, adaptive parameter estimation mechanism and neural network estimation mechanism; In step one, Formation dynamics model in, i is the agent's index, t represents time, and p i (t), v i (t) and a i (t) represent the position, velocity, and acceleration of the i-th agent, respectively, and α i (t) represents the unknown time-varying control coefficient, u i (t-τ i (t) represents the time delay τ. i The control input for (t), f i (v i (t), a i (t) represents an unknown nonlinear function. This is due to external interference. In step two Gaussian error function approximation strategy Where exp is the exponential function, Z is the neural network input vector, and φ k (Z) Gaussian radial basis functions, where n is the number of neurons, and c k With b k These represent the center and width of the k-th hidden layer neural node, respectively. f(Z)=θ *T φ(Z)+ε(Z) (3) Where f(Z) is an arbitrarily continuous function, θ * Z represents the optimal weight vector, and ε(Z) represents the minimum approximation error. In step three Formation tracking error Among them, e i m (t) represents the position tracking error, d i (t)=p i-1 (t)-p i (t)-l i l represents the actual distance between adjacent agents. i It is the length of the agent, Δ i Represents the stationary distance, h is the front end of the agent, and K i It is a delayed compensation adjustment factor. Let τ represent the integral from time 0 to time t, where τ is the integration variable and β is the integral. i (t) Delay compensation variable, w i It is an auxiliary variable for correcting tracking error, Ψ i (ψ i ) is a smooth, monotonically increasing function, ρ(t) represents the preset performance function, and ψ i It represents the unconstrained transformation error, and ln is the natural logarithm. and These are design parameters. In step four, First-order sliding mode error Among them, s i λ is the sliding mode error, and λ is the design parameter. Coupled sliding mode error Among them, Π i It is the coupling sliding mode error, 0 < δ ≤ 1, and N represents the number of agents. In step five, Reinforcement learning optimization algorithm based on actuator-evaluator neural network system in, This represents the learning law of the criterion. F ci >0 is the learning rate. These are design parameters, I m It is the identity matrix. yes The estimated value. in, F represents the learning law of the actuator. ai >0 is the learning rate, and satisfies F ai >F ci . In step six, Adaptive parameter estimation mechanism in, tanh represents the hyperbolic tangent function, F ψi >0 and σ ψi >0 is a design parameter, ζ i It is a constant. It is ψ i The estimated value. Neural network estimation mechanism Among them, F αi >0 and σ αi >0 is a design parameter. α i >0 is α i The minimum value, yes The estimated value. Among them, F fi >0 and σ fi >0 is a design parameter. yes The estimated value. In step seven, Adaptive sliding mode control strategy for formation based on reinforcement learning, Gaussian error function approximation strategy, adaptive parameter estimation mechanism, and neural network estimation mechanism. in, ξ i >0 and These are design parameters.