Fixed-time safety tracking control method for bolting robot under multi-learning tunnel performance

By combining composite learning control and tunnel-type preset performance functions with nonlinear fast integral terminal sliding mode control, the problem of over-tightening or under-tightening during the tightening process of the bolt operation robot is solved, safe tracking control and optimization within a fixed time are achieved, and the accuracy and safety of bolt tightening are improved.

CN119828456BActive Publication Date: 2025-09-26CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411818398.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-09-26
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Existing bolt tightening methods rely on manual operation, which is prone to over-tightening or under-tightening, leading to potential hidden dangers. In addition, existing automated control methods are difficult to achieve accurate fixed time parameter estimation and safety tracking control in practical applications.

Method used

Composite learning control is adopted in combination with tunnel-type preset performance function and nonlinear fast integral terminal sliding mode control to eliminate the continuous excitation condition and design a fixed-time convergence controller. Combined with the reinforcement learning method of the Actor-Critic framework, the control strategy is optimized to ensure safety and optimality.

Benefits of technology

The bolt operation robot can accurately tighten the bolts within a limited time, avoid over-tightening or under-tightening, ensure the safety of the system and the optimality of the control scheme, and simplify the implementation process of the controller.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119828456B_ABST
    Figure CN119828456B_ABST
Patent Text Reader

Abstract

The present invention relates to a fixed-time safety tracking control method for a bolt-operating robot under multi-learning tunnel performance, which belongs to the field of robot adaptation. The method first establishes a system model of the bolt-operating robot, and designs a series of important conversion functions, introduces the tunnel-type preset performance function into the tracking error constraint control, and realizes the global control effect. The continuous excitation condition is eliminated by composite learning control, and the weak excitation condition of the interval excitation is realized. A new type of nonlinear fast integral terminal sliding mode fixed-time convergence controller is designed to ensure the fixed-time convergence of the tracking error while avoiding the singularity problem. Finally, the above control scheme is combined with the Actor-Critic architecture of reinforcement learning to form a new type of safe reinforcement learning method, which realizes the optimization of the control scheme while ensuring the safety of the robot's tracking performance. This method improves the accuracy and efficiency of the bolt-operating robot and ensures the safety of the operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of robot adaptation and relates to a fixed time safety tracking control method for a bolt operation robot under multiple learning tunnel performances. Background Art

[0002] During the bolt tightening process, the tightening effect of the bolts will directly affect the production quality. However, in some current manufacturing processes, bolt tightening methods are still traditional manual operations. When workers rely on experience to tighten bolts, it is very easy to "over-tighten" or "under-tighten" the bolts, resulting in potential hazards such as fastener loosening and breakage. More serious cases may cause incalculable consequences. Applying intelligent assembly technology to production and manufacturing will improve production and manufacturing levels while saving human resources. This work is of great social value. At present, the precision performance and intelligent operation of some bolt robots in operation are not high and need to be improved. Therefore, it is a general trend to study a high-precision, high-efficiency, and high-reliability automated bolt tightening control algorithm.

[0003] To achieve satisfactory performance, many robot control methods have been studied over the past few decades, such as sliding mode control, model predictive control, and adaptive control. Among these methods, adaptive control has been well applied due to its online learning capabilities to update control parameters and handle the uncertainty of robot models.

[0004] Generally, parameter estimation approaches in adaptive control include two different approaches: indirect approaches that use the prediction error of the filtered control torque to generate parameter estimates, and direct approaches that use the tracking error to drive parameter adaptation. However, these approaches can only guarantee asymptotic control and parameter estimation under continuous excitation conditions, which is very demanding and often impractical in practice. Composite adaptive control is a control strategy that combines direct and indirect adaptation, in which the parameter adaptation law is driven by both the tracking error and the prediction error. Using composite adaptive control, smaller tracking errors and faster parameter estimation can be achieved through higher adaptive gains without exciting high-frequency unmodeled dynamics. However, a drawback of composite adaptive control is that it still requires continuous excitation conditions to ensure accurate parameter estimation. In this regard, researchers have proposed a new technique called composite learning control to achieve parameter convergence under more achievable conditions, called interval excitation conditions, rather than strict continuous excitation conditions. In composite learning, the time interval integral of the filtered regressor is used to construct the prediction error, and the prediction error and tracking error are used to update the parameter estimate, resulting in accurate and smooth parameter estimates. Researchers have proposed composite learning robot control for robotic systems to ensure parameter convergence. Other researchers have proposed a least-squares modulated composite learning control based on the Moore-Penrose pseudo-inverse to achieve a balanced and easily adjustable parameter convergence rate. Composite learning control has also been applied to further improve parameter convergence in hydraulic systems. Combining composite learning control with neural networks has led to the development of a neural network-based robot control system that can handle unknown disturbances.

[0005] Although the aforementioned composite learning control schemes can guarantee accurate parameter estimation and system control under weak interval excitation conditions, in most cases, convergence is asymptotic, which in principle implies an infinite convergence time. Clearly, both theoretically and practically, finite / fixed-time convergence is preferable to exponential convergence due to its faster convergence performance. Terminal sliding mode control is considered an effective method for achieving finite-time convergence. However, when terminal sliding mode control is combined with adaptive control to estimate unknown system parameters, potential singularity issues may arise. A two-phase robust control strategy can be used to address this issue, further maintaining a conservative control gain. Some researchers have proposed an adaptive optimal finite-time parameter estimation and tracking control method for a benchmark servo system. However, these methods cannot guarantee finite-time parameter estimation without strict continuous excitation conditions. In response, a finite-time composite learning control method for a robotic manipulator based on interval excitation has been proposed. The parameter estimates of this method converge to a small tight set of their real values ​​within finite time, while the tracking error converges exponentially to a small tight set. Currently, there is no feasible composite control scheme based on interval excitation to ensure accurate parameter estimation and robot control within a finite time, and fixed-time schemes are even more difficult.

[0006] In robot tracking control with tracking error constraints, the concept of a preset performance function has been widely used and analyzed. This approach constrains the robot's error to a predetermined performance function, thereby ensuring the robot's safety. However, most preset performance control methods implicitly assume that the performance constraint is satisfied from the outset of system operation, implying perfect knowledge of the initial conditions. These methods are ineffective if initial condition conflicts occur, as they are not defined outside the allowed set. From a global perspective, they produce local tracking results. However, in practical scenarios, delayed error constraints are often encountered. Tracking errors may initially violate the performance constraint and then adjust to the allowed set within a reasonable user-specified time. A typical example is when a robot first traverses an open area without any obstacles and then, after a certain period of time, enters a space surrounded by obstacles. In this case, it is necessary to impose certain constraints on the robot's behavior to regulate its path, avoid obstacles, and ensure the safety of the robot system. Most preset performance control schemes require that the robot's initial error be constrained within the preset performance function, and many performance functions, such as the funnel-shaped performance function, are symmetrically distributed on both sides of the error. This results in the performance function being only applicable to symmetrical performance constraints and having strict restrictions on the initial tracking conditions. In addition, the funnel-shaped performance leads to loose control performance. As the actual requirements of engineering projects continue to increase, more stringent requirements are often placed on its transient performance.

[0007] Ensuring safety and good performance under various circumstances remains a key technical and practical challenge for the widespread deployment of robotic systems. Considering both adaptability and safety in robotics has been a research hotspot. Reinforcement learning, as a learning-based sequential control method, has attracted strong interest and attention when system uncertainty exists. However, various limitations exist in practical applications, especially for robotic systems with safety requirements. One of the main reasons for these limitations is the lack of safety guarantees. When inevitable system uncertainty exists in robotic systems, it will lead to unpredictable and harmful situations. To optimize the control solution, the optimal control problem must be considered. The traditional solution is to calculate the optimal control law by solving the Hamilton-Jacobi-Bellman equation, which adopts the Bellman optimality principle. However, the HJB equation contains inherent nonlinearities and is often quite complex, making it impossible to directly calculate its unique solution in most cases.

[0008] Reinforcement learning is a widely discussed and researched method that uses neural networks to iteratively approximate the optimal solution of the HJB function, avoiding the direct solution of the HJB equation. For systems with strict feedback, some researchers have proposed optimized backstepping control, which involves learning virtual control of each subsystem in a linked nonlinear system. Within the actor-critic framework, this method employs optimized backstepping techniques to optimize overall system control, achieved by optimizing the virtual control of each backstepping subsystem. Using reinforcement learning, the critic and actor components are independently approximated. Through iterative learning, a program is developed that satisfies the Bellman optimality principle for the continuous-time HJB equation. While this method and structure are relatively simple, they fail to consider the system's constraints. For most reinforcement learning-based systems designed without safety guarantees, even if ultimate control performance is achieved through learning, the state variables may escape the safe zone during the learning process or when the system is perturbed, leading to failure.

[0009] Some researchers have combined the Lyapunov obstacle method with reinforcement learning, calling it an adaptive control method for safety reinforcement learning. Safety reinforcement learning is a promising approach for safety-critical systems, optimizing control schemes while satisfying safety constraints. Research on safety reinforcement learning falls into two main categories: model-free and model-based approaches. Model-free safety reinforcement learning methods achieve safety performance only after a sufficient number of learning episodes; this can be violated during learning interactions, especially in the early stages. This characteristic also makes them difficult to respond safely to changing environments. Model-based safety reinforcement methods either directly determine actions during the learning process or supervise them through a backup controller to ensure safety. This means that safety assurance and action exploration are separated in the design. Model predictive control is a typical model-based approach that utilizes precise models for safe action exploration in reinforcement learning, often attempting to pre-compute actions in advance to ensure safety performance.

[0010] Compared to the obstacle Lyapunov function, the preset performance function has better constraint performance, and the constraint performance is set in advance, which can effectively ensure the safety of the robot system during movement. However, there is currently no robot adaptive safe tracking control method that combines the preset performance function with reinforcement learning. Summary of the Invention

[0011] In view of this, the object of the present invention is to provide a fixed time safety tracking control method for a bolt operation robot under multiple learning tunnel performances.

[0012] The intelligent bolting robot contemplated by this patent is designed to achieve bolt tightening in the intelligent manufacturing process. Considering typical scenarios often encountered by bolting robots during bolt tightening, namely, due to the heavy weight and multiple threads of bolts in large equipment, the bolting robot is not subject to any error constraints in the initial tightening phase. However, after a certain period of time, in order to achieve the goal of tightening the bolts without overtightening, it is necessary to impose certain constraints on the force and time of the bolt tightening robot.

[0013] This patent considers both the performance and speed of tracking error convergence for a bolt-working robot. A series of tunnel-type preset performance functions are designed to address error convergence. These functions ensure a tighter convergence set for the bolt-tightening tracking error, resulting in better control effectiveness. The function design also considers the globality and latency of the error. This means that regardless of the initial error or when the constraints are applied, the tracking error is constrained, ensuring that the function meets the control requirements. To address the error convergence speed, a composite learning fixed-time sliding mode controller is designed using compound learning control. This controller eliminates the continuous excitation condition required by traditional controllers, enhancing the feasibility of the solution. A nonlinear fast integral terminal sliding mode control block is also designed. This module not only ensures fixed-time convergence speed but also eliminates the singularity problem associated with traditional sliding mode control blocks. To address the robot's safety and optimization issues, this patent combines preset performance control with reinforcement learning, utilizing an actor-critic framework to obtain solutions to the HJB equation. This approach ensures the optimality of the solution while ensuring the operational safety of the bolt-working robot, improving the robustness of the robot system and conserving computational resources.

[0014] In order to achieve the above object, the present invention provides the following technical solutions:

[0015] The fixed time safety tracking control method of the bolt operation robot under multi-learning tunnel performance includes the following steps:

[0016] Step 1: Establish a typical n-link bolting robot system model based on the Euler-Lagrange equation, select appropriate parameters as state variables, and obtain the state space equation of the bolting robot system;

[0017] Step 2: Design a series of important transfer functions to introduce the tunnel-type preset performance function into the tracking error constraint control to achieve a global control effect that is independent of the initial value of the tracking error;

[0018] Step 3: Use compound learning control to eliminate the strong continuous excitation condition required in the controller design process and realize the weak excitation condition of interval excitation;

[0019] Step 4: Combining composite learning and sliding mode control methods, a novel nonlinear fast integral terminal sliding mode fixed-time convergence controller is designed;

[0020] Step 5: Combine the above control scheme with the Actor-Critic architecture of reinforcement learning to form a new safe reinforcement learning method based on a preset performance function, achieving the safety of the robot's tracking performance and the optimality of the control scheme.

[0021] Furthermore, in step 2, important conversion functions include a scaling function and an error conversion function.

[0022] Furthermore, the scaling function is used to achieve unconstrainedness on the initial value of the error, and the error transfer function is used to introduce a global control effect that is independent of the initial value of the tracking error into the tracking error constrained control.

[0023] Furthermore, in step three, the composite learning control uses the time interval integral of the filtered regressor to construct a prediction error, and uses the prediction error and the tracking error to update the parameter estimation, thereby obtaining an accurate and smooth parameter estimation.

[0024] Furthermore, in step 4, the nonlinear fast integral terminal sliding mode controller includes a sliding module and an auxiliary variable. The design of the sliding module and the auxiliary variable can avoid the singularity problem and achieve fixed-time convergence of the tracking error.

[0025] Furthermore, in step five, the secure reinforcement learning method includes defining a performance indicator function, constructing a Hamiltonian function, establishing an HJB equation, using a neural network to approximate the optimal solution of the HJB equation, and designing an update law for the Actor-Critic network.

[0026] Furthermore, the defining of the performance indicator function includes defining a cost function and a Hamiltonian function.

[0027] Furthermore, constructing the Hamiltonian function includes adding the cost function and the Hamiltonian function.

[0028] Furthermore, establishing the HJB equation includes bringing the optimal control strategy and the optimal performance index function into the Hamiltonian function, and obtaining the partial derivative of the optimal control strategy.

[0029] Furthermore, the update law of the actor-critic network is designed by constructing a positive definite function using the square of the Bellman residual, and obtaining the update law by calculating the negative gradient of the positive definite function.

[0030] The beneficial effects of the present invention are:

[0031] (1) This design considers the problem of delayed tracking error constraints faced by bolt-tightening robots during the bolt-tightening process, which requires neither “overtightening” nor “undertightening”. Combined with the tunnel-type preset performance function, a tighter convergence set is provided for the error constraints. Regardless of the initial error and when the process starts, parameter convergence performance can be achieved.

[0032] (2) Aiming at the continuous excitation conditions required by traditional controller design, a composite learning method is used to construct a tracking error and prediction error to jointly drive the parameter adaptive law update, thus realizing the weak excitation conditions of interval excitation, which provides assistance for the implementation of the controller;

[0033] (3) To address the singularity problem existing in traditional terminal sliding mode control, the present invention designs a novel nonlinear fast integral terminal sliding mode control block. The controller designed based on this sliding module can achieve fixed-time convergence of the tracking error and eliminate the singularity problem encountered during the execution of the controller.

[0034] (4) This paper innovatively combines a preset performance function with the optimal backstepping method to design a new type of safe reinforcement learning algorithm. While ensuring the constraint performance of the robot system's tracking error, it also achieves the optimization of the control scheme.

[0035] (5) In order to solve the problem that the update law of the traditional reinforcement learning Actor-Critic network is relatively complex, the present invention uses the square of the Bellman residual to construct a new positive definite function. The positive definite function is equivalent to the HJB equation. The update law designed based on the positive definite function has simpler calculations, which is conducive to the real-time control of the control scheme and relaxes the continuous incentive conditions.

[0036] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0038] Figure 1 This is the technical roadmap of the present invention;

[0039] Figure 2 Preset performance fixed-time secure reinforcement learning sliding mode controller for tunnels. DETAILED DESCRIPTION

[0040] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0041] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0042] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0043] This paper considers both the performance and speed of tracking error convergence in a bolting robot. Under the premise that the bolting robot system has error delay control requirements, a tunnel-type preset performance fixed-time composite learning sliding mode control method is designed, which is independent of the initial error conditions. This method achieves tracking error constraints within a finite time. The introduction of the composite learning method eliminates the continuous excitation condition required by traditional control methods, and only requires interval excitation to implement the control method. The introduction of the tunnel-type preset function eliminates the initial error requirement and also brings more compact constraint performance.

[0044] After achieving convergence within a finite time for the bolting robot system, an actor-critic learning network based on reinforcement learning was introduced to ensure the safety of the robot system and optimize its control scheme. By combining a preset performance control method with reinforcement learning optimization, the HJB equation for the robot system was derived using a cost function and the Hamiltonian equation. Due to the inherent strong coupling nonlinearity of the HJB equation, the optimal control scheme cannot be directly derived. Using the actor-critic learning network, the optimal solution of the HJB equation was continuously approximated, thereby optimizing the system's control strategy while ensuring the tracking safety of the robot system.

[0045] See also Figure 1 and Figure 2 The present invention is designed to consist of the following steps:

[0046] (1) Based on the Euler-Lagrange equation, a typical n-link bolting robot system model is established. Appropriate parameters are selected as state variables to obtain the state space equation of the bolting robot system. From the state space equation, it can be seen that the robot system is a nonlinear strict feedback system. Therefore, in the subsequent controller design process, classic controller design methods such as backstepping can be used to design the control law, laying the foundation for subsequent controller design and the research and establishment of preset performance control.

[0047] (2) A series of important transfer functions are designed to introduce tunnel-type preset performance functions into tracking error constraint control, achieving a global control effect that is independent of the initial value of the tracking error. Composite learning control is used to eliminate the strong continuous excitation conditions required in the controller design process and realize the weak excitation conditions of interval excitation. Combining composite learning and sliding mode control methods, a new nonlinear fast integral terminal sliding mode fixed-time convergence controller is designed.

[0048] (3) By introducing the Actor-Critic architecture of reinforcement learning and combining it with preset performance control, a new type of safe reinforcement learning optimization control method is designed to achieve the safety of the robot system and the optimality of the control scheme.

[0049] To achieve the above object, the technical solution of the present invention is as follows:

[0050] The bolt-working robot preset safety fixed-time tracking error convergence scheme designed by the present invention mainly consists of the following parts: composite learning uses the time interval integration of the filtered regressor to construct the prediction error, and uses the prediction error and tracking error to update the parameter estimation, thereby obtaining accurate and smooth parameter estimation, eliminating the influence of the continuous excitation condition on the controller implementation, and only requiring the interval excitation condition to achieve the convergence of the tracking error; the tunnel-type preset performance function provides a new performance function for the constraint of the tracking error, which has better convergence performance for the error; the nonlinear fast integral terminal sliding mode controller realizes the fixed-time convergence of the parameters while avoiding the singularity problem; the above control scheme is combined with the Actor-Critic architecture of reinforcement learning to form a new safety reinforcement learning method based on the preset performance function, which realizes the safety of the robot tracking performance and the optimality of the control scheme.

[0051] Part I Model establishment and system description

[0052] AModel establishment

[0053] Consider a type of n-link bolt-operating robot system, which can be expressed using the Euler-Lagrange equation as follows:

[0054]

[0055] Where q(t)=[q1(t),…q n (t)], is the angular position of the joint, M(q)∈R n×n , represents the inertia matrix, represents the centripetal torque and the Lilio torque, is the viscous friction torque, g(q)R n represents the gravitational torque, τ∈R n Represents the input control torque of the robot system.

[0056] Select appropriate state variables, let x1(t)=q, The state space equation of the bolting robot system obtained by the Euler-Lagrange equation is as follows:

[0057]

[0058] B. System Description

[0059] Both models can be used to design control laws for robotic systems, each with its own advantages. The equations of these two models will be used flexibly in the subsequent control scheme design process. To facilitate subsequent design, the following lists some important properties of the dynamic equations of the robotic system expressed in terms of the Euler-Lagrange equations:

[0060] Property 1: M(q) is a symmetric positive definite matrix. There are two constants m1 and m2 such that m1I<<M(q)< <Im2.

[0061] Nature 2 is a skew symmetric matrix. That is, the following equation holds:

[0062]

[0063] Property 2 shows that internal forces do not do work and is applicable to any type of arm robot system, that is, to bolt operation robot systems.

[0064] Property 3 The robot's dynamic system can be used with an unknown constant parameter vector θ∈R m , linearize, that is:

[0065]

[0066] where ξ is an arbitrary vector, represents the first-order derivative of ξ, Represents the dynamic regression matrix.

[0067] For the convenience of expression, the above formula can be rewritten as:

[0068]

[0069] Filter out the The filtered expression is obtained:

[0070]

[0071] The continuous excitation and interval excitation of the signal are defined as follows:

[0072] Definition 1 If (σ is the excitation length), in the interval t∈[T e -t d ,T e ] makes If it holds, then the signal Φ∈R m×n , is called an interval-excited signal.

[0073] Definition 2 If At any t>0, we have:

[0074] If it holds, then the signal Φ∈R m×n , is called a bounded continuous incentive signal.

[0075] Lemma 1 For a general dynamical system x(0)=x0, where x∈Rn , assuming that the origin is an equilibrium point of the system, if there exists a continuously differentiable, positive definite, fundamentally unbounded function V(x) that satisfies:

[0076]

[0077] Among them, α, β>0, 0<γ1<1, γ2>1. Then the stability of the system is global fixed-time stability, and the stable time T can be estimated by formula (8):

[0078]

[0079] Lemma 2 For any x i ≥0, and there exists a real number 0<γ1<1,γ2>1, the following inequality holds:

[0080]

[0081] Assumption 1: The expected tracking trajectory q d (t) and its first-order derivative are bounded.

[0082] The tracking error of the robot system is defined as follows:

[0083] e(t)=x1(t)-q d (t) (10)

[0085] Among them, q d (t) represents the expected target signal.

[0086] The parameter update error of the robot system is defined as follows:

[0087]

[0088] The control goal of this invention is to develop a safe preset performance finite time sliding mode control algorithm to achieve the following control objectives:

[0089] All error signals e(t) and Under the premise of satisfying interval excitation and non-singularity problems, it converges to 0 in a fixed time;

[0090] P2 For any initial tracking condition (including initial constraint violation), the tracking error e(t) can be adjusted to a pre-specified tunnel performance set within the user-specified stabilization time, with globally safe control performance.

[0091] Part II Description of Important Technical Means

[0092] A composite learning control

[0093] In adaptive control, parameter convergence is desirable because it improves the overall stability and robustness of the closed-loop system. However, in traditional adaptive control, ensuring parameter convergence requires a strict condition: continuous excitation. However, this strict continuous excitation condition is extremely difficult to achieve in real-world control, making the control scheme difficult to implement in practice. Compound learning control, on the other hand, utilizes the time-interval integral of the filtered regressor to construct a prediction error, eliminating the need for time-dependent state derivation. The prediction error and the filter tracking error are then used to update the parameter estimates. Global stability of the closed-loop system can be achieved under the conditions of interval excitation.

[0094] In order to relax the continuous incentive conditions, the composite learning control constructs the following variable:

[0095]

[0096] where τ d is a continuous integration time, multiply both sides of the above formula by get:

[0097]

[0098] Where Θ(t) is the excitation matrix. If the interval excitation condition is satisfied, according to Definition 1, we have a positive constant So that Θ(T e )≥σI, then, define the prediction error ε as:

[0099]

[0100] The modified incentive matrix is ​​further defined as:

[0101]

[0102] It is easy to find the following relationship:

[0103]

[0104] Define a projection operator as follows:

[0105]

[0106] According to the above theorems and contents, the parameter update law can be designed.

[0107] B tunnel preset performance

[0108] Compared to other performance functions such as funnel-type, tunnel preset performance provides a tighter set of allowable values, thus achieving better transient performance. The constraint problem considered in this invention is that the system tracking error is not subject to any constraints in the initial stage, but after a certain period of time, it is constrained to the specified tunnel preset performance:

[0109]

[0110] e l (t), e u (t) gives the upper and lower bounds of the tunnel's preset performance, and τ is the given time. The upper and lower bounds are defined as follows:

[0111]

[0112] Where 0<δ<1, e0=e(0), represents the initial value of the tracking error, ρ(t)=(ρ0-ρ ∞ )e -lt +ρ ∞ , and 0<ρ ∞ <ρ0,l>0.

[0113] The underlying mechanism of tunnel preset performance is to unify different initial conditions into a concise performance description by using the properties of symbolic functions. Most previous preset performance controls are symmetrically distributed on both sides, and due to their loose performance bounds, they cannot impose stricter constraints. However, tunnel preset performance, due to its asymmetric performance bounds and concise expression, has better finite-time convergence and flexible performance.

[0114] C Nonlinear Fast Integral Terminal Sliding Mode Control

[0115] Terminal sliding mode control is considered an important technology for achieving fixed-time convergence of tracking errors in nonlinear strict feedback systems. In traditional terminal sliding mode control, the sliding module is designed as follows:

[0116] s=ζ+λsig γ (ζ),0<γ<1 (20)

[0118] When using the traditional terminal sliding module for regression matrix adaptive robot control, the corresponding auxiliary variables should be designed as:

[0119]

[0120] However, the design of the auxiliary variables mentioned above will inevitably lead to singularity problems. To solve this problem, the present invention proposes a nonlinear fast integral terminal sliding mode control block. The sliding module and auxiliary variables are designed as follows:

[0121]

[0122] Among them, λ1, λ2, λ3, λ4 are positive constants, 0<α<1, -1<β<0,

[0123]

[0124] The above sliding mode control block can not only ensure fixed-time tracking control, but also avoid the singularity problem. The above characteristics make the proposed control strategy suitable for fixed-time parameter estimation and tracking control tasks.

[0125] Part III Design of a Fixed-Time Preset Performance Sliding Mode Controller

[0126] AGlobal tunnel type preset performance function

[0127] In order to achieve the globality and good tracking performance of tracking error constraints, several important conversion functions are first introduced to better introduce the tunnel-type preset performance function into the error-constrained controller design.

[0128] ① Scaling function

[0129] In order to achieve unconstrained initial value of error, a scaling function is defined.

[0130] Definition 1 Scaling function b f (t) have the following unique properties:

[0131] 1) When t≥0, b f (t) Any C n+1 The order is bounded;

[0132] 2) In the interval t∈[0,τ), b f (t) Strictly monotonically increasing;

[0133] 3)b f (0)=0, and when t≥τ, b f (t) = 1, for any time t, b f (t)∈[0,1], when t→τ, b f The first derivative of (t) tends to 0.

[0134] Select b f (t) The function is as follows:

[0135]

[0136] Where l is a defined positive definite constant.

[0137] In practical applications, robotic systems are subject to a wide range of initial condition constraints. If the system violates the performance constraints, it may cause singularity problems, which may lead to failure of the control design. In order to deal with this limitation, this paper proposes a scaling function and its properties. From Definition 1, we can see that b f (t) is multiplied by e(t), because b f (0) = 0, so that the initial value is always guaranteed to be 0, and thus there is a unified solution for any initial condition. Since the tracking error must be subject to performance constraints immediately after a finite time τ, for t ≥ τ, b f (t)e(t) = e(t). For time 0≤t<τ, the scaling function is required to increase smoothly from 0 to 1 and remain constant thereafter. Such a design mechanism can lead to extensive research combining various scaling functions with alternative control methods.

[0138] ② Error transfer function

[0139] For any initial tracking condition, the tracking error e(t) is scaled using the scaling function b f (t), construct the following error transfer function:

[0140]

[0141] Among them, e u (t), e l (t) is the upper and lower bounds of the tunnel preset performance.

[0142] According to the definitions of the above functions, we can see

[0143]

[0144] Therefore, no matter how much the given initial error is (including initial constraint violations), the error constraint can be guaranteed.

[0145] To ensure that the tracking error is always constrained within the pre-set tunnel performance boundary, an auxiliary variable is introduced based on the previous functions:

[0146]

[0147] The tunnel-type preset performance function that is independent of the initial value is established as:

[0148]

[0149] An important property of the zeta function is that it is bounded, ensuring that the tracking error remains within the tunnel's predefined performance. Furthermore, this tracking error function provides a unified solution for handling both symmetric and asymmetric constraints, eliminating the need to redevelop the transformation function.

[0150] Another outstanding feature of the ζ function is that no matter what the initial value and initial direction of the robot system are, the initial value of the ζ function is 0, which makes the initial control signal applicable to any initial conditions.

[0151] The proof is as follows:

[0152] For any initial error value e(0), according to the scaling function b f (t) and the definition of the auxiliary variable χ, we can get χ(0) = 0, which leads to b f (0) = 0; According to the definition of the tunnel error transfer function, the following inequality always holds:

[0153]

[0154] That is, no matter what the initial tracking error of the robot system is and when the error constraint is started, it can be guaranteed that the initial error is always limited within the preset tunnel-type performance function.

[0155] Next, as long as the ζ function is bounded, it can be achieved:

[0156]

[0157] Right now:

[0158] -e l <e s (t) <e u (32)

[0160] The proof is complete.

[0161] Initial value-independent tunnel-type preset performance functions describe an explicit relationship between the zeta function and the auxiliary variable χ, a crucial relationship for designing control solutions for delay-constrained error problems. However, the commonly used logarithmic and hyperbolic tangent transformation functions are inherently complex and coupled, making such an explicit relationship difficult to obtain.

[0162] Finding the first-order derivative of the ζ function and sorting out the related equations are:

[0163]

[0164] in,

[0165] Substituting the above formula into the dynamic expression of the robot system is:

[0166]

[0167] Multiply both sides by get:

[0168]

[0169] in:

[0170]

[0171] Through the above transformation, the robot system (1) subject to tracking error constraints is transformed into an equivalent "unconstrained" robot system (35). In the next section, a composite learning fixed-time controller will be designed. One of the tasks of the designed controller is to ensure the boundedness of ζ.

[0172] B. Design of compound learning fixed-time controller

[0173] Using the composite learning method, we construct an incentive matrix for interval incentives:

[0174]

[0175] Based on the previously designed sliding module auxiliary variables and nonlinear fast integration terminal sliding module, the update law and control law of the designed system are as follows:

[0176]

[0177] Substituting the above control law and update law into the dynamic equation, the dynamic closed-loop system of the bolting robot is obtained as follows:

[0178]

[0179] Part IV System Stability and Convergence Performance Analysis

[0180] Consider the following Lyapunov-like positive definite function:

[0181]

[0182] The function V is differentiated with respect to time, and the closed-loop expression of the bolting robot dynamic system is substituted into it. According to the properties of the robot's Euler-Lagrange equation, we can obtain:

[0183]

[0184] According to the inequality:

[0185]

[0186] Substituting the result of the inequality into:

[0187]

[0188] According to the revised incentive matrix and combined with Lemma 2, we have:

[0189]

[0190] in,

[0191]

[0192] Secondly, considering the stability on t∈[0,∞], ignoring the last two terms of inequality (45), we can obtain:

[0193]

[0194] According to the above results, we can get s and are bounded, so we can deduce that ζ and Therefore, we conclude that the sliding module converges asymptotically to zero under interval excitation.

[0195] Based on the above analysis, we can say that the system always converges before the interval excitation conditions are met. When the interval excitation conditions are met, we have:

[0196]

[0197] Combining the above results and using Lemma 2, we have:

[0198]

[0199] According to Lemma 1, we can conclude that s and It converges to 0 within a fixed time. Therefore, it can be concluded that the tracking error converges to 0 within a fixed time. Based on the boundedness of ζ, it can be inferred that the tracking error is always constrained within the preset tunnel-type preset performance function, that is, the tracking error constraint is guaranteed.

[0200] Part V: Optimizing Secure Reinforcement Learning Controllers

[0201] In the controller design process described above, we obtained a sliding mode controller with global fixed-time convergence and preset performance. However, the controller optimization problem was not considered during the design process. Traditional safety reinforcement learning controller design incorporates the barrier Lyapunov function into the controller design of the optimal backstepping method, which ensures both tracking error constraints and optimal control schemes.

[0202] However, this error constraint method using the barrier Lyapunov function has a hidden prerequisite: the initial tracking error must be constrained by the barrier Lyapunov function from the outset, and the controller design must be real-time. This design method fails when there is no tracking error constraint initially, and only after a certain period of time, because the definition of the barrier Lyapunov function has been violated.

[0203] We have proposed a tunnel-type preset performance control method that is independent of the initial error value and the time when the error constraint begins. If the preset performance control is combined with the optimal backstepping method, a new error-constrained optimization method will be created.

[0204] During implementation, the optimal backstepping method requires the use of a reinforcement learning actor-critic network to approximate the HJB equation, thereby obtaining an optimal controller. Traditionally, the update law solution is obtained by gradient descent of the square of the Bellman residual equation, which is equivalent to estimating the HJB equation and inevitably introduces many nonlinear terms. To simplify the solution of the update law, the present invention designs a simple positive definite function. The update law obtained by taking its negative gradient is simple in form, and the continuous excitation condition can also be relaxed, making it easier to use in real-time controllers.

[0205] Define the performance indicator function:

[0206]

[0207] Among them, r(ζ(τ),u(ζ))=ζ(t)Q(x)ζ(t)+u T u is the cost function, Q(x) is a positive definite matrix, and u is the system input. The Hamiltonian function is as follows:

[0208]

[0209] Let u * is the optimal control strategy, and the optimal performance index function is:

[0210]

[0211] Where Ψ(Ω) is a tight set of feasible control strategies, and (50) and u * Substituting (51) into the HJB equation, we obtain:

[0212]

[0213] Find u for (52) * The partial derivative of , we can get the optimal control strategy:

[0214]

[0215] Substituting (53) back into (52) we have:

[0216]

[0217] By solving equation (54), we can obtain Then obtain u * .

[0218] However, this equation is highly nonlinear and difficult to solve through numerical calculations. Therefore, it is necessary to introduce the reinforcement learning Actor-critic architecture approximated by neural networks to solve this problem.

[0219] V * (z) is divided into two parts:

[0220] V * (ζ)=β||ζ(t)|| 2 -β||ζ(t)|| 2 +V * (ζ)

[0221] =β||ζ(t)|| 2 +V o (ζ) (55)

[0222] Where β is a design constant, V o (ζ)=V * (ζ)-β||ζ(t)|| 2

[0223] According to the universal approximation theorem of neural networks, the continuous function V o (ζ) can be approximated by a neural network as:

[0224] V o (ζ)=W *T S(ζ)+ε(ζ) (56)

[0225] W * is the ideal m-dimensional weight vector, S(ζ) is the basis function vector, and m is the number of neurons.

[0226] From (55), we can get u approximated by neural network * ,

[0227]

[0228] However, the ideal m-dimensional weight vector W * Is an unknown value, we need to learn its estimated value through the Actor-Critic architecture of reinforcement learning:

[0229]

[0230] in, and is the estimated value obtained by reinforcement learning. Substituting (58) into the HJB equation obtained by reinforcement learning

[0231]

[0232] The Bellman residual φ(t) is defined as the difference between equation (52) and equation (59):

[0233]

[0234] Define Φ(t)=φ(t) 2 , is a positive definite function.

[0235] In order to simplify the design of the update law, a new positive definite function is considered. beg The partial derivative of , we can get:

[0236]

[0237] The above equation is equivalent to the HJB equation, that is, Equivalent to Therefore, the function after partial derivative can be used to design the update law, which will greatly simplify the calculation of the update law and improve the real-time performance of the control scheme.

[0238] Define a new positive definite function:

[0239]

[0240] It can be clearly seen that P(t)=0, and it can be inferred that Thus, it is concluded Therefore, the excellent properties of the P(t) function provide great convenience for the design of the update law.

[0241] The update law of the Actor-Critic network is designed as follows:

[0242]

[0243] Among them, γ c and γ a are the learning rates of the Critic network and the Actor network respectively.

[0244] So far obtained and According to the definition, Φ(t)>0, Therefore, it is asymptotically stable, that is, and Will approach the final true value. In this way, the W problem is solved by the reinforcement learning Actor-Critic architecture. * Unknown problems, so as to solve the optimal control strategy u * , to achieve the purpose of safe tracking and control.

[0245] Compared with the traditional update law, the update law designed in this way makes the optimized control algorithm significantly simpler, and at the same time, it can also release the strict conditions of persistent incentives. Generally, reinforcement learning for optimal control is an iterative process, and the critic and actor are trained at the same time. Therefore, the high complexity of the control design mainly comes from the construction of the critic update method and the actor update method. In the published optimization methods, both update laws are designed based on the square of the Bellman residual. Since this equation is a complex nonlinear equation, it inevitably increases the complexity of the control design. In the present invention, the reinforcement learning algorithm is designed based on a simple positive function, which is equivalent to the HJB equation, which is very meaningful for reducing the complexity of the control design.

[0246] The safety reinforcement learning control method proposed in this paper appropriately integrates a tunnel-type preset performance function into an optimization backstepping method to maintain state variables within a designed safety range. During the learning process, the variance of control performance under system uncertainty is significantly reduced. This technique rigorously considers the important factors of safety assurance and exploration, without the need for additional safety guarantees such as uncertainty estimation, auxiliary safety controllers, or a priori safety sets.

[0247] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A fixed-time safety tracking control method for a bolting robot under multi-learning tunnel performance, characterized by: The following steps are involved: Step 1: Establish a typical n-link bolting robot system model based on the Euler-Lagrange equation, select appropriate parameters as state variables, and obtain the state space equation of the bolting robot system; Step 2: Design a series of important transfer functions to introduce the tunnel-type preset performance function into the tracking error constraint control to achieve a global control effect that is independent of the initial value of the tracking error; Step 3: Use compound learning control to eliminate the strong continuous excitation condition required in the controller design process and realize the weak excitation condition of interval excitation; Step 4: Combining composite learning and sliding mode control methods, a novel nonlinear fast integral terminal sliding mode fixed-time convergence controller is designed; Step 5: Combine the above control scheme with the Actor-Critic framework of reinforcement learning to form a new safe reinforcement learning method based on a preset performance function, achieving the safety of the robot's tracking performance and the optimality of the control scheme; In the step 2, the important conversion functions include a scale function and an error transfer function; Scaling function b f (t) are as follows: Where l is a defined positive constant and τ is a finite time; The error transfer function is designed as follows: Among them, e u (t), e l (t) Preset upper and lower bounds for tunnel performance; Introduce auxiliary variables: Establish a tunnel-type preset performance function ζ that is independent of the initial value: Design the sliding module and auxiliary variables as follows: Among them, λ1, λ2, λ3, λ4 are positive constants, 0<α<1, -1<β<0, 2. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performance conditions according to claim 1 is characterized by: The scaling function is used to achieve unconstrainedness on the initial value of the error, and the error transfer function is used to introduce the global control effect that is irrelevant to the initial value of the tracking error into the tracking error constrained control.

3. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performance conditions according to claim 1 is characterized in that: In the step three, the composite learning control uses the time interval integral of the filtered regressor to construct a prediction error, and uses the prediction error and the tracking error to update the parameter estimation, thereby obtaining an accurate and smooth parameter estimation.

4. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performance conditions according to claim 1 is characterized in that: In the step 4, the nonlinear fast integral terminal sliding mode controller includes a sliding module and an auxiliary variable. The design of the sliding module and the auxiliary variable can avoid the singularity problem and achieve fixed-time convergence of the tracking error.

5. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performances according to claim 1 is characterized in that: In step five, the secure reinforcement learning method includes defining a performance indicator function, constructing a Hamiltonian function, establishing an HJB equation, using a neural network to approximate the optimal solution of the HJB equation, and designing an update law for the Actor-Critic network.

6. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performance conditions according to claim 5 is characterized by: Defining the performance indicator function includes defining a cost function and a Hamiltonian function.

7. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performance conditions according to claim 5 is characterized by: The constructing of the Hamiltonian function includes adding the cost function and the Hamiltonian function.

8. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performances according to claim 5 is characterized by: The establishment of the HJB equation includes bringing the optimal control strategy and the optimal performance index function into the Hamiltonian function, and obtaining the partial derivative of the optimal control strategy.

9. The method for secure tracking control of a bolting robot during fixed time under multiple learning tunnel performance conditions according to claim 5 is characterized by: The update law of the actor-critic network is designed by constructing a positive definite function using the square of the Bellman residual, and obtaining the update law by calculating the negative gradient of the positive definite function.

Citation Information

Patent Citations

  • Mechanical arm input saturation fixed time trajectory tracking control method and system

    CN111496792A

  • Limited mechanical arm fixed time self-adaptive parameter identification and control method and device

    CN116587279A