Control method and system for a multi-arm robot system with specified performance constraints

By designing an approximately optimal virtual controller based on a global performance function and a cost function, and combining dynamic surface techniques and nonlinear filters, the problems of poor control performance and high control cost of multi-arm manipulator systems in complex environments are solved, and the stability, performance constraints and robustness are improved.

CN116968029BActive Publication Date: 2026-02-24SHENZHEN INSTITUTE OF INFORMATION TECHNOLOGY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311016991.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2026-02-24
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

Multi-arm robotic systems are susceptible to external influences in complex environments, resulting in poor control performance. At the same time, designing a distributed optimal controller is difficult, making it hard to reduce control costs while meeting stability and specified control performance requirements.

Method used

By establishing a state dynamics model, an approximately optimal virtual controller based on a global performance function and a cost function is designed. By combining dynamic surface techniques and nonlinear filters, the controller design is optimized to reduce control costs. Furthermore, reinforcement learning algorithms are introduced to improve robustness and control performance.

Benefits of technology

This approach ensures the stability and specified control performance of multi-arm robotic systems in complex environments, while reducing control costs, improving system robustness and control effectiveness, broadening applicability, and avoiding the computational explosion problem in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116968029B_ABST
    Figure CN116968029B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of industrial process control, and provides a control method and system of a multi-single-arm manipulator system with specified performance constraints. When a consistency error is obtained, coordinate transformation is performed according to a consistency error model of the single-arm manipulator and a constructed global performance function, and the global performance function is added, so that the performance requirements are met, the problem of poor control performance caused by the influence of the external environment due to the complexity of the working environment is avoided, and the control effect is guaranteed. On this basis, a cost function and an optimal cost function are designed based on the consistency error and the optimization target, and an approximate optimal controller is designed based on the cost function, so that the purpose of reducing the control cost is achieved. In the control process, the global performance function and the reinforcement learning algorithm are added, so that the specified performance constraints are met and the control cost is reduced. On the basis of meeting the stability of the multi-single-arm manipulator system and meeting the specified control performance, the purpose of reducing the control cost is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of industrial process control technology, and in particular relates to a control method and system for a multi-arm robotic system with specified performance constraints. Background Technology

[0002] Single-arm robotic systems are widely used in civilian and military fields due to their wide applicability, high safety, and easy expansion, replacing humans in performing tasks in complex or hazardous environments. Compared to a single single-arm robotic system, a multi-single-arm robotic system, because it comprises multiple single-arm robotic systems, can be used to handle or complete more complex industrial tasks.

[0003] The inventors discovered that in the control process of a multi-arm manipulator system, not only is system stability desirable, but certain performance requirements are also in place, such as overshoot, convergence speed, and convergence range. However, due to the complexity of the working environment, multi-arm manipulator systems are easily affected by external factors, which may result in deficiencies in control performance. Besides ensuring system stability and meeting specified control performance requirements, there are also certain requirements regarding the control cost; often, it is desirable to achieve better control results with less control cost. However, since the performance of multi-arm manipulator systems needs to be within a certain range, designing a distributed optimal controller presents certain challenges. Summary of the Invention

[0004] To address the aforementioned problems, this invention proposes a control method and system for a multi-arm manipulator system with specified performance constraints. While ensuring system consistency, this invention reduces control costs by adding a global performance function and a design cost function, thereby guaranteeing the stability of the multi-arm manipulator system and meeting the specified control performance.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solution:

[0006] In a first aspect, the present invention provides a control method for a multi-arm robotic system with specified performance constraints, comprising:

[0007] Establish a state-dynamic model for a multi-single-arm robotic arm;

[0008] Based on the established state dynamics model, coordinate transformation is performed according to the consistency error model of the single-arm manipulator and the constructed global performance function to obtain the transformed consistency error;

[0009] Based on the obtained consistency error, an approximate optimal virtual controller based on reinforcement learning is designed; this includes designing the cost function and the optimal cost function based on the consistency error and the optimization objective.

[0010] Based on the approximate optimal virtual controller, new state variables are obtained by introducing dynamic surface techniques and nonlinear filters;

[0011] Based on the new state variables, an error transformation is designed, and an approximate optimal controller is obtained based on the error transformation and reinforcement learning algorithm.

[0012] A near-optimal controller is used to control a multi-arm robotic system.

[0013] Furthermore, a communication topology diagram describing the communication relationships between multiple single-arm manipulators is established. Based on the communication topology diagram, considering the unknown nonlinear effects on the multiple single-arm manipulator system, the motion model of the multiple single-arm manipulator system is transformed into a state dynamics model.

[0014] Furthermore, the global performance function is:

[0015]

[0016] Where, χ i Let i be a positive constant; i is the i-th single-arm robotic arm. P i It is a monotonically increasing function value.

[0017] Furthermore, the cost function is:

[0018]

[0019] in, z i,1 For conversion error; α i,1 This is a virtual control signal; t is a time constant.

[0020] When the cost function is optimal, we can obtain:

[0021]

[0022] in, For optimal virtual controller, It is a compact set that includes the origin.

[0023] Furthermore, the nonlinear filter is:

[0024]

[0025] Among them, ι i and ρ i These are positive design parameters; i represents the i-th single-arm robotic arm; For filtering error, a i,c The output signal of the filter. It is an approximately optimal virtual controller; For αi,c The derivative; ζ i express The upper bound; This is an estimated value.

[0026] Furthermore, an error transformation is designed based on the new state variables, and an approximately optimal controller is obtained based on the error transformation, prediction model, reinforcement learning algorithm, and neural network.

[0027] Furthermore, the approximate optimal controller is:

[0028]

[0029] Among them, M i Represents the moment of inertia; β i,2 For positive design parameters, z i,2 This is a virtual error. To execute neural network weights; S i,2 are basis functions; For α i,c The derivative; for The estimated value of S; i,f are basis functions; This is an estimate of the composite disturbance.

[0030] The second transformation module is configured to: design an error transformation based on the new state variables, and obtain an approximate optimal controller based on the error transformation and reinforcement learning algorithm;

[0031] The control module is configured to control the multi-arm robotic system using an approximately optimal controller.

[0032] Secondly, the present invention also provides a control system for a multi-arm robotic system with specified performance constraints, comprising:

[0033] The state dynamics model building module is configured to: build state dynamics models for multiple single-arm manipulators;

[0034] The transformation module is configured to: perform coordinate transformation based on the established state dynamics model, the consistency error model of the single-arm manipulator, and the constructed global performance function to obtain the transformed consistency error;

[0035] The first transformation module is configured to: design an approximate optimal virtual controller based on reinforcement learning according to the obtained consistency error; wherein, based on the consistency error and the optimization objective, design a cost function and an optimal cost function;

[0036] The state variable calculation module is configured to obtain new state variables based on the approximate optimal virtual controller by introducing dynamic surface techniques and nonlinear filters.

[0037] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the control method for a multi-arm robotic system with specified performance constraints as described in the first aspect.

[0038] Fourthly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the control method for a multi-arm robotic system with specified performance constraints as described in the first aspect.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] 1. When obtaining the transformed consistency error, this invention performs coordinate transformation based on the consistency error model of the single-arm manipulator and the constructed global performance function. The addition of the global performance function satisfies the performance requirements, avoids the problem of being affected by the external environment due to the complexity of the working environment, and ensures the control effect. On this basis, an approximately optimal virtual controller is designed to ensure stability, meet performance constraints, and reduce control costs. In other words, while satisfying the stability and specified control performance of the multi-single-arm manipulator system, the goal of reducing control costs is achieved.

[0041] 2. This invention introduces a prediction error to construct an identification model, enabling the neural network system to accurately approximate the unknown nonlinear function of the system. This results in a more robust controller and better control performance based on the identification model. By obtaining an error variation model based on the consistency error of the single-arm manipulator and the constructed global performance function, the consistency error is made to meet the constraints, thus ensuring that multiple single-arm manipulators can effectively track the given signal. Furthermore, this method has wider applicability and does not require additional requirements for the initial error. By introducing a simplified reinforcement learning algorithm in each step of the backstepping method, the optimality of the system is ensured while simplifying the control design, achieving a trade-off between control performance and control cost. The control signal of the virtual controller is filtered, and nonlinear dynamic surface technology is introduced to solve the "computational explosion" problem of the traditional backstepping method and eliminate the influence of filtering errors, while relaxing the requirements for filter parameters. Attached Figure Description

[0042] The accompanying drawings, which form part of this embodiment, are used to provide a further understanding of this embodiment. The illustrative embodiments and their descriptions are used to explain this embodiment and do not constitute an improper limitation of this embodiment.

[0043] Figure 1 This is a flowchart of Embodiment 1 of the present invention;

[0044] Figure 2 This is a communication topology diagram between single-arm robotic arms in Embodiment 1 of the present invention;

[0045] Figure 3 This is a consistency effect diagram of the simulation test of Embodiment 1 of the present invention;

[0046] Figure 4 The simulation test state x of Embodiment 1 of the present invention i,2 Line graph;

[0047] Figure 5 This is a convergence curve of the consistency error in the simulation experiment of Embodiment 1 of the present invention;

[0048] Figure 6 The unknown nonlinear function and time-varying disturbance f of the simulation test system in Embodiment 1 of the present invention i +d i A schematic diagram of the estimation results. Detailed Implementation

[0049] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0050] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0051] Example 1:

[0052] When ensuring the stability and meeting specified control performance requirements of a multi-arm manipulator system, certain requirements are also placed on the control cost. Often, it is desirable to achieve better control results with less control cost. However, since the performance of a multi-arm manipulator system needs to be within a certain range, designing a distributed optimal controller presents certain difficulties. To resolve the conflict between stability and meeting specified control performance requirements and control cost, this embodiment provides a control method for a multi-arm manipulator system with specified performance constraints, including:

[0053] Establish a state-dynamic model for a multi-single-arm robotic arm;

[0054] Based on the established state dynamics model, coordinate transformation is performed according to the consistency error model of the single-arm manipulator and the constructed global performance function to obtain the transformed consistency error;

[0055] Based on the obtained consistency error, an approximate optimal virtual controller based on reinforcement learning is designed; this includes designing the cost function and the optimal cost function based on the consistency error and the optimization objective.

[0056] Based on the approximate optimal virtual controller, new state variables are obtained by introducing dynamic surface techniques and nonlinear filters;

[0057] Based on the new state variables, an error transformation is designed, and an approximate optimal controller is obtained based on the error transformation and reinforcement learning algorithm.

[0058] A near-optimal controller is used to control a multi-arm robotic system.

[0059] Specifically, when obtaining the transformed consistency error, coordinate transformation is performed based on the consistency error model of the single-arm manipulator and the constructed global performance function. The addition of the global performance function satisfies the performance requirements, avoids the problem of being affected by the external environment due to the complexity of the working environment, and ensures the control effect. On this basis, an approximately optimal virtual controller and an approximately optimal controller are designed to ensure stability, meet performance constraints, and reduce control costs. In other words, while satisfying the stability and specified control performance of the multi-single-arm manipulator system, the goal of reducing control costs is achieved. The multi-single-arm manipulator system includes one leader manipulator and N follower manipulators. The follower manipulators are single-arm manipulators with unknown nonlinearity and time-varying disturbances, such as... Figure 1 The specific steps in this embodiment are as follows:

[0060] S1. Establish a state dynamics model for multiple single-arm manipulators.

[0061] Optionally, the single-arm robotic arm system model as a follower robotic arm has been widely recognized in related fields. The system model of the i-th single-arm robotic arm in the follower robotic arm is as follows:

[0062]

[0063] Where, q i , and These represent the angular position, angular velocity, and angular acceleration of the connecting rod, respectively; M i Indicates moment of inertia; m i The mass of the connecting rod is represented by g; g represents the acceleration due to gravity; l i and B i These represent the link length and damping coefficient, respectively; τ i This indicates the input for single robotic arm control.

[0064] The steps for establishing a state dynamics model are as follows:

[0065] S1.1 Establish graph theory knowledge to describe the communication relationships between multiple single-arm robotic arms:

[0066] like Figure 2The diagram illustrates the communication relationships between robotic arms in a topological graph. This embodiment incorporates relevant knowledge of algebraic graph theory. Considering there are N robotic arm systems in the network, using... Let represent a directed communication topology in graph theory, where Represents a vertex set. Let B represent the edge set, where B = [b i,j ]∈R N×N Let b be the adjacency matrix. (j,i)∈ξ indicates that the j-th robotic arm system can transmit information to the i-th robot. When (j,i)∈ξ, b i,j =1, otherwise b i,j =0. If the leader can transmit information to agent i, then b i,0 =1, otherwise b i,0 =0.

[0067] S1.2. Considering that multi-arm robotic systems in complex environments may be affected by unknown nonlinearities and external disturbances, let... and The motion model of a multi-arm robotic system can be transformed into the following state-space equations:

[0068]

[0069] in, y i Indicates the output signal; d i This indicates an external disturbance.

[0070] S2. Based on the consistency error model of the single-arm robot and the constructed global performance function, perform coordinate transformation to obtain the transformed consistency error.

[0071] S2.1 Combining graph theory and system models, the following consistency error is defined:

[0072]

[0073] Among them, y r A reference trajectory for leaders.

[0074] S2.2 Design the global performance function:

[0075]

[0076] Where, χ i η is a positive constant. i ∈(-1,1). Choose a monotonically increasing function P. i (t) makes P i (0)=P i,0 =1 and P i (∞)=P i,∞ >1. Let achievable and Pick

[0077] S2.3, Perform the following coordinate transformation:

[0078]

[0079] in, The conversion error is:

[0080] S3. Based on the changed consistency error obtained in step S2, design an approximate optimal virtual controller based on reinforcement learning.

[0081] S3.1. Based on the consistency error and optimization objective, design the cost function and the optimal cost function:

[0082]

[0083] in, z i,1 For conversion error; α i,1 This is a virtual control signal; t is a time constant.

[0084] When the cost function is optimal, we can obtain

[0085]

[0086] in, For optimal virtual controller, It is a compact set that includes the origin.

[0087] S3.2, Order The distributed Hamilton-Jacobi-Bellman equations are derived as follows:

[0088]

[0089] in, π i,1 =ω i P i m i ,

[0090] S3.3, will Decomposed into the following form:

[0091]

[0092] in,

[0093] S3.4, By the Bellman optimality principle, through The optimal virtual control signal can be obtained as follows:

[0094]

[0095] S3.5, due to It exhibits strong nonlinearity, therefore a radial basis function neural network is used to approximate it as follows:

[0096]

[0097] in, Let S be the ideal weight vector. i,1 ε is a basis function. i,1 To approximate the residual difference, satisfy |ε i,1 |≤ε i,1 * , ε i,1 * Let be an arbitrarily small constant.

[0098] Therefore, the optimal virtual controller It can be written in the following form

[0099]

[0100] S3.6, due to the optimal weight vector Since the optimal virtual controller is unknown, it is not available. Therefore, a reinforcement learning strategy with an execution-evaluation structure is adopted to obtain an approximate optimal function and an approximate optimal virtual controller:

[0101]

[0102]

[0103] in, To evaluate the network, a function is used to obtain the approximate optimal function. To evaluate the network estimated weights. To execute the network, used to obtain an approximately optimal virtual controller. To perform network weight estimation.

[0104] S3.7, The weight update rate for network estimation is:

[0105]

[0106] Among them, Y i,a1 To execute network design parameters and satisfy

[0107] S3.8, The evaluation network estimated weight update rate is:

[0108]

[0109] Among them, Y i,c1 To evaluate network design parameters and satisfy Y i,c1 >Y i,a1 .

[0110] S4. Based on the approximate optimal virtual controller obtained in step S3, this embodiment introduces dynamic surface technology and designs nonlinear filters to avoid the "computational explosion" problem inherent in the backstepping method, while relaxing the requirements for filter parameters.

[0111] S4.1 Design the following nonlinear filter:

[0112]

[0113] Among them, ι i and ρ i These are positive design parameters. This represents the filtering error. ζ i express The upper realm, Its estimated value.

[0114] S4.2, Design the following update rate:

[0115]

[0116] Among them, κ i,1 and κ i,2 These are design parameters.

[0117] S5. Based on the new state variables obtained in step S4, design an error transformation, and obtain an approximate optimal controller based on the error transformation, prediction model, reinforcement learning algorithm and neural network technology.

[0118] S5.1 Based on the new state variables, design the following coordinate transformation:

[0119] z i,2 =x i,2 -a i,c

[0120] S5.2. Based on coordinate transformation and control objectives, design the cost function and the optimal cost function:

[0121]

[0122] in,

[0123] When the cost function is optimal, we can obtain

[0124]

[0125] in, For the optimal controller, Ω zi,2 It is a compact set that includes the origin.

[0126] S5.3 The distributed Hamilton-Jacobi-Bellman equation is derived as follows:

[0127]

[0128] S5.4, will Decomposed into the following form:

[0129]

[0130] in,

[0131] S5.5, Based on the Bellman optimality principle, through The optimal control signal can be obtained as follows:

[0132]

[0133] S5.6, due to and f i It exhibits strong nonlinearity, therefore a radial basis function neural network is used to approximate it as follows:

[0134]

[0135]

[0136] in, and Let S be the ideal weight vector. i,2 and S i,f D is a basis function. i =d i +ε i,f ε i,2 and ε i,f To approximate the residual difference, satisfy |ε i,2 |≤ε i,2 * and |ε i,f |≤ε i,f * ε i,2 * and ε i,f * Let be an arbitrarily small constant.

[0137] Therefore, the optimal controller It can be written in the following form:

[0138]

[0139] S5.7 Design prediction error:

[0140]

[0141] in, For x i,2 The estimated value.

[0142] Its update rate is designed as follows:

[0143]

[0144] in, and They are respectively and D i The estimated value, h i These are positive design parameters.

[0145] By introducing prediction error, the following neural network update law and perturbation observer are designed:

[0146]

[0147]

[0148]

[0149] S5.8, due to the optimal weight vector Since the optimal controller is unknown, it is not available. Therefore, a reinforcement learning strategy with an execution-evaluation structure is adopted to obtain an approximate optimal function and an approximate optimal controller.

[0150]

[0151]

[0152] in, To evaluate the network, a function is used to obtain the approximate optimal function. To evaluate the network estimated weights. To implement the network, used to obtain an approximate optimal controller. To perform network weight estimation.

[0153] S5.9, The weight update rate for network estimation is:

[0154]

[0155] Among them, Y i,a2 To execute network design parameters and satisfy

[0156] S5.10, The evaluation network estimated weight update rate is:

[0157]

[0158] Among them, Y i,c2 To evaluate network design parameters and satisfy Y i,c2 >Y i,a2 .

[0159] In a simulation experiment of this embodiment:

[0160] The control objective of the simulation experiment is to achieve the angular position q of the robotic arm. i Tracking the given trajectory signal y r =sin(t), and in the second subsystem, the influence of external disturbances is considered, with the disturbance signal being d. i = 2cos(t). Considering the actual system, the system parameters are designed as follows: Moment of inertia M i =1kg / m 2 The acceleration due to gravity is g = 9.8 m / s². 2 The mass, length, and damping coefficient of the connecting rod are respectively m i =1kg,l i =1m, B i =1.

[0161] The initial values ​​for the simulation system state and the neural network weights are: x i,1 (0) = [0.1, 0.1, 0.1, 0.1], x i,2 (0) = [0.1, 0.1, 0.1, 0.1], r i,2 (0) = [0.1, 0.1, 0.1, 0.1], a i,c (0) = [0.1, 0.1, 0.1, 0.1],

[0162] Other parameters are designed as follows: β i,1 =β i,2 =[20,20,20,20], Y i,a1 =Y i,a2 =[4,4,4,4],Y i,c1 =Y i,c2 =[20,20,20,20]. h i =[20,20,20,20],Γ i =[5,5,5,5], r i,f =[2,2,2,2],δ i =[20,20,20,20], L i =[10,10,10,10].

[0163] Simulation results are as follows Figure 3-6 As shown. Figure 3 The diagram shows the consistency effect of multiple single-arm robotic arms. It can be seen that all the single-arm robotic arms can track the given leader signal, thus achieving the goal of consistent control.

[0164] Figure 4 For state x i,2 The trajectory diagram shows that state x... i,2 It is bounded. Figure 5 The convergence curve of the consistency error shows that after applying the improved global specification performance technique, the consistency error of each robot can converge to the predetermined range, and the proposed method does not have too many requirements on the initial value of the consistency error. Figure 6 The graph shows the estimation of the unknown nonlinear function and external disturbance signal of the robot. It can be seen that after introducing the prediction error design identification network and the adaptive law of the disturbance observer, a good estimation effect can be obtained.

[0165] Example 2:

[0166] This embodiment provides a control system for a multi-arm robotic arm system with specified performance constraints, including:

[0167] The state dynamics model building module is configured to: build state dynamics models for multiple single-arm manipulators;

[0168] The transformation module is configured to: perform coordinate transformation based on the established state dynamics model, the consistency error model of the single-arm manipulator, and the constructed global performance function to obtain the transformed consistency error;

[0169] The first transformation module is configured to: design an approximate optimal virtual controller based on reinforcement learning according to the obtained consistency error; wherein, based on the consistency error and the optimization objective, design a cost function and an optimal cost function;

[0170] The state variable calculation module is configured to obtain new state variables based on the approximate optimal virtual controller by introducing dynamic surface techniques and nonlinear filters.

[0171] The operating method of the system is the same as the control method of the multi-arm manipulator system with specified performance constraints in Embodiment 1, and will not be repeated here.

[0172] Example 3:

[0173] This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the control method for a multi-arm robotic system with specified performance constraints as described in Embodiment 1.

[0174] Example 4:

[0175] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the control method for a multi-arm robotic system with specified performance constraints as described in Embodiment 1.

[0176] The above description is merely a preferred embodiment of this practice and is not intended to limit the scope of this practice. Various modifications and variations can be made to this practice by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this practice should be included within the protection scope of this practice.

Claims

1. A control method for a multi-arm robotic system with specified performance constraints, characterized in that, include: Establish a state-dynamic model for a multi-single-arm robotic arm; Based on the established state dynamics model, coordinate transformation is performed according to the consistency error model of the single-arm manipulator and the constructed global performance function to obtain the transformed consistency error; Based on the obtained consistency error, an approximate optimal virtual controller based on reinforcement learning is designed; this includes designing the cost function and the optimal cost function based on the consistency error and the optimization objective. Based on the approximate optimal virtual controller, new state variables are obtained by introducing dynamic surface techniques and nonlinear filters; Based on the new state variables, an error transformation is designed, and an approximate optimal controller is obtained based on the error transformation and reinforcement learning algorithm. A near-optimal controller is used to control a multi-arm robotic system; The global performance function is: in, It is a positive number; i For the first A single-arm robotic arm; , The value is a monotonically increasing function. Cost function: in, , This is the conversion error; For virtual control signals; t It is a time constant; i For the first A single-arm robotic arm; When the cost function is optimal, we can obtain: in, For optimal virtual controller, It is a compact set that includes the origin.

2. The control method for a multi-arm robotic system with specified performance constraints as described in claim 1, characterized in that, Establish a communication topology diagram describing the communication relationships between multiple single-arm robotic arms; Based on the communication topology diagram, considering the unknown nonlinear effects on the multi-arm manipulator system, the motion model of the multi-arm manipulator system is transformed into a state dynamics model.

3. The control method for a multi-arm robotic system with specified performance constraints as described in claim 1, characterized in that, The nonlinear filter is: in, and These are positive design parameters; i For the first A single-arm robotic arm; For filtering error, For an approximately optimal virtual controller; for The derivative; express The upper bound; This is an estimated value.

4. The control method for a multi-arm robotic system with specified performance constraints as described in claim 1, characterized in that, Based on the new state variables, an error transformation is designed, and an approximate optimal controller is obtained based on the error transformation, prediction model, reinforcement learning algorithm, and neural network.

5. The control method for a multi-arm robotic system with specified performance constraints as described in claim 1, characterized in that, The approximate optimal controller is: in, Indicates the moment of inertia; For positive design parameters, This is a virtual error. To execute neural network weights; are basis functions; for The derivative; for The estimated value; are basis functions; This is the estimated value of the composite disturbance.

6. A control system for a multi-arm manipulator system with specified performance constraints, employing the control method for a multi-arm manipulator system with specified performance constraints as described in any one of claims 1-5, characterized in that, include: The state dynamics model building module is configured to: build state dynamics models for multiple single-arm manipulators; The transformation module is configured to: perform coordinate transformation based on the established state dynamics model, the consistency error model of the single-arm manipulator, and the constructed global performance function to obtain the transformed consistency error; The first transformation module is configured to: design an approximate optimal virtual controller based on reinforcement learning according to the obtained consistency error; wherein, based on the consistency error and the optimization objective, design a cost function and an optimal cost function; The state variable calculation module is configured to obtain new state variables based on the approximate optimal virtual controller by introducing dynamic surface techniques and nonlinear filters. The second transformation module is configured to: design an error transformation based on the new state variables, and obtain an approximate optimal controller based on the error transformation and reinforcement learning algorithm; The control module is configured to control the multi-arm robotic system using an approximately optimal controller.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps of the control method for a multi-arm robotic system with specified performance constraints as described in any one of claims 1-5.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the control method for a multi-arm robotic system with specified performance constraints as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-single-arm manipulator system control method and system

    CN113814983A