Robotic system intelligent control method and system based on multi-player reinforcement learning

By linearizing and iterating the policy of the robot system using a multi-player reinforcement learning method, the stability and robustness problems of the robot system in high-degree-of-freedom and multi-actuator control are solved, and the control policy solution is solved quickly and efficiently.

CN122185169APending Publication Date: 2026-06-12SHUNDE INNOVATION SCHOOL UNIVERSITY OF SCIENCE & TECHNOLOGY BEIJING
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHUNDE INNOVATION SCHOOL UNIVERSITY OF SCIENCE & TECHNOLOGY BEIJING
Filing Date
2026-03-05
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing robot systems suffer from stability, robustness, and real-time issues in high-degree-of-freedom and multi-actuator control, making them particularly difficult to control effectively under different working conditions and task requirements.

Method used

A multi-player reinforcement learning approach is adopted, which linearizes the nonlinear robot system through first-order Taylor expansion and forward Euler method, divides players into reachable and unreachable, constructs a cost function, and uses an improved policy iteration method to solve for the optimal control policy.

Benefits of technology

This improved the stability and robustness of the robot system, reduced the number of control iterations and data sampling costs, and increased the solution speed and consistency of the control strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122185169A_ABST
    Figure CN122185169A_ABST
Patent Text Reader

Abstract

The application provides a kind of robot system intelligent control method and system based on multi-player reinforcement learning, comprising: first-order Taylor expansion is carried out to continuous time nonlinear robot system, continuous time linearization robot system is obtained, and each control input is regarded as a player, forming continuous time multi-player decision system;Forward Euler method is used for discretization and substituted into the decision system to obtain a discrete time linear system model;Other players are divided into: unreachable players, adjacent reachable players and non-adjacent reachable players, and different control strategies are set for different players;A cost function is constructed for each player, and an optimal control objective for each player is established;An improved policy iteration method is used to separate the strategy data collection from the strategy evaluation / policy improvement, iteratively solve the unknown cost function, obtain the optimal control strategy, and use the optimal control strategy to control the robot system.The application can intelligently control the robot system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot control technology, and specifically refers to an intelligent control method and system for robot systems based on multiplayer reinforcement learning. Background Technology

[0002] With the widespread application of robotics in industrial manufacturing, collaborative assembly, service robots, and special operations, robot systems are developing towards higher degrees of freedom, more actuators, and greater collaboration. Multi-degree-of-freedom robotic arms, collaborative robots, and mobile manipulators typically exhibit significant nonlinear dynamic characteristics, and their multiple control input channels have complex coupling relationships. In engineering practice, these systems often need to operate under different working conditions and task requirements, placing higher demands on the stability, robustness, and real-time performance of control algorithms. Summary of the Invention

[0003] To address the technical problems existing in the prior art, the present invention provides an intelligent control method and system for robot systems based on multi-player reinforcement learning, the technical solution of which is as follows: On the one hand, a method for intelligent control of robot systems based on multi-player reinforcement learning is provided, the method comprising: S1. Select the desired equilibrium point of the robot system, and perform a first-order Taylor expansion on the continuous-time nonlinear robot system at the desired equilibrium point to obtain the continuous-time linearized robot system. S2. Each control input channel of the continuous-time linearized robot system is regarded as an independent player. Each player acts on the same continuous linear system through its input, forming a continuous-time multi-player decision-making system. S3. The continuous-time linearized robot system is discretized using the forward Euler method and substituted into the continuous-time multiplayer decision system to obtain a discrete-time linear system model, which is used to describe the state evolution relationship under the sampling period. S4. Based on the communication topology between players, and according to whether there is a directed path from other players to the current player, other players are divided into: reachable players and unreachable players. Reachable players are further divided into: adjacent reachable players and non-adjacent but reachable players based on whether they are directly adjacent to the current player. Different control strategies are set for different players. S5. Construct a cost function representing the performance index for each player, and establish the optimal control objective for each player; S6. An improved strategy iteration method is adopted to separate strategy data acquisition from strategy evaluation / strategy improvement, iteratively solve the unknown cost function to obtain the optimal control strategy, and use the optimal control strategy to control the robot system.

[0004] Optionally, S1 specifically includes: The nonlinear robot system is represented as: (1) in, It is the joint angle. It's joint velocity. It is joint acceleration. It is the inertia matrix. It is the Coriolis and centrifugal terms. This is the gravity term, describing the equivalent torque of gravity on each joint. It is the joint driving force, and each component corresponds to an actuator input; The equivalent state space form is organized as follows: (2) Select a desired equilibrium point for the system ,satisfy: (3) And perform a first-order Taylor expansion of the system near the desired equilibrium point: (4) After ignoring higher-order terms, the continuous-time linearized robot system is obtained: (5) Among them, matrix , It can be obtained analytically from the robot dynamics model, or calculated near the desired equilibrium point in engineering implementation using numerical difference methods. To control the input.

[0005] Optionally, S2 specifically includes: Due to control input Depend on It consists of several independent actuator control input channels, which can be decomposed into: (6) in Indicates the first Control input channels for each joint or actuator; Define the input matrix corresponding to each control input channel: (7) Therefore, the continuous-time linearized robot system (5) is rewritten as the continuous-time multiplayer decision system: (8) In this representation, each scalar input It is regarded as an independent decision-making entity, that is, a player.

[0006] Optionally, S3 specifically includes: Assume the system sampling period is The continuous-time linearized robot system (5) is discretized using the forward Euler method, and has (9) Substituting into the continuous-time multiplayer decision system (8), we get: (10) definition:

[0007] The discrete-time linear system model is then obtained: (11).

[0008] Optionally, S4 specifically includes: Let the player set be The communication topology is a directed graph. Description, in which Indicates player Able to receive from players Information for any target player Based on directed paths and adjacency relationships, the remaining players are classified as follows: Reachable and unreachable players are defined as follows: if there exists at least one path from player... To the player A directed path, denoted as Then the player is called For players It is reachable; otherwise it is inaccessible. Based on this definition: (12) in Represents a set of reachable players. Indicates an unreachable set of players; Due to players The quantity and information are unknown, and its impact on players... The effects are uniformly treated as bounded uncertainty terms in subsequent modeling. ; Adjacency repartitioning of reachable players: in the set of reachable players In the middle, further based on whether or not it is related to the player Divide into directly adjacent sections: Define Player The set of adjacent players is: (13) The player can then be broken down into: (14) in Indicates the set of reachable adjacent players. The table represents a set of reachable but non-adjacent players. ; Strategy setting principles: For players Player Its status / policy information can be obtained directly, so it is included in the collaboration item and a cooperative approach is adopted in control updates; For players Player Since its strategy cannot be directly observed or coordinated, but its effects are transmitted through system coupling, it is treated as a source of adversarial / perturbation, and its worst-case effects are conservatively characterized using a minimax structure.

[0009] Optionally, S5 specifically includes: For the discrete-time linear system model (11), from the player From the perspective of: (15) in For the first The neighboring players of each player The control input, For players The equivalent perturbation term for reachable non-adjacent players, For a bounded uncertain term, The input matrix is ​​of compatible dimension; For each player Construct local performance indices that include cooperative / confrontational / uncertain influences, and provide corresponding optimal control objectives: (16) in For cooperation weight, It is a weight matrix for robust compensation terms for unknown player inputs. To represent the weights for the adversarial / uncertainty terms, the positive terms in equation (16) correspond to cooperation / energy consumption, while the negative terms are used to express the competitor's weights under the minimax structure. The worst possible impact; The cost function for constructing a performance metric for each player is as follows: (17) Under information constraints, players The local optimal objective is defined as: (18).

[0010] Optionally, S6 specifically includes: S6.1 Off-strategy data acquisition: Initialize the permissible behavior policy for each player. Selecting permissible behavior strategies And add a small amount of exploratory noise to enhance data richness: Online operation and data logging: Strategies running on the discrete-time linear system model (11) Collection trajectory ,in This is the state trajectory of the nominal system, excluding uncertainties. The nominal system is as follows: (19) Create a regression dataset and reuse it across iterations: This dataset is reused in subsequent policy iterations, thereby reducing the cost of repeated sampling and improving data utilization efficiency; For any execution strategy and perturbation strategies The nominal system is written as: (20) Substituting the nominal system (20) into the cost function (17) yields: (twenty one) in , They represent control and disturbance respectively. Control and disturbance strategies during the iteration phase; S6.2 Evaluate the current player's strategy and solve for the corresponding cost function: First, due to the kernel matrix in equation (21) It is unknown and coupled with other player controls in (21), making it impossible to solve directly. Therefore, the Kronecker product property is used to parameterize it here: For each player The cost function (21) is rearranged into a vector product using the Kronecker product: (twenty two) in In order to be in The parameter vector to be estimated during the iteration phase, This indicates that the matrix is ​​vectorized; Substitute (22) into (21) and compare with the kernel matrix. Perform similar parameterization operations (22) on the relevant items, and rearrange (21) to obtain: (twenty three) in The Kronecker product of different inputs and states is defined as:

[0011] The corresponding unknown components are:

[0012] Rearranging equation (23) into a compact form, we get: (twenty four) in This represents the vector of unknown parameters to be solved, derived from the unknown kernel matrix in (21). composition, , Indicates the number of adjacent players; The recursive least squares method is used to estimate the unknown parameter matrix, and the formula is as follows: (25) in The adaptive weight matrix is ​​updated according to the following recursive relationship: (26) The initial values ​​of the weight matrix are set to... ,in It is a sufficiently large positive number; S6.3 Strategy Enhancement: Evaluation results based on S6.2 Components in Update current player strategy: against Players Exchange necessary information with neighboring collaborators to achieve joint improvement based on neighborhood information; make and They respectively represent the passage through the first The value matrix is ​​used in the next iteration. and proximity strategy Control and disturbance strategies obtained through single-step strategy improvement: (27) in and , , , express The inverse operation; The process of recursively optimizing the strategy based on the previously improved strategy is described by the following multi-policy improvement equation: (28) in Indicates player In the In the next iteration, through The strategy obtained through this strategy improvement. ; The multi-strategy improvement satisfies the convergence condition. or The time ends, among which For the preset threshold, This represents the maximum number of iterations. The final control strategy is set as follows: ; S6.4 Termination Condition: Set the end threshold The process terminates when the following condition is met for any player: (29) After the termination, each player The final strategy is taken as the optimal / approximate optimal control strategy, and the optimal control strategy of the player is output. and disturbance ; If not satisfied, then let Repeat S6.2 and S6.3 until equation (29) is satisfied; Using the optimal control strategy Controlling robots.

[0013] On the other hand, a robot system intelligent control system based on multi-player reinforcement learning is provided, the system comprising: The expansion module is used to select the desired equilibrium point of the robot system, and to perform a first-order Taylor expansion on the continuous-time nonlinear robot system at the desired equilibrium point to obtain a continuous-time linearized robot system. A forming module is used to treat each control input channel of the continuous-time linearized robot system as an independent player, and each player acts together on the same continuous linear system through its input, forming a continuous-time multi-player decision-making system. The discrete module is used to discretize the continuous-time linearized robot system using the forward Euler method, and substitute it into the continuous-time multiplayer decision system to obtain a discrete-time linear system model, which is used to describe the state evolution relationship under the sampling period. The segmentation module is used to classify other players into reachable players and unreachable players based on the communication topology between players and the existence of directed paths from other players to the current player. Reachable players are further divided into adjacent reachable players and non-adjacent reachable players based on whether they are directly adjacent to the current player. Different control strategies are set for different players. The building module is used to construct a cost function representing performance metrics for each player and to establish the optimal control objective for each player; The solution module is used to separate policy data acquisition from policy evaluation / policy improvement using an improved policy iteration method, iteratively solve the unknown cost function, obtain the optimal control policy, and use the optimal control policy to control the robot system.

[0014] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described intelligent control method for a robot system based on multi-player reinforcement learning.

[0015] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-described intelligent control method for a robot system based on multi-player reinforcement learning.

[0016] The beneficial effects of the technical solution provided by this invention include at least the following: This invention has strong engineering feasibility: by linearizing the desired equilibrium point and discretizing with Euler, a discrete linear-time system can be directly obtained, which is convenient for embedded / digital control implementation.

[0017] This invention is adapted to a directed information structure: players are divided into reachable / unreachable based on directed paths, and further distinguished between adjacent / non-adjacent, so that the control design is consistent with the information constraints; it uniformly handles "cooperation-confrontation-uncertainty": adjacent players cooperate, non-adjacent reachable players minimax, and unreachable players are treated as uncertainties, resulting in a clear and scalable structure.

[0018] This invention features model-free and high data efficiency: Improved policy iteration separates policy data acquisition from evaluation / improvement, allowing for the reuse of historical data and significantly reducing the number of evaluations; Fast solution speed: During the policy improvement phase, iterative updates are made by utilizing neighbor collaboration and combining current evaluation results, which can significantly accelerate the convergence of the optimal policy and reduce the number of iterations and evaluation overhead. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of an intelligent control method for a robot system based on multi-player reinforcement learning, provided by an embodiment of the present invention. Figure 2 This is a general block diagram of an intelligent control method for a robot system based on multi-player reinforcement learning, provided by an embodiment of the present invention. Figure 3 This is a block diagram of an intelligent control system for a robot system based on multi-player reinforcement learning, provided in an embodiment of the present invention. Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0022] This invention provides an intelligent control method for a robot system based on multi-player reinforcement learning. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of this method is shown below. Figure 2 This is a general block diagram of an intelligent control method for a robot system based on multi-player reinforcement learning, provided by an embodiment of the present invention. The processing flow may include the following steps: S1. Select the desired equilibrium point of the robot system, and perform a first-order Taylor expansion on the continuous-time nonlinear robot system at the desired equilibrium point to obtain the continuous-time linearized robot system. Optionally, S1 specifically includes: The nonlinear robot system is represented as: (1) in, It is the joint angle. It's joint velocity. It is joint acceleration. It is the inertia matrix. It is the Coriolis and centrifugal term (velocity-related nonlinear term). This is a gravity term (pose-dependent), describing the equivalent torque of gravity on each joint. It is the joint driving force, and each component corresponds to an actuator input; The equivalent state space form is organized as follows: (2) Select a desired equilibrium point for the system ,satisfy: (3) And perform a first-order Taylor expansion of the system near the desired equilibrium point: (4) After ignoring higher-order terms, the continuous-time linearized robot system is obtained: (5) Among them, matrix , It can be obtained analytically from the robot dynamics model, or calculated near the desired equilibrium point in engineering implementation using numerical difference methods. To control the input.

[0023] S2. Each control input channel of the continuous-time linearized robot system is regarded as an independent player. Each player acts on the same continuous linear system through its input, forming a continuous-time multi-player decision-making system. Optionally, S2 specifically includes: Due to control input Depend on It consists of several independent actuator control input channels, which can be decomposed into: (6) in Indicates the first Control input channels for each joint or actuator; Define the input matrix corresponding to each control input channel: (7) Therefore, the continuous-time linearized robot system (5) is rewritten as the continuous-time multiplayer decision system: (8) In this representation, each scalar input It is regarded as an independent decision-making entity, that is, a player.

[0024] S3. The continuous-time linearized robot system is discretized using the forward Euler method and substituted into the continuous-time multiplayer decision system to obtain a discrete-time linear system model, which is used to describe the state evolution relationship under the sampling period. Optionally, S3 specifically includes: Assume the system sampling period is The continuous-time linearized robot system (5) is discretized using the forward Euler method, and has (9) Substituting into the continuous-time multiplayer decision system (8), we get: (10) definition:

[0025] The discrete-time linear system model is then obtained: (11).

[0026] S4. Based on the communication topology between players, and according to whether there is a directed path from other players to the current player, other players are divided into: reachable players and unreachable players. Reachable players are further divided into: adjacent reachable players and non-adjacent but reachable players based on whether they are directly adjacent to the current player. Different control strategies are set for different players. Optionally, S4 specifically includes: Let the player set be The communication topology is a directed graph. Description, in which Indicates player Able to receive from players Information for any target player Based on directed paths and adjacency relationships, the remaining players are classified as follows: Distinguishing between reachable and unreachable players: If there exists at least one path from player... To the player A directed path, denoted as Then the player is called For players It is reachable; otherwise it is inaccessible. Based on this definition: (12) in Represents a set of reachable players. Indicates an unreachable set of players; Due to players The quantity and information are unknown, and its impact on players... The effects are uniformly treated as bounded uncertainty terms in subsequent modeling. ; Adjacency repartitioning of reachable players: in the set of reachable players In the middle, further based on whether or not it is related to the player Divide into directly adjacent sections: Define Player The set of adjacent players is: (13) The player can then be broken down into: (14) in Indicates the set of reachable adjacent players. The table represents a set of reachable but non-adjacent players. ; Strategy setting principles: For players Player Its state / policy information can be directly obtained, so it is included in the coordination term and a cooperative approach is adopted in the control update (e.g., sharing local estimates and policy parameters, or joint adjustment according to a consistent objective). For players Player Since its strategy cannot be directly observed or coordinated, but its effects are transmitted through system coupling, it is treated as a source of adversarial / perturbation, and its worst-case effects are conservatively characterized using a minimax structure.

[0027] S5. Construct a cost function representing the performance index for each player, and establish the optimal control objective for each player; Optionally, S5 specifically includes: For the discrete-time linear system model (11), from the player From the perspective of: (15) in For the first The neighboring players of each player The control input, For players The equivalent perturbation term for reachable non-adjacent players, For a bounded uncertain term, The input matrix is ​​of compatible dimension (it can be a known constant matrix or a matrix with...). Consistent structure matrix); For each player Construct local performance indices that include cooperative / confrontational / uncertain influences, and provide corresponding optimal control objectives: (16) in For cooperation weight, It is a weight matrix for robust compensation terms for unknown player input. To represent the weights for the adversarial / uncertainty terms, the positive terms in equation (16) correspond to cooperation / energy consumption, while the negative terms are used to express the competitor's weights under the minimax structure. The worst-case scenario (equivalently, can be understood as) (to "maximize"); The cost function for constructing a performance metric for each player is as follows: (17) Under information constraints, players The local optimal objective is defined as: (18).

[0028] S6. An improved strategy iteration method is adopted to separate strategy data acquisition from strategy evaluation / strategy improvement, iteratively solve the unknown cost function to obtain the optimal control strategy, and use the optimal control strategy to control the robot system.

[0029] Optionally, S6 specifically includes: S6.1 Off-policy policy data collection: Initialize the permissible behavior policy for each player. Choose a permissible (bounded / stable) behavioral strategy And add a small amount of exploratory noise to enhance data richness: Online operation and data logging: Strategies running on the discrete-time linear system model (11) Collection trajectory ,in This is the state trajectory of the nominal system, excluding uncertainties. The nominal system is as follows: (19) Create a regression dataset and reuse it across iterations: This dataset is reused in subsequent policy iterations, thereby reducing the cost of repeated sampling and improving data utilization efficiency; For any execution strategy and perturbation strategies The nominal system is written as: (20) Substituting the nominal system (20) into the cost function (17) yields: (twenty one) in , They represent control and disturbance respectively. Control and disturbance strategies during the iteration phase; S6.2 Evaluate the current player's strategy and solve for the corresponding cost function: First, due to the kernel matrix in equation (21) It is unknown and coupled with other player controls in (21), making it impossible to solve directly. Therefore, the Kronecker product property is used to parameterize it here: For each player The cost function (21) is rearranged into a vector product using the Kronecker product: (twenty two) in In order to be in The parameter vector to be estimated during the iteration phase, This indicates that the matrix is ​​vectorized; Substitute (22) into (21) and compare with the kernel matrix. Perform similar parameterization operations (22) on the relevant items, and rearrange (21) to obtain: (twenty three) in The Kronecker product of different inputs and states is defined as:

[0030] The corresponding unknown components are:

[0031] Rearranging equation (23) into a compact form, we get: (twenty four) in This represents the vector of unknown parameters to be solved, derived from the unknown kernel matrix in (21). composition, , Indicates the number of adjacent players; The recursive least squares method is used to estimate the unknown parameter matrix (since the standard least squares method requires batch data processing and is not suitable for the online implementation of this embodiment, the recursive least squares method is used in this embodiment), and its formula is as follows: (25) in The adaptive weight matrix is ​​updated according to the following recursive relationship: (26) The initial value of the weight matrix is ​​set to... ,in It is a sufficiently large positive number; S6.3 Strategy Enhancement: Evaluation results based on S6.2 Components in Update current player strategy: against Players Exchange necessary information with neighboring collaborators (such as...) (Local parameters, or estimated gains), thereby achieving joint improvement based on neighborhood information; make and They respectively represent the passage through the first The value matrix is ​​used in the next iteration. and proximity strategy The control strategy and disturbance strategy obtained by performing single-step strategy improvement: (27) in and , , , express The inverse operation; The process of recursively optimizing the strategy based on the previously improved strategy is described by the following multi-policy improvement equation: (28) in Indicates player In the In the next iteration, through The strategy obtained through this strategy improvement. ; The multi-strategy improvement satisfies the convergence condition. or The time ends, among which For the preset threshold, This represents the maximum number of iterations. The final control strategy is set as follows: ; S6.4 Termination Condition: Set the end threshold The process terminates when the following condition is met for any player: (29) After the termination, each player The final strategy is taken as the optimal / approximate optimal control strategy, and the optimal control strategy of the player is output. and disturbance ; If not satisfied, then let Repeat S6.2 and S6.3 until equation (29) is satisfied; Using the optimal control strategy Controlling robots.

[0032] like Figure 3 As shown, this embodiment of the invention also provides an intelligent control system for a robot system based on multi-player reinforcement learning, the system comprising: The expansion module 310 is used to select the desired equilibrium point of the robot system, and perform a first-order Taylor expansion on the continuous-time nonlinear robot system at the desired equilibrium point to obtain a continuous-time linearized robot system. The forming module 320 is used to treat each control input channel of the continuous-time linearized robot system as an independent player, and each player acts on the same continuous linear system through its input to form a continuous-time multi-player decision-making system. Discrete module 330 is used to discretize the continuous-time linearized robot system using the forward Euler method and substitute it into the continuous-time multiplayer decision system to obtain a discrete-time linear system model, which is used to describe the state evolution relationship under the sampling period. The partitioning module 340 is used to divide other players into reachable players and unreachable players based on the communication topology between players and whether there is a directed path from other players to the current player. Reachable players are further divided into adjacent reachable players and non-adjacent reachable players based on whether they are directly adjacent to the current player. Different control strategies are set for different players. Module 350 is used to construct a cost function representing performance metrics for each player and to establish the optimal control objective for each player. The solver module 360 ​​is used to employ an improved strategy iteration method to separate the strategy data acquisition from the strategy evaluation / strategy improvement, iteratively solve the unknown cost function, obtain the optimal control strategy, and use the optimal control strategy to control the robot system.

[0033] The intelligent control system for a robot system based on multiplayer reinforcement learning provided in this embodiment of the invention has a functional structure that corresponds to the intelligent control method for a robot system based on multiplayer reinforcement learning provided in this embodiment of the invention, and will not be described again here.

[0034] Figure 4This is a schematic diagram of the structure of an electronic device 400 provided in an embodiment of the present invention. The electronic device 400 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 401 and one or more memories 402. The memory 402 stores at least one instruction, which is loaded and executed by the processor 401 to implement the steps of the above-described intelligent control method for a robot system based on multi-player reinforcement learning.

[0035] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the aforementioned intelligent control method for a robot system based on multi-player reinforcement learning. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0036] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0037] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for intelligent control of a robot system based on multi-player reinforcement learning, characterized in that, The method includes: S1. Select the desired equilibrium point of the robot system, and perform a first-order Taylor expansion on the continuous-time nonlinear robot system at the desired equilibrium point to obtain the continuous-time linearized robot system. S2. Each control input channel of the continuous-time linearized robot system is regarded as an independent player. Each player acts on the same continuous linear system through its input, forming a continuous-time multi-player decision-making system. S3. The continuous-time linearized robot system is discretized using the forward Euler method and substituted into the continuous-time multiplayer decision system to obtain a discrete-time linear system model, which is used to describe the state evolution relationship under the sampling period. S4. Based on the communication topology between players, and according to whether there is a directed path from other players to the current player, other players are divided into: reachable players and unreachable players. Reachable players are further divided into: adjacent reachable players and non-adjacent but reachable players based on whether they are directly adjacent to the current player. Different control strategies are set for different players. S5. Construct a cost function representing the performance index for each player, and establish the optimal control objective for each player; S6. An improved strategy iteration method is adopted to separate strategy data acquisition from strategy evaluation / strategy improvement, iteratively solve the unknown cost function to obtain the optimal control strategy, and use the optimal control strategy to control the robot system.

2. The method according to claim 1, characterized in that, S1 specifically includes: The nonlinear robot system is represented as: (1) in, It is the joint angle. It's joint velocity. It is joint acceleration. It is the inertia matrix. It is the Coriolis and centrifugal terms. This is the gravity term, describing the equivalent torque of gravity on each joint. It is the joint driving force, and each component corresponds to an actuator input; The equivalent state space form is organized as follows: (2) Select a desired equilibrium point for the system ,satisfy: (3) And perform a first-order Taylor expansion of the system near the desired equilibrium point: (4) After ignoring higher-order terms, the continuous-time linearized robot system is obtained: (5) Among them, matrix , It can be obtained analytically from the robot dynamics model, or calculated near the desired equilibrium point in engineering implementation using numerical difference methods. For controlling input.

3. The method according to claim 2, characterized in that, S2 specifically includes: Due to control input Depend on It consists of several independent actuator control input channels, which can be decomposed into: (6) in Indicates the first Control input channels for each joint or actuator; Define the input matrix corresponding to each control input channel: (7) Therefore, the continuous-time linearized robot system (5) is rewritten as the continuous-time multiplayer decision system: (8) In this representation, each scalar input It is regarded as an independent decision-making entity, that is, a player.

4. The method according to claim 3, characterized in that, S3 specifically includes: Assume the system sampling period is The continuous-time linearized robot system (5) is discretized using the forward Euler method, and has (9) Substituting into the continuous-time multiplayer decision system (8), we get: (10) definition: The discrete-time linear system model is then obtained: (11)。 5. The method according to claim 4, characterized in that, S4 specifically includes: Let the player set be The communication topology is a directed graph. Description, in which Indicates player Able to receive from players Information for any target player Based on directed paths and adjacency relationships, the remaining players are classified as follows: Reachable and unreachable players are defined as follows: if there exists at least one path from player... To the player A directed path, denoted as Then the player is called For players It is reachable; otherwise it is inaccessible. Based on this definition: (12) in Represents a set of reachable players. Indicates an unreachable set of players; Due to players The quantity and information are unknown, and its impact on players... The effects are uniformly treated as bounded uncertainty terms in subsequent modeling. ; Adjacency repartitioning of reachable players: in the set of reachable players In the middle, further based on whether or not it is related to the player Divide into directly adjacent sections: Define Player The set of adjacent players is: (13) The player can then be broken down into: (14) in Indicates the set of reachable adjacent players. The table represents a set of reachable but non-adjacent players. ; Strategy setting principles: For players Player Its status / policy information can be obtained directly, so it is included in the collaboration item and a cooperative approach is adopted in control updates; For players Player Since its strategy cannot be directly observed or coordinated, but its effects are transmitted through system coupling, it is treated as a source of adversarial / perturbation, and its worst-case effects are conservatively characterized using a minimax structure.

6. The method according to claim 5, characterized in that, S5 specifically includes: For the discrete-time linear system model (11), from the player From the perspective of: (15) in For the first The neighboring players of each player The control input, For players The equivalent perturbation term for reachable non-adjacent players, For a bounded uncertain term, The input matrix is ​​of compatible dimension; For each player Construct local performance indices that include cooperative / confrontational / uncertain influences, and provide corresponding optimal control objectives: (16) in For cooperation weight, It is a weight matrix for robust compensation terms for unknown player input. To represent the weights for the adversarial / uncertainty terms, the positive terms in equation (16) correspond to cooperation / energy consumption, while the negative terms are used to express the competitor's weights under the minimax structure. The worst possible impact; The cost function for constructing a performance metric for each player is as follows: (17) Under information constraints, players The local optimal objective is defined as: (18)。 7. The method according to claim 6, characterized in that, S6 specifically includes: S6.1 Off-strategy data collection: Initialize the permissible behavior policy for each player. Selecting permissible behavior strategies And add a small amount of exploratory noise to enhance data richness: Online operation and data logging: Strategies running on the discrete-time linear system model (11) Collection trajectory ,in This is the state trajectory of the nominal system, excluding uncertainties. The nominal system is as follows: (19) Create a regression dataset and reuse it across iterations: This dataset is reused in subsequent policy iterations, thereby reducing the cost of repeated sampling and improving data utilization efficiency; For any execution strategy and perturbation strategies The nominal system is written as: (20) Substituting the nominal system (20) into the cost function (17) yields: (twenty one) in , They represent control and disturbance respectively. Control and disturbance strategies during the iteration phase; S6.2 Evaluate the current player's strategy and solve for the corresponding cost function: First, due to the kernel matrix in equation (21) It is unknown and coupled with other player controls in (21), making it impossible to solve directly. Therefore, the Kronecker product property is used to parameterize it here: For each player The cost function (21) is rearranged into a vector product using the Kronecker product: (22) in In order to be in The parameter vector to be estimated during the iteration phase, This indicates that the matrix is ​​vectorized; Substitute (22) into (21) and compare with the kernel matrix. Perform similar parameterization operations (22) on the relevant items, and rearrange (21) to obtain: (23) in The Kronecker product of different inputs and states is defined as: The corresponding unknown components are: Rearranging equation (23) into a compact form, we get: (24) in This represents the vector of unknown parameters to be solved, derived from the unknown kernel matrix in (21). composition, , Indicates the number of adjacent players; The recursive least squares method is used to estimate the unknown parameter matrix, and its formula is as follows: (25) in The adaptive weight matrix is ​​updated according to the following recursive relationship: (26) The initial values ​​of the weight matrix are set to... ,in It is a sufficiently large positive number; S6.3 Strategy Enhancement: Evaluation results based on S6.2 Components in Update current player strategy: against Players Exchange necessary information with neighboring collaborators to achieve joint improvement based on neighborhood information; make and They respectively represent the passage through the first The value matrix is ​​used in the next iteration. and proximity strategy The control strategy and disturbance strategy obtained by performing single-step strategy improvement: (27) in and , , , express The inverse operation; The process of recursively optimizing the strategy based on the previously improved strategy is described by the following multi-policy improvement equation: (28) in Indicates player In the In the next iteration, through The strategy obtained through this strategy improvement. ; The multi-strategy improvement satisfies the convergence condition. or The time ends, among which For the preset threshold, This represents the maximum number of iterations. The final control strategy is set as follows: ; S6.4 Termination Condition: Set an end threshold The process terminates when the following condition is met for any player: (29) After the termination, each player The final strategy is taken as the optimal / approximate optimal control strategy, and the optimal control strategy of the player is output. and disturbance ; If not satisfied, then let Repeat S6.2 and S6.3 until equation (29) is satisfied; Using the optimal control strategy Controlling robots.

8. An intelligent control system for a robot system based on multi-player reinforcement learning, characterized in that, The system includes: The expansion module is used to select the desired equilibrium point of the robot system, and to perform a first-order Taylor expansion on the continuous-time nonlinear robot system at the desired equilibrium point to obtain a continuous-time linearized robot system. A forming module is used to treat each control input channel of the continuous-time linearized robot system as an independent player, and each player acts on the same continuous linear system through its input, forming a continuous-time multi-player decision-making system. The discrete module is used to discretize the continuous-time linearized robot system using the forward Euler method, and substitute it into the continuous-time multiplayer decision system to obtain a discrete-time linear system model, which is used to describe the state evolution relationship under the sampling period. The segmentation module is used to classify other players into reachable players and unreachable players based on the communication topology between players and the existence of directed paths from other players to the current player. Reachable players are further divided into adjacent reachable players and non-adjacent reachable players based on whether they are directly adjacent to the current player. Different control strategies are set for different players. The building module is used to construct a cost function representing performance metrics for each player and to establish the optimal control objective for each player; The solution module is used to separate policy data acquisition from policy evaluation / policy improvement using an improved policy iteration method, iteratively solve the unknown cost function, obtain the optimal control policy, and use the optimal control policy to control the robot system.

9. An electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, characterized in that, The at least one instruction is loaded and executed by the processor to implement the intelligent control method for a robot system based on multi-player reinforcement learning as described in any one of claims 1-7.

10. A computer-readable storage medium storing at least one instruction, characterized in that, The at least one instruction is loaded and executed by the processor to implement the intelligent control method for a robot system based on multi-player reinforcement learning as described in any one of claims 1-7.