Multi-robot assembly line method and system based on temporal logic control strategy

Through a multi-robot assembly line method based on temporal logic control strategy, utilizing parity check game and reward automaton decomposition of potential energy function, combined with value iteration and decentralized Q-learning algorithm, the problems of slow convergence of multi-agent reinforcement learning and low efficiency of traditional assembly line are solved, and efficient and high-quality assembly line automation is achieved.

CN116787136BActive Publication Date: 2025-09-12CHANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310520398.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-10
Publication Date
2025-09-12
Estimated Expiration
2043-05-10

AI Technical Summary

Technical Problem

Existing multi-agent reinforcement learning methods suffer from slow convergence and sparse rewards when dealing with multi-task specifications. Traditional assembly line technology relies on manual operation, which is inefficient and costly, and makes it difficult to ensure quality.

Method used

A temporal logic-based control strategy is adopted. The temporal logic control strategy is synthesized through parity check game to express the robot's task specification. A reward automaton with potential energy function is constructed, which is decomposed into a reward automaton for each robot. The reward automaton is expanded on the Markov decision process, and the reward shaping algorithm of value iteration and the distributed Q-learning algorithm are combined to improve the learning speed.

Benefits of technology

It shortens the time for the robot group to learn the optimal strategy, improves the overall reward value, and improves the efficiency and quality of assembly line assembly, avoiding the local optimal problem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116787136B_ABST
    Figure CN116787136B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-robot assembly line method and system based on a temporal logic control strategy, comprising: expressing the robot's task specification based on a control strategy synthesized by temporal logic of a parity check game, constructing a reward automaton with a potential energy function according to the acceptance condition of the synthesized strategy to assign a reward value to the robot's behavior; decomposing the comprehensive strategy of a robot group consisting of a generalized reactive specification of rank 1 into a reward automaton for each robot, and expanding the reward automaton with potential energy on the MDP; proposing a reward shaping algorithm based on value iteration and a distributed Q-learning algorithm to improve the speed at which the robot group learns the optimal strategy; the present invention captures the temporal attributes of the task based on temporal logic, decomposes the generated comprehensive strategy into multiple individual reward automata to guide the robots to learn the optimal strategy, and proposes a reward shaping algorithm based on value iteration to improve the efficiency of the robot group in learning the optimal assembly strategy and avoid falling into the problem of falling into local optimality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-agent multi-task continuous control, and specifically to a multi-robot assembly line method and system based on a temporal logic control strategy. Background Art

[0002] Multi-agent systems are distributed computing systems in which multiple agents interact in a cooperative or adversarial manner within the same environment to maximize task completion and achieve specific goals. Multi-agent reinforcement learning is widely used in sequential decision-making problems in multi-agent systems. Reactive multi-agent reinforcement learning, a subfield of multi-agent reinforcement learning, focuses on how to train multiple agents to interact with each other to achieve common goals. However, the design complexity of reactive multi-agent reinforcement learning strategies typically increases with the complexity of the task to be completed. In addition, current multi-agent reinforcement learning methods suffer from slow convergence and sparse rewards when dealing with multi-task specifications.

[0003] LTL (Linear Temporal Logic) is a formal language that can describe complex non-Markov specifications. LTL is introduced to design task specifications in multi-agent multi-task learning. It can capture the temporal properties of the environment and tasks to express complex task constraints. This paper considers using a fragment of LTL, GR(1), to generate a comprehensive strategy for a robot group. The comprehensive strategy is then decomposed into a set of potential-based reward automata for individual robots to guide their learning.

[0004] Smart factories leverage advanced information technology, automated equipment, and intelligent manufacturing systems to digitize, intelligentize, and flexibly manage production processes, improving efficiency and quality while reducing costs and energy consumption. Assembly line technology is a crucial component of smart factories. Traditional assembly line technology primarily relies on manual labor, resulting in low production efficiency, high labor costs, and difficulty ensuring quality. In modern smart factories, assembly line technology has achieved a high degree of automation, utilizing industrial robots and automated equipment to complete product assembly, testing, and packaging, making the production process digital, intelligent, and flexible. Summary of the Invention

[0005] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.

[0006] In view of the above-mentioned problems, the present invention is proposed.

[0007] The first aspect of an embodiment of the present invention provides a multi-robot assembly line method based on a temporal logic control strategy, including: expressing the robot's task specification based on a control strategy of temporal logic synthesized by a parity check game, constructing a reward automaton with a potential energy function according to the acceptance condition of the synthetic strategy to assign a reward value to the robot's behavior; decomposing the comprehensive strategy of the robot group composed of a generalized reactive specification of rank 1 into a reward automaton for each robot, and expanding the reward automaton with potential energy on the MDP; proposing a reward shaping algorithm based on value iteration and a distributed Q-learning algorithm to improve the speed at which the robot group learns the optimal strategy.

[0008] As a preferred solution of the multi-robot assembly line method based on temporal logic control strategy described in the present invention, the expression of the robot's task specification includes:

[0009] Converting the control strategy formula of the parity check game synthetic temporal logic into a Buchi automaton with a single accepting state;

[0010] Constructing deterministic finite automata to guide the system to an acceptable state, where a rank-1 generalized reactive specification is a fragment of a control strategy of temporal logic, and a transformation of the formal language that can model the environment and task specifications of reactive robotic systems;

[0011] The calculation of the rank-1 generalized reactivity specification φ includes,

[0012]

[0013] in, A Boolean formula representing the initial state of the environment, A Boolean formula representing the initial state of the system, and The union of control strategy formulas representing temporal logic of system invariants, and represents the union of control policy formulas of temporal logic with active policies, and The Boolean formula representing the transformation relationship is valid at all times. and Indicates that the Boolean formula can always be established at some point in the future, ψ i A Boolean formula representing the system transformation relationship, Represents a Boolean formula.

[0014] As a preferred solution of the multi-robot assembly line method based on temporal logic control strategy described in the present invention, the reward value assigned to the robot's behavior includes:

[0015] Based on synthetic strategy Define a reward automaton with potential energy to assign reward values ​​to the robot's behavior, where ε represents a finite state set, ε0∈ε represents the initial state, and Γ represents the acceptable state set. Represents a set of actions, represents the transition function between states;

[0016] The reward automaton is defined as N= <E,E0,T,F,δ e ,δ r ,Ψ>, where E represents a finite state set, E0∈E represents the initial state, represents the set of accepted states, F represents the set of actions, δ e ∈E×F→E represents the transition function between states, represents the state reward function with the transition function, represents the potential energy function;

[0017] The synthesis strategy and the parameters of the reward automaton have a one-to-one correspondence, where δ r The calculation of (e, a), Ψ(e, a) depends on the state of the robot performing action a, where e∈E.

[0018] As a preferred solution of the multi-robot assembly line method based on temporal logic control strategy described in the present invention, it also includes:

[0019] When the state obtained by the inter-state transfer function does not belong to the set of accepted states, the robot is given a reward of 0, and Ψ(e, a) takes a value between 0 and rv. When the state obtained by the inter-state transfer function belongs to the set of accepted states, the robot is given a continuous reward rv, and Ψ(e, a) takes a value of pv. The formula is as follows:

[0020]

[0021]

[0022] Among them, rv and pv represent the reward values ​​given to the robot.

[0023] As a preferred solution of the multi-robot assembly line method based on temporal logic control strategy described in the present invention, the process of decomposing into a reward automaton for each robot includes:

[0024] Given the reward automaton N= <E,E0,T,F,δ e , δ r , Ψ> and local action set F i , define the mapping function A set of states from state to e∈E The formulas for the projection state of the mapping function and the reward automaton under the local action set are:

[0025]

[0026]

[0027] The reward automaton is decomposed into a single reward automaton using the local action set, and the state, initial state and final state of the single reward automaton are expressed using the mapping function Definition: If the transition state is the final state, the reward is set to rv and the potential value is set to pv; otherwise, the reward is set to 0 and the potential value is set to (0, pv);

[0028] The single reward automaton is defined as N i =<E i , E 0i , T i , F i , δ e i , δ r i ,Ψ i >, where:

[0029]

[0030]

[0031]

[0032] If and only if and e′ = δ e (e, a), δ e i ∈E i ×F i →E i is defined as

[0033] When satisfied is defined as otherwise

[0034] When e is satisfied i ∈T i , is defined as Ψ i (e i )=pv, otherwise Ψ i (e i )∈(0,pv).

[0035] As a preferred solution of the multi-robot assembly line method based on temporal logic control strategy described in the present invention, wherein: the expansion of the reward automaton with potential energy on the MDP includes:

[0036] The single reward automaton and the Markov decision process share the label function The expanded MDP is defined as in:

[0037]

[0038]

[0039]

[0040]

[0041] Among them, S represents the state set, s0 represents the initial state, A represents the action set, P represents the state transition probability, R represents the reward function of state transition, and γ represents the discount factor.

[0042] As a preferred solution of the multi-robot assembly line method based on temporal logic control strategy described in the present invention, wherein: improving the speed at which the robot group learns the optimal strategy includes:

[0043] A reward shaping algorithm based on value iteration is proposed to assign the potential function to each reward automaton. The goal of the reward shaping algorithm is to find the optimal strategy to maximize the expected cumulative reward. Initially, Ψ i The potential value of is 0. The value function of each state e is updated in each iteration. When the change of the state value is negligible, the algorithm terminates and uses the calculated value function as the potential energy function of the reward automaton.

[0044] It proposes to apply a distributed Q-learning algorithm to a robot group to learn the optimal strategy. The distributed Q-learning algorithm adopts MDP The individual reward automaton, learning rate and discount factor are used as inputs to output the Q-value function of each robot. When an event is a shared event between multiple robots, the conversion of the robot's individual reward automaton is triggered, that is, during the decentralized training process of the individual robot, an action will be taken with a predetermined probability p to observe the next state <s′, e′ i >, then calculate the reward based on the reward function and potential energy function, and calculate each state e according to the Bellman equation i The Q-value function of .

[0045] A second aspect of an embodiment of the present invention provides a multi-robot assembly line system based on a temporal logic control strategy, comprising:

[0046] a reward value assigning unit, configured to express the robot's task specification based on a control strategy of a parity check game synthetic temporal logic, and to assign a reward value to the robot's behavior by constructing a reward automaton with a potential energy function according to an acceptance condition of the synthetic strategy;

[0047] A policy decomposition unit is used to decompose the comprehensive policy of the robot group specified by the rank-1 generalized reactivity into the reward automaton of each robot and expand the reward automaton with potential energy on the MDP;

[0048] The learning speed improvement unit is used to propose a reward shaping algorithm based on value iteration and a decentralized Q-learning algorithm to improve the speed at which the robot group learns the optimal strategy.

[0049] According to a third aspect of an embodiment of the present invention, a device is provided, comprising:

[0050] processor;

[0051] a memory for storing processor-executable instructions;

[0052] The processor is configured to call the instructions stored in the memory to execute the method described in any embodiment of the present invention.

[0053] According to a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, including:

[0054] When the computer program instructions are executed by a processor, the method according to any embodiment of the present invention is implemented.

[0055] Beneficial effects of the present invention:

[0056] ① The present invention proposes to use the GR(1) specification to set the comprehensive strategy of the robot group, thereby shortening the time required for the robot group to learn the optimal strategy and improving the overall reward value obtained by the robot group;

[0057] ② In order to ensure that the decomposed individual reward automaton can ensure that the interaction of the entire robot still satisfies GR(1), this paper proposes to use parallel combination and forward simulation to verify the rationality of the method, and transition the definition of the overall reward automaton to the individual reward automaton for each robot to learn through a mapping function; since the reward automaton is related to the historical state of the label function, and the MDP in reinforcement learning only depends on the current state, we propose a method to expand the reward automaton on the MDP to solve this problem;

[0058] ③The present invention proposes to use a reward shaping algorithm based on value iteration to set a potential energy function for each state of the robot group to avoid the robot from transitioning from a high potential energy state to a low potential energy state, thereby ensuring that the robot group can learn the optimal strategy more effectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:

[0060] Figure 1 The overall framework diagram of the multi-robot assembly line method and system based on the temporal logic control strategy provided by the present invention;

[0061] Figure 2 This is the overall framework diagram of the reactive multi-robot reinforcement learning of the multi-robot assembly line method and system based on the temporal logic control strategy provided by the present invention;

[0062] Figure 3 State transition diagram and decomposition diagram of the multi-robot assembly line method and system based on temporal logic control strategy provided by the present invention under the φ control strategy;

[0063] Figure 4 A comparison chart of the training effects of different algorithms for three robot groups in the multi-robot assembly line method and system based on the temporal logic control strategy provided by the present invention;

[0064] Figure 5 This is a schematic diagram of the beer filling operation of multiple AGV carts in the multi-robot assembly line assembly method and system based on the temporal logic control strategy provided by the present invention. DETAILED DESCRIPTION

[0065] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0066] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0067] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0068] The present invention is described in detail with reference to schematic diagrams. For ease of illustration, cross-sectional views of device structures may be partially enlarged and not to scale when describing embodiments of the present invention. Furthermore, the schematic diagrams are merely illustrative and should not limit the scope of the present invention. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.

[0069] In the description of the present invention, it should be noted that the terms "upper, lower, inner, and outer" and other references to orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first, second, or third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0070] In this disclosure, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be interpreted broadly. For example, they may refer to fixed, removable, or integral connections. They may also refer to mechanical, electrical, or direct connections, indirect connections through an intermediary, or internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in this disclosure.

[0071] Example 1

[0072] Reference Figures 1 to 3 In one embodiment of the present invention, a multi-robot assembly line method based on a temporal logic control strategy is provided, comprising:

[0073] S1: Based on the control strategy of the parity check game synthetic temporal logic, the robot's task specification is expressed. According to the acceptance conditions of the synthetic strategy, a reward automaton with a potential energy function is constructed to assign reward values ​​to the robot's behavior. It should be noted that:

[0074] The robot's task specification includes:

[0075] The formula of control strategy of synthetic temporal logic of parity check game is transformed into Buchi automaton with single accepting state;

[0076] Construct a deterministic finite automaton to guide the system to an acceptable state, where the rank-1 generalized reactive specification is a fragment of the control strategy of temporal logic, which can transform the formal language for modeling the environment and task specifications of the reactive robot system. The GR(1) specification can express the impact of the environment specification on the task specification.

[0077] Specifically, such as Figure 3 As shown, the calculation of GR(1) reduction φ includes,

[0078]

[0079] in, A Boolean formula representing the initial state of the environment, A Boolean formula representing the initial state of the system, and The union of control strategy formulas representing temporal logic of system invariants, and represents the union of control policy formulas of temporal logic with active policies, and The Boolean formula representing the transformation relationship is valid at all times. and Indicates that the Boolean formula can always be established at some point in the future, ψ i A Boolean formula representing the system transformation relationship, represents a Boolean formula;

[0080] It should be explained that Indicates assumptions about the environment, It represents the assumption about the system. If the environment can satisfy the specification described by the formula of the control strategy in temporal logic, then the system can also satisfy this specification.

[0081] Furthermore, the reward values ​​assigned to the robot's behavior include:

[0082] Based on synthetic strategy Define a reward automaton with potential energy to assign reward values ​​to the robot's behavior, where ε represents a finite state set, ε0∈ε represents the initial state, and Γ represents the acceptable state set. Represents a set of actions, represents the transition function between states;

[0083] The reward automaton is defined as N = <E,E0,T,F,δ e , δ r ,Ψ>, where E represents a finite state set, E0∈E represents the initial state, represents the set of accepted states, F represents the set of actions, δ e∈E×F→E represents the transition function between states, represents the state reward function with the transition function, represents the potential energy function;

[0084] The parameters of the synthesis strategy and reward automaton correspond one to one, where δ r The calculation of (e, a) and Ψ(e, a) depends on the state of the robot performing action a, where e∈E;

[0085] Furthermore, when the state obtained by the inter-state transfer function does not belong to the set of accepted states, the robot is given a reward of 0, and Ψ(e, a) takes a value between 0 and rv. When the state obtained by the inter-state transfer function belongs to the set of accepted states, the robot is given a continuous reward rv, and Ψ(e, a) takes a value of pv. The formula is as follows:

[0086]

[0087]

[0088] Among them, rv and pv represent the reward values ​​given to the robot.

[0089] S2: Decompose the comprehensive strategy of the robot group specified by GR(1) into a reward automaton for each robot, and expand the reward automaton with potential energy on the MDP to solve the problem that the MDP only depends on the current state of the robot. It should be noted that:

[0090] The process of breaking down the reward automaton for each robot involves,

[0091] Given a reward automaton N=<E,E0,T,F,δ e , δ r , Ψ> and local action set F i , define the mapping function A set of states from state to e∈E Among them, the formulas of the projection state of the mapping function and the reward automaton under the local action set are:

[0092]

[0093]

[0094] The reward automaton is decomposed into a single reward automaton using the local action set. The state, initial state and final state of the single reward automaton are mapped using the mapping function Definition: The transition relationship is preserved only when there is a transition between different projected states. If the transition state is the final state, the reward is set to rv and the potential value is set to pv. Otherwise, the reward is set to 0 and the potential value is set to (0, pv).

[0095] It should be noted that a single reward automaton is defined as Ni = <E i , E oi , T i , F i , δ e i , δ r i ,Ψ i >, where:

[0096]

[0097]

[0098]

[0099] If and only if and e′ = δ e (e, a), δ e i ∈E i ×F i →E i is defined as

[0100] When satisfied is defined as otherwise

[0101] When e is satisfied i ∈T i , is defined as Ψ i (e i )=pv, otherwise Ψ i (e i )∈(0,pv);

[0102] It should be noted that by decomposing the comprehensive strategy into multiple reward automata for a single robot to learn, and then giving a local action set F i , compare the passed state with F\F iThe event merging in . The reward automaton N is constructed by a comprehensive strategy of the GR(1) specification, which ensures that the corresponding system behavior satisfies the GR(1) specification. Specifically, if the state is accepted, the operation of the reward automaton N is guaranteed to satisfy the GR(1) specification and can be visited infinitely. At the same time, we use forward simulation and parallel combination to verify that the parallel combination of the decomposed reward automaton can still guarantee the GR(1) specification.

[0103] Further, the expansion of reward automata with potential energy on MDP includes,

[0104] Single Reward Automata and Markov Decision Processes Shared Label Function The expanded MDP is defined as in:

[0105]

[0106]

[0107]

[0108]

[0109] Among them, S represents the state set, s0 represents the initial state, A represents the action set, P represents the state transition probability, R represents the reward function of state transition, and γ represents the discount factor;

[0110] It should be noted that if the next state is an accepting state, the reward function is updated to If the next state is unacceptable, the reward function is updated to For the sake of simplicity in calculation, and Ψ i Set under the same scalar;

[0111] It should be explained that Figure 2 The overall framework of reactive multi-robot reinforcement learning is presented. First, a comprehensive strategy generation tool such as gr1c is used to construct the GR(1) specification and the comprehensive strategy of the robot group. Considering the scenario where multiple robots cooperate with each other to complete the overall task, the entire warehouse area is divided into multiple grids, where x, y, and z are the three assembly line areas where robots need to complete the assembly. In the GR(1) specification, 0, 1, 2, and 3 represent the state sets corresponding to the robots, and we use the action annotations of each robot to connect the transitions between states.

[0112] If the comprehensive strategy defined by GR(1) is directly applied to the robot group, the robot group will need to spend a long time to learn the optimal strategy and may easily fall into local optimality during the learning process. This is because when a single robot learns its own strategy, the actions of other robots will not only affect the actions of the robot itself, but also affect the overall learning environment of the robot group. Therefore, we decompose the comprehensive strategy generated by GR(1) into multiple reward automata for each robot to learn. Each robot has its own sub-strategy, so that it will not be affected by other robots during the learning process, greatly improving the efficiency of the robot learning the optimal strategy and effectively avoiding the problem of falling into local optimality. However, when the comprehensive strategy generated by GR(1) is decomposed into multiple individual reward automata, there may be a combination of multiple reward automata that does not meet the original comprehensive strategy. Therefore, we use forward simulation and parallel combination of reward automata to verify that the interaction between the decomposed robots still meets the GR(1) specification. We also ensure that the original reward automata can be in an acceptable state only if and only if all individual reward automata reach an acceptable state. Finally, we expand the MDP to a reward automata for distributed training.

[0113] S3: We propose a reward shaping algorithm based on value iteration and a distributed Q-learning algorithm to improve the speed at which the robot group learns the optimal strategy.

[0114] If a typical reinforcement learning method is directly applied to the control of a robot group, the robot group usually needs to complete the entire task to obtain a reward. Usually, such methods are not applicable to situations with multiple subtasks. Reward shaping algorithms are applied to reinforcement learning, usually adding additional rewards to guide the robot group to achieve a specific goal. In reward shaping, a potential energy function is used to motivate the robot to perform behaviors with high potential energy states and prevent behaviors that lead to low potential energy states. Therefore, a reward shaping algorithm based on value iteration is proposed to assign a potential energy function to each reward automaton. The goal of the reward shaping algorithm is to find the optimal strategy to maximize the expected cumulative reward. Initially, Ψ is set. i The potential value of is 0. The value function of each state e is updated in each iteration. When the change of the state value is negligible, the algorithm terminates and uses the calculated value function as the potential energy function of the reward automaton. The specific implementation is shown in Table 1; Table 1: Reward shaping algorithm based on value iteration.

[0115]

[0116] Furthermore, although the centralized training-distributed execution algorithm can effectively deal with the non-stationary problem of the robot group in the learning interaction process, its scalability is poor. Therefore, it is proposed to apply the distributed Q-learning algorithm to the robot group to learn the optimal strategy. The distributed Q-learning algorithm adopts MDP The individual reward automaton, learning rate and discount factor are used as inputs to output the Q-value function of each robot. When an event is a shared event between multiple robots, the conversion of the robot's individual reward automaton is triggered, that is, during the decentralized training process of the individual robot, an action will be taken with a predetermined probability p to observe the next state. <s′,e′ i >, then calculate the reward based on the reward function and potential energy function, and calculate each state e according to the Bellman equation i The Q-value function of , the specific implementation is shown in Table 2; Table 2: Decentralized Q-learning algorithm for extended MDP.

[0117]

[0118] It should be noted that, ① although most reinforcement learning algorithms have been verified to be able to handle multi-task scenarios of robot groups, the algorithms generally have the following two problems: (1) The action of a single robot will change the state of the environment, and the change in the state of the environment will cause the actions of other robots to change accordingly. (2) At any moment, a single robot not only needs to consider the optimal strategy under the current state, but also needs to consider the actions of other robots; this greatly increases the time required for the entire robot group to learn the optimal strategy. The present invention proposes to use the GR (1) specification to set the comprehensive strategy of the robot group, thereby shortening the time required for the robot group to learn the optimal strategy and increasing the reward value obtained by the robot group as a whole;

[0119] ② To ensure that the decomposed individual reward automata can ensure that the overall robot interaction still satisfies GR(1), this paper proposes to use parallel combination and forward simulation to verify the rationality of the method. Through the mapping function, the definition of the overall reward automata is transferred to the individual reward automata for each robot to learn. Since the reward automata are related to the historical state of the label function, while the MDP in reinforcement learning only depends on the current state, we propose a method to expand the reward automata on the MDP to solve this problem.

[0120] ③ When using traditional reinforcement learning methods to train robots to complete assembly tasks on industrial assembly lines, robots typically only receive their rewards after completing the entire task, which results in a long period of time for the robots to learn the optimal strategy. This paper proposes a reward shaping algorithm based on value iteration, assigning a potential energy function to each state of the robot group. This prevents the robots from transitioning from high-potential to low-potential states, thereby ensuring that the robot group can more effectively learn the optimal strategy.

[0121] The second aspect of the present invention is disclosed.

[0122] Provides a multi-robot assembly line system based on temporal logic control strategy, including:

[0123] A reward value assignment unit is used to express the robot's task specifications based on a control strategy of parity check game synthetic temporal logic, and to construct a reward automaton with a potential energy function according to the acceptance conditions of the synthetic strategy to assign reward values ​​to the robot's behavior;

[0124] A policy decomposition unit is used to decompose the comprehensive policy of the robot group specified by the rank-1 generalized reactivity into the reward automaton of each robot and expand the reward automaton with potential energy on the MDP;

[0125] The learning speed improvement unit is used to propose a reward shaping algorithm based on value iteration and a distributed Q-learning algorithm to improve the speed at which the robot group learns the optimal strategy.

[0126] The third aspect of the present invention is disclosed.

[0127] Provided is a device comprising:

[0128] processor;

[0129] a memory for storing processor-executable instructions;

[0130] The processor is configured to call instructions stored in the memory to execute any one of the aforementioned methods.

[0131] The fourth aspect of the present invention is disclosed.

[0132] A computer-readable storage medium is provided, on which computer program instructions are stored, including:

[0133] When the computer program instructions are executed by a processor, any of the above methods is implemented.

[0134] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0135] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0136] Example 2

[0137] Reference Figures 4-5 This is the second embodiment of the present invention. Different from the first embodiment, this embodiment provides a verification test of a multi-robot assembly line method and system based on a temporal logic control strategy to verify and illustrate the technical effects adopted in this method.

[0138] This example uses three AGVs as examples to illustrate the method flow. The AGVs move and cooperate based on the GR(1) protocol to complete the overall task. We set the state space of the AGVs completing the assembly task as discrete, while the action space of the AGVs is continuous. Regarding the state space, this example decomposes the entire warehouse into multiple grids. Regarding the action space, this example uses sensors to detect the speed of the AGVs and the relative position intervals between the AGVs to avoid collisions during training.

[0139] In this embodiment, each robot has different task areas, namely x, y, and z, where x represents the filling area, y represents the capping area, and z represents the box sealing area. Figure 5 As shown in Figure 1, the AGV team completes the beer filling production through mutual cooperation; the system constraints of the three robots generated by GR (1) can be expressed as follows:

[0140] ENV:x;

[0141] SYS:yz;

[0142] ENVINIT:!x;

[0143] ENVTRANS:[](!x->x');

[0144] ENVGOAL:[]<>x;

[0145] SYSINIT: !y&!z;

[0146] SYSTRANS:[](x&!y->y')

[0147] &[](x&y&!z->z')

[0148] &[](!x&!z->x');

[0149] SYSGOAL:[]<>(y&z);

[0150] Initially, AGV 1 is not in area x, AGV 2 is not in area y, and AGV 3 is not in area z. If AGV 1 is not currently in area x, it needs to first go to the assembly line in area x for filling operations; if AGV 1 stays at the assembly line in area x for filling operations, AGV 2 needs to reach the assembly line in area y for capping operations; if AGV 1 and AGV 2 have already been performing filling and capping operations on the assembly lines in area x and area y respectively, AGV 3 needs to go to the assembly line in area z for box sealing operations; if AGV 1 is not performing filling operations on the assembly line in area x, and AGV 3 is not performing box sealing operations on the assembly line in area z, AGV 1 needs to first go to the assembly line in area x for filling operations; each AGV needs to visit the corresponding assembly line area infinitely frequently to complete the corresponding operations.

[0151] This embodiment addresses the robot group assembly problem and proposes a robot group reinforcement learning method based on a temporal logic control strategy. The proposed comprehensive strategy is decomposed into multiple reward automata for each robot to learn, thereby improving the overall training speed. The present invention also proposes applying a reward shaping algorithm based on value iteration and a distributed Q-learning algorithm with extended MDP to the robot group assembly training process.

[0152] In order to verify the proposed method, this example compares the training effects of the proposed method (decentralized Q-learning method with reward-shaped individual reward automata (DQ-iRM+RS)) with the following four methods: ① decentralized Q-learning method with individual reward automata (DQ-iRM), ② hierarchical reinforcement learning method with individual reward automata (H-IL+RS), ③ independent Q-learning method with individual reward automata (IQL+RS), and ④ centralized Q-learning method with individual reward automata (CQRM).

[0153] The comparison results are as follows Figure 4 As shown in the figure, it can be seen that the method provided by the present invention can enable the robot group to learn the optimal strategy faster than other methods, and can also obtain higher cumulative rewards. Therefore, the present invention captures the temporal properties of the task based on temporal logic, decomposes the generated comprehensive strategy into multiple individual reward automata to guide the robots to learn the optimal strategy, and proposes a reward shaping algorithm based on value iteration to improve the efficiency of the robot group learning the optimal assembly strategy and avoid falling into the problem of local optimality.

[0154] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A multi-robot assembly line method based on temporal logic control strategy, characterized in that: include: A control strategy based on parity check game synthetic temporal logic is used to describe the robot's task specification, and a reward automaton with a potential energy function is constructed according to the acceptance conditions of the synthetic strategy to assign reward values ​​to the robot's behavior. Decompose the comprehensive strategy of a robot group specified by a rank-1 generalized reactivity into a reward automaton for each robot, and extend the reward automaton with potential energy on the MDP; A reward shaping algorithm based on value iteration and a decentralized Q-learning algorithm are proposed to improve the speed at which the robot group learns the optimal strategy.

2. The multi-robot assembly line method based on temporal logic control strategy according to claim 1, characterized in that: The robot's task specification includes: Converting the control strategy formula of the parity check game synthetic temporal logic into a Buchi automaton with a single accepting state; Constructing deterministic finite automata to guide the system to an acceptable state, where a rank-1 generalized reactive specification is a fragment of a control strategy of temporal logic, and a transformation of the formal language that can model the environment and task specifications of reactive robotic systems; The calculation of the rank-1 generalized reactivity specification φ includes, in, A Boolean formula representing the initial state of the environment, A Boolean formula representing the initial state of the system, and The union of control strategy formulas representing temporal logic of system invariants, and represents the union of control policy formulas of temporal logic with active policies, and The Boolean formula representing the transformation relationship is valid at all times. and It means that a Boolean formula will always be true at some point in the future.

3. The multi-robot assembly line method based on temporal logic control strategy according to claim 2, characterized in that: Assigning reward values ​​to the robot's behavior includes: Based on synthetic strategy Define a reward automaton with potential energy to assign reward values ​​to the robot's behavior, where ε represents a finite state set, ε0∈ε represents the initial state, and Γ represents the acceptable state set. Represents a set of actions, represents the transition function between states; The reward automaton is defined as N= <E,E0,T,T,δ e ,δ r ,Ψ>, where E represents a finite state set, E0∈E represents the initial state, represents the set of accepted states, F represents the set of actions, δ e ∈E×F→E represents the transition function between states, represents the state reward function with the transition function, represents the potential energy function; The synthesis strategy and the parameters of the reward automaton have a one-to-one correspondence, where δ r The calculation of (e, a), Ψ(e, a) depends on the state of the robot performing action a, where e∈E.

4. The multi-robot assembly line method based on temporal logic control strategy according to claim 3, characterized in that: Also includes, When the state obtained by the inter-state transfer function does not belong to the set of accepted states, the robot is given a reward of 0, and Ψ(e,a) takes a value between 0 and rv. When the state obtained by the inter-state transfer function belongs to the set of accepted states, the robot is given a continuous reward rv, and Ψ(e,a) takes a value of pv. The formula is as follows: Among them, rv and pv represent the reward values ​​given to the robot.

5. The multi-robot assembly line method based on temporal logic control strategy according to claim 4, characterized in that: The decomposition into a reward automaton for each robot involves, Given the reward automaton N= <E,E0,T,F,δ e ,δ r ,Ψ> and local action set F i , define the mapping function Indicates mapping a state e∈E of the reward machine to a set of states The formulas for the projection state of the mapping function and the reward automaton under the local action set are: The reward automaton is decomposed into a single reward automaton using the local action set, and the state, initial state and final state of the single reward automaton are expressed using the mapping function Definition: If the transition state is the final state, the reward is set to rv and the potential value is set to pv; otherwise, the reward is set to 0 and the potential value is set to (0, pv); The single reward automaton is defined as N i = <E i ,E 0i ,T i ,F i ,δ e i ,δ r i ,Ψ i >, where: If and only if and e′ = δ e (e,a),δ e i ∈E i ×F i →E i is defined as When satisfied is defined as otherwise When e is satisfied i ∈T i , is defined as Ψ i (e i )=pv, otherwise Ψ i (e i )∈(0,pv).

6. The multi-robot assembly line method based on temporal logic control strategy according to claim 5, characterized in that: The expansion of the reward automaton with potential energy on the MDP includes, The single reward automaton and the Markov decision process share the label function The expanded MDP is defined as in: Among them, S represents the state set, s0 represents the initial state, A represents the action set, P represents the state transition probability, R represents the reward function of state transition, and γ represents the discount factor.

7. The multi-robot assembly line method based on temporal logic control strategy according to claim 6, characterized in that: Improving the speed at which the robot group learns the optimal strategy includes, A reward shaping algorithm based on value iteration is proposed to assign the potential function to each reward automaton. The goal of the reward shaping algorithm is to find the optimal strategy to maximize the expected cumulative reward. Initially, Ψ i The potential value of is 0. The value function of each state e is updated in each iteration. When the change of the state value is negligible, the algorithm terminates and uses the calculated value function as the potential energy function of the reward automaton. It proposes to apply a distributed Q-learning algorithm to a robot group to learn the optimal strategy. The distributed Q-learning algorithm adopts MDP The individual reward automaton, learning rate and discount factor are used as inputs to output the Q-value function of each robot. When an event is a shared event between multiple robots, the conversion of the robot's individual reward automaton is triggered, that is, during the decentralized training process of the individual robot, an action will be taken with a predetermined probability p to observe the next state. <s′,e′ i >, then calculate the reward based on the reward function and potential energy function, and calculate each state e according to the Bellman equation i The Q-value function of .

8. A multi-robot assembly line system based on a temporal logic control strategy, applying the multi-robot assembly line method based on a temporal logic control strategy according to any one of claims 1 to 7, characterized in that: include: a reward value assigning unit, configured to express the robot's task specification based on a control strategy of a parity check game synthetic temporal logic, and to assign a reward value to the robot's behavior by constructing a reward automaton with a potential energy function according to an acceptance condition of the synthetic strategy; A policy decomposition unit is used to decompose the comprehensive policy of the robot group specified by the rank-1 generalized reactivity into the reward automaton of each robot and expand the reward automaton with potential energy on the MDP; The learning speed improvement unit is used to propose a reward shaping algorithm based on value iteration and a decentralized Q-learning algorithm to improve the speed at which the robot group learns the optimal strategy.

9. A device, characterized in that The device comprises, processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Sequential logic task planning method based on reinforcement learning

    CN110014428A

  • Method for achieving robot square part assembling based on deep reinforcement learning

    CN110666793A