A Multi-Robot Dynamic Task Planning Method Based on Meta-Reinforcement Learning
Through a method based on meta reinforcement learning, a multi-robot dynamic task planning algorithm is designed, which solves the problem that multi-robot systems in the prior art are difficult to adapt quickly in a dynamic environment, and achieves optimization of obtaining efficient planning solutions in a short time and maintaining performance when scene changes.
Patent Information
- Application Number
- CN202410901603.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-07-05
AI Technical Summary
Existing multi-robot task planning algorithms are difficult to quickly adapt to scene changes in dynamic environments where tasks cannot be predetermined, and existing algorithms need to be adjusted or redesigned under performance guarantees.
Using a method based on meta-reinforcement learning, a task planning method that can quickly adapt to a dynamic environment is designed by establishing mathematical models of multiple representative task planning scenarios, pre-training and deep reinforcement learning are performed, and general task planning algorithm parameters are obtained, and fine-tuning is performed in the target scenario.
In dynamic task planning scenarios where tasks cannot be predetermined, efficient task planning schemes can be obtained in a short time, and performance can be maintained through few updates when scene changes, improving the algorithm's adaptability to dynamic environments.
Smart Images

Figure CN118859952B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of task scheduling of robot systems, and in particular relates to a multi-robot dynamic task planning method based on meta-reinforcement learning. Background Art
[0002] With the development of science and technology, intelligent manufacturing has become an important development direction of China's industry. Among them, multi-robot systems have been increasingly widely used in industrial production, and the development and optimization of their task planning methods have become the focus of research.
[0003] A multi-robot system consists of multiple robot units that cooperate to complete complex tasks such as assembly, handling, inspection, and packaging. Industrial robots, with their efficient, precise, and stable operation capabilities, significantly improve production efficiency and product quality, and reduce human errors and production costs. In addition, multi-robot systems have high flexibility and parallelism, and can greatly improve the execution efficiency of tasks through reasonable regulation of robot resources. When unexpected emergencies occur in the system, multi-robot systems can also quickly adapt to new situations through reasonable regulation of robot resources to ensure the smooth progress of tasks.
[0004] Facing the rapidly growing scale of industrial scenarios, the complexity of multi-robot systems also poses higher and more complex requirements for scheduling algorithms to fully exploit the potential of the system. Most existing multi-robot task planning algorithms require all tasks to be known before the planning starts, but in actual industrial scenarios, tasks may arrive at any time, which cannot meet this assumption; when the industrial system changes, most existing algorithms cannot quickly migrate to a new task planning scenario while ensuring the performance of the algorithm, but need to adjust or redesign the algorithm.
[0005] Therefore, it is necessary to provide a multi-robot dynamic task planning method based on meta-reinforcement learning to solve the above problems. Summary of the Invention
[0006] The object of the present invention is to provide a multi-robot dynamic task planning method based on meta-reinforcement learning, and design a task planning algorithm based on meta-reinforcement learning to solve the optimization problem on the task planning model. In a dynamic task planning scenario where tasks cannot be determined in advance, the algorithm can obtain a task planning scheme with higher efficiency in a short time, and when the scenario changes, the algorithm can reach the same performance level as before after a few updates, greatly improving the adaptability of the algorithm to dynamic environments.
[0007] To achieve the above object, the present invention provides a multi-robot dynamic task planning method based on meta-reinforcement learning, including the following steps:
[0008] S1: Establish mathematical models for multiple representative task planning scenarios;
[0009] S2: Apply the meta-reinforcement learning method to perform pre-training in the task planning scenarios established in step S1 to obtain general task planning algorithm parameters;
[0010] S3: Establish a mathematical model for the target task planning scenario;
[0011] S4: Apply the deep reinforcement learning method to fine-tune based on the algorithm parameters obtained in step S2 to obtain the optimal task planning method suitable for the target scenario.
[0012] Preferably, in step S1, select multiple representative task planning scenarios as pre-training scenarios, model them, define the key variables and constraints in the scenarios, and transform the actual planning problem into a combinatorial optimization problem, specifically expressed as:
[0013] Divide the time range of the task planning problem into P discrete time periods with equal lengths. Define that the time-related information of the task is discretized into integer multiples of the time period length. Then the time-related constraints of the task are shown in formulas (1) to (3).
[0014]
[0015] T s,m,n -T e,m,n =l g,m +l d,m (2)
[0016]
[0017] Among them, formula (1) means that the actual assigned time period of the task must be within the available time period of the task and the available time period of the task for the robot; formula (2) means that the length of the actual assigned time period of the task should be the sum of the time required to move to the task and the execution time of the task; formula (3) means that different tasks executed by the same robot should not overlap.
[0018] In formulas (1)-(3), m and m′ respectively represent the serial numbers of any two different tasks; n represents the serial number of any one robot; s represents the start time of the time period, e represents the end time of the time period; l d,m is the execution time of the mth task t m ; l g,m is the time required to move to the mth task t m ; t s,m and t e,m respectively represent the start time and end time of the available time period of the task t m ; t s,m,n and te,m,n respectively represent task t m for robot r n the start time and end time of the available time period; T s,m,n and T e,m,n represent task t m actually assigned to robot r n the start time and end time of the time period;
[0019] Define the objective function r as shown in formula (4):
[0020]
[0021] represents the set of robots of size N, represents the set of tasks of size M, P is the number of time segments; m represents the serial number of any task in; n represents the serial number of any robot in; p m represents task t m priority; o m,n represents task t m whether it is assigned to robot r n , when the value is 1, it means task t m is assigned to robot r n , when it is 0, it is the opposite; the more tasks planned successfully, the higher the priority of the planned tasks, and the larger the r value; while maximizing the r value, the algorithm should also obey the constraints between robots, and the constraint formula is as shown in formula (5):
[0022]
[0023] The constraint means that the number of times any task is executed by different robots should not be greater than 1;
[0024] When the algorithm interacts with the environment, it obtains the state from the environment and uses a state matrix of size N (the number of robots) multiplied by P (the number of time segments) to represent the occupancy of each robot in each time segment, as shown in formula (6):
[0025]
[0026] where s m,n,p represents that in the m-th step, robot r n is occupied in the p-th time segment, and all possible states of the environment are defined as the state set state matrix S m all possible values of Among them, under the condition of satisfying the constraints, in the m-th step, the action a taken by the algorithm m As shown in formula (7):
[0027] Where represents the action set of the algorithm, and a m = 0 means not taking any action and discarding task t m , while when a m = n, it means assigning the task to robot r n , and n is a positive integer not greater than N.
[0028] Preferably, in step S2, it specifically includes the following steps:
[0029] S21: Initialize the reinforcement learning model. The reinforcement learning model is a decision-making model constructed based on a deep neural network, which can select the plan for a certain task according to the state of the task planning scenario and obtain the optimal plan;
[0030] S22: In multiple pre-training scenarios, apply the deep reinforcement learning method to train multiple reinforcement learning models respectively. The deep reinforcement learning model is a deep Q-network. The neural network takes the state matrix as the input and outputs the value function Q of the state as the decision basis. In each decision-making, the deep reinforcement learning model estimates the Q value for the state matrix corresponding to each feasible plan of the task. The decision corresponding to the state matrix with the highest Q value is the best decision, and the selection method of the planning action is as shown in formula (7):
[0031]
[0032] Among them, ε is the exploration rate, indicating the degree to which the algorithm tends to explore new strategies, a m represents the action selected by the algorithm in the m-th step, and a r represents a random action sampled from the uniform distribution U(A) of the action space A;
[0033] The target value of the Q value can be expressed as the expected value of the weighted sum of the future reward values of the algorithm, and its calculation formula is as shown in (8):
[0034]
[0035] Among them, γ represents the decay value, represents the expectation operation.
[0036] After obtaining the Q value target value calculated by formula (8) and the estimated value output by the network, update the network parameters, as shown in formula (9):
[0037]
[0038] Among them, θj is the parameter of the inner-loop network after the j-th update; α is the learning rate of the inner-loop network, which determines the update speed of the network parameters; is the loss function, which is used to measure the difference between the Q-value target and the estimated value;
[0039] S23: In the pre-training scenario, generate experience information. Each inner-loop model interacts with the environment again, conducts task planning, and stores the experience information for subsequent steps;
[0040] S24: Update the outer-loop network using the experience information stored in step S23. Update the parameters through the experience stored in the inner-loop network, and learn the planning method in the pre-training scenario. The parameter update process of the outer-loop network is shown in formula (10):
[0041]
[0042] where θ i is the parameter of the outer-loop network after the i-th update, and β is the learning rate of the outer-loop network;
[0043] S25: Repeat steps S22 to S24, set the parameters of the inner-loop network to the same as those of the outer-loop network, and perform inner-loop network update, experience collection, and outer-loop network update.
[0044] Preferably, in step S4, in the task planning scenario where the desired application algorithm is used, use the deep reinforcement learning method to train the outer-loop network again. After repeated training, the outer-loop network should be able to obtain the optimal task planning solution in this scenario. When evaluating the task planning solution using formula (4), the r value can reach the highest; when the task scenario changes, train and update the pre-trained outer-loop network again to obtain a task planning algorithm suitable for the new environment.
[0045] Therefore, the present invention adopts the above-mentioned multi-robot dynamic task planning method based on meta-reinforcement learning, designs a task planning algorithm based on meta-reinforcement learning, and solves the optimization problem on the task planning model. In the dynamic task planning scenario where the task cannot be determined in advance, this algorithm can obtain a task planning solution with high efficiency in a short time, and when the scenario changes, this algorithm can reach the same performance level as before after a few updates, greatly improving the adaptability of the algorithm to the dynamic environment.
[0046] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0047] Figure 1 is the flowchart of a multi-robot dynamic task planning method based on meta-reinforcement learning of the present invention;
[0048] Figure 2 It is a schematic diagram of the task planning method in the embodiment of the present invention;
[0049] Figure 3 It is a schematic diagram of the structure of the deep Q - network in the embodiment of the present invention;
[0050] Figure 4 It is a result display diagram of the values obtained in different scenarios and different planning times in the embodiment of the present invention. Detailed implementation manners
[0051] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0052] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs.
[0053] The words such as "including" or "comprising" used in the present invention mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation to the present invention. When the absolute position of the object being described changes, the relative position relationship may also change accordingly. In the present invention, unless otherwise clearly defined and limited, terms such as "attached" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium. It can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0054] Embodiment
[0055] As Figures 1-4 shown, the present invention provides a multi - robot dynamic task planning method based on meta - reinforcement learning, including the following steps:
[0056] S1: Establish mathematical models of multiple representative task planning scenarios; In step S1, select multiple representative task planning scenarios as pre - training scenarios, and model them, define the key variables and constraint conditions in the scenarios, and transform the actual planning problem into a combinatorial optimization problem, which is specifically expressed as:
[0057] The time range of the task planning problem is divided into P discrete time periods of equal length. All time-related information of the tasks is discretized into integer multiples of the time period length. Then, the time-related constraints of the tasks are shown in Formulas (1) to (3).
[0058]
[0059] T s,m,n -T e,m,n =l g,m +l d,m (2)
[0060]
[0061] Among them, Formula (1) means that the actual assigned time period of the task must be within the available time period of the task and the available time period of the task for the robot; Formula (2) means that the length of the actual assigned time period of the task should be the sum of the time required to move to the task and the execution time of the task; Formula (3) means that different tasks executed by the same robot should not overlap.
[0062] In Formulas (1)-(3), m and m′ respectively represent the serial numbers of any two different tasks; n represents the serial number of any one robot; s represents the start time of the time period, and e represents the end time of the time period; l d,m is the execution time of the m-th task t m ; l g,m is the time required to move to the m-th task t m ; t s,m and t e,m respectively represent the start time and end time of the available time period of the task t m ; t s,m,n and t e,m,n respectively represent the start time and end time of the available time period of the task t m for the robot r n ; T s,m,n and T e,m,n represent the start time and end time of the time period when the task t m is actually assigned to the robot r n .
[0063] Define the objective function r as shown in Formula (4):
[0064]
[0065] represents a set of robots of size N, represents a set of tasks of size M, P is the number of time segments; m represents The serial number of any one of the tasks; n represents The serial number of any one of the robots; p m Indicates task t m 's priority; o m,n Indicates task t m Whether it is assigned to robot r n When the value is 1, it means task t m Is assigned to robot r n When it is 0, it is the opposite; the more tasks with successful planning, the higher the priority of the planned tasks, and the larger the r value; while maximizing the r value, the algorithm should also obey the constraints between robots, and the constraint formula is shown in formula (5):
[0066]
[0067] The constraint means that the number of times any task is executed by different robots should not be greater than 1;
[0068] When the algorithm interacts with the environment, it obtains the state from the environment and uses a state matrix of size N (the number of robots) multiplied by P (the number of time segments) to represent the occupancy of each robot in each time segment, as shown in formula (6):
[0069]
[0070] Among them, s m,n,p Indicates that in the m-th step, robot r n The occupancy situation on the p-th time segment, and all possible states of the environment are defined as the state set State matrix S m All possible values of Are in the state set m Under the condition of meeting the constraints, in the m-th step, the action a taken by the algorithm
[0071] Among them Represents the action set of the algorithm, a m =0 means not taking any action and discarding task t m , while a m =n means assigning the task to robot r n , and n is a positive integer not greater than N.
[0072] S2: Apply the meta-reinforcement learning method to perform pre-training in the task planning scenario established in step S1 to obtain the general task planning algorithm parameters; in step S2, it specifically includes the following steps:
[0073] S21: Initialize the reinforcement learning model. The reinforcement learning model is a decision-making model based on a deep neural network, which can select a plan for a certain task according to the state of the task planning scenario and obtain the optimal planning solution;
[0074] S22: In multiple pre-training scenarios, apply the deep reinforcement learning method to train multiple reinforcement learning models respectively. The deep reinforcement learning model is a deep Q-network. The neural network takes the state matrix as input and outputs the value function Q of the state as the decision basis. At each decision-making time, the deep reinforcement learning model estimates the Q value for the state matrix corresponding to each feasible planning solution of the task. The decision corresponding to the state matrix with the highest Q value is the best decision. The selection method of the planning action is shown in formula (7):
[0075]
[0076] where ε is the exploration rate, indicating the degree to which the algorithm tends to explore new strategies, a m represents the action selected by the algorithm in the m-th step, and a r represents a random action sampled from the uniform distribution U(A) of the action space A;
[0077] The target value of the Q value can be expressed as the expected value of the weighted sum of the future reward values, and its calculation formula is shown in (8):
[0078]
[0079] where γ represents the decay value, represents the expectation operation.
[0080] After obtaining the Q value target value calculated by formula (8) and the estimated value output by the network, update the network parameters, as shown in formula (9):
[0081]
[0082] where θ j is the parameter of the inner loop network after the j-th update; α is the learning rate of the inner loop network, which determines the update speed of the network parameters; is the loss function, which is used to measure the difference between the Q value target value and the estimated value;
[0083] S23: Generate experience information in the pre-training scenario. Each inner loop model interacts with the environment again, performs task planning, and stores the experience information for use in subsequent steps;
[0084] S24: Update the outer loop network using the experience information stored in step S23, and update the parameters through the experience stored in the inner loop network to learn the planning method in the pre-training scenario. The parameter update process of the outer loop network is shown in formula (10):
[0085]
[0086] where θ i is the parameter of the outer loop network after the i-th update, and β is the learning rate of the outer loop network;
[0087] S25: Repeat steps S22 to S24, set the parameters of the inner loop network to be the same as those of the outer loop network, and perform inner loop network update, experience collection, and outer loop network update.
[0088] S3: Establish a mathematical model for the target task planning scenario; the model establishment method is similar to step S1. Define the key variables and constraints in the scenario as shown in formulas (1) to (3) and formula (5), and establish the objective function according to formula (5).
[0089] S4: Apply the deep reinforcement learning method and fine-tune based on the algorithm parameters obtained in step S2 to obtain the optimal task planning method suitable for the target scenario.
[0090] In step S4, in the task planning scenario where the algorithm is expected to be applied, use the deep reinforcement learning method to train the outer loop network again. Similar to step S22, after repeated training, the outer loop network should be able to obtain the optimal task planning scheme in this scenario. When evaluating the task planning scheme using formula (4), the r value can reach the highest; when the task scenario changes, train and update the pre-trained outer loop network again to obtain the task planning algorithm suitable for the new environment.
[0091] The technical solution of the present invention is described in detail for the simulation model simulating the real environment.
[0092] Consider using 3 different pre-training scenarios to pre-train the algorithm.
[0093] In pre-training scenario 1, N = 40 robots cooperate to complete M = 50 tasks within 15 hours;
[0094] In pre-training scenario 2, N = 10 robots cooperate to complete M = 20 tasks within 15 hours;
[0095] In pre-training scenario 3, N = 10 robots cooperate to complete M = 50 tasks within 15 hours.
[0096] The execution time of each task ranges from 10 to 200 minutes, and it takes 10 minutes to move to the task location. The actual execution period of the task should be within the available time period of the task, and its length is the sum of the execution time and the moving time. Tasks cannot be divided or interrupted, and cannot be canceled after planning. To reduce the complexity of the problem, the time range is divided into equal time segments, each time segment is 10 minutes, then the task length is 1 to 20 time segments, and the moving time is 1 time segment. Each robot can only perform one task at a time, that is, there should be no overlap between different tasks performed by the robot. Since robot resources are limited and there may be conflicts between tasks, the algorithm needs to weigh the priorities of the tasks, select higher priority tasks as much as possible, and execute more tasks to obtain higher rewards as shown in formula (4).
[0097] Tasks are planned in the following ways: Figure 2 As shown in the figure. In the state matrix, the planning of tasks is carried out step by step. At each step, a decision is made on a task to determine whether to assign the task and to which robot the task is assigned. If the algorithm decides to assign the current task to a certain robot, the corresponding row of the robot in the state matrix needs to have an idle time period that satisfies the constraints (1) to (3) and (5) in order to schedule the task. After each decision is made, the algorithm updates the corresponding position in the state matrix, and repeats this M times until all tasks are planned and a complete task planning scheme is formed.
[0098] In the three pre-training scenarios, after pre-training as described in step S2, a general task planning method, i.e., the parameters of the outer loop network, is obtained. Subsequently, further training is performed in different environments as described in step S4 to obtain the best planning method suitable for the target environment. In the six different target scenarios, the r values obtained by the algorithm at different update times are as follows Figure 4 As shown, the horizontal axis is the number of updates and the vertical axis is the r value. Figure 4 It can be seen that in all scenarios, the algorithm can achieve a higher r value within 15 updates, that is, obtain better task planning effects and achieve the purpose of quickly adapting to new planning scenarios.
[0099] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
[0100] Therefore, the present invention adopts the above-mentioned multi-robot dynamic task planning method based on meta-reinforcement learning, designs a task planning algorithm based on meta-reinforcement learning, and solves the optimization problem on the task planning model. In the dynamic task planning scenario where the task cannot be determined in advance, this algorithm can obtain a highly efficient task planning scheme in a short time, and when the scenario changes, this algorithm can reach the same performance level as before after a few updates, greatly improving the adaptability of the algorithm to the dynamic environment.
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-robot dynamic task planning method based on meta-reinforcement learning, characterized in that: It includes the following steps: S1: Establish mathematical models for multiple representative task planning scenarios; Select multiple representative task planning scenarios as pre-training scenarios, model them, define the key variables and constraints in the scenarios, and transform the actual planning problem into a combinatorial optimization problem, specifically expressed as: Divide the time range of the task planning problem into discrete time periods, each with an equal length. Define that the time-related information of the task is discretized into integer multiples of the time period length. Then, the time-related constraints of the task are shown in Formulas (1) to (3). (1) (2) (3) Among them, formula (1) indicates that the actual assigned time period of a task must be within the available time period of the task and the available time period of the task for the robot; formula (2) indicates that the length of the actual assigned time period of a task should be the sum of the time required to move to the task and the execution time of the task; formula (3) indicates that different tasks executed by the same robot should not overlap; In Formulas (1)-(3), and respectively represent the serial numbers of any two different tasks; represents the serial number of any one robot; represents the start time of a time period, and represents the end time of the time period; is the execution duration of the th task ; is the duration required to move to the th task ; and respectively represent the start time and end time of the available time period of Task ; and respectively represent the start time and end time of the available time period of Task for Robot ; and represent the start time and end time of the time period when Task is actually assigned to Robot ; Define the objective function As shown in formula (4): (4) Denote a set of robots with size ; ; Denote a set of tasks with size ; ; is the number of time segments; Denote the serial number of any task in ; Denote the serial number of any robot in ; Denote the priority of task ; Denote whether task is assigned to robot , when the value is 1, it means that task is assigned to robot , and when the value is 0, it is the opposite; the more tasks with successful planning, the higher the priority of the planned tasks, the larger the value; while maximizing the value, the algorithm should also comply with the constraints among robots, and the constraint formula is shown in formula (5): (5) The constraint means that the number of times any task is executed by different robots should not be greater than 1; When the algorithm interacts with the environment, it obtains the state from the environment and uses a state matrix of size equal to the number of robots multiplied by the number of time segments to represent the occupancy of each robot in each time segment, as shown in Equation (6): (6) Among them, indicates that in the step, the occupancy of the robot in the th time segment. Define all possible states of the environment as the state set , and all possible values of the state matrix are within the state set ; S2: Apply the meta-reinforcement learning method to perform pre-training in the task planning scenarios established in step S1 to obtain general task planning algorithm parameters; S3: Establish a mathematical model for the target task planning scenario; S4: Apply the deep reinforcement learning method to fine-tune based on the algorithm parameters obtained in step S2 to obtain the optimal task planning method suitable for the target scenario.
2. The multi-robot dynamic task planning method based on meta-reinforcement learning according to claim 1, wherein: In step S2, it specifically includes the following steps: S21: Initialize the reinforcement learning model. The reinforcement learning model is a decision-making model constructed based on a deep neural network, which can select the planning for a certain task according to the state of the task planning scenario to obtain the optimal planning scheme; S22: In multiple pre-training scenarios, apply the deep reinforcement learning method to train multiple reinforcement learning models respectively. The deep reinforcement learning model is a deep Q network. The neural network takes the state matrix as the input and outputs the value function Q of the state as the decision-making basis. At each decision-making time, the deep reinforcement learning model estimates the Q value for the state matrix corresponding to each feasible planning scheme of the task. The decision corresponding to the state matrix with the highest Q value is the best decision. The selection method of the planning action is shown in formula (7): (7) Among them, represents the action set of the algorithm, is the exploration rate, indicating the degree to which the algorithm tends to explore new strategies, represents the action selected by the algorithm in the step, represents the average distribution in the action set and is a random action drawn from it; The target value of the value is expressed as the expected value of the weighted sum of future reward values, and its calculation formula is shown in (8): (8) Among them, represents the attenuation value, represents the desired operation; After obtaining the value of the target value calculated by Equation (8) and the estimated value output by the network, the network parameters are updated as shown in Equation (9): After obtaining the value of the target value calculated by Equation (8) and the estimated value output by the network, the network parameters are updated as shown in Equation (9): (9) Among them, is the parameter of the inner loop network after the -th update; is the learning rate of the inner loop network, which determines the update speed of the network parameters; is the loss function, which is used to measure the difference between the target value and the estimated value; S23: Generate experience information in the pre-training scenario. Each inner-loop model interacts with the environment again to perform task planning and store the experience information for use in subsequent steps; S24: Use the experience information stored in step S23 to update the outer-loop network. Update the parameters through the experience stored in the inner-loop network and learn the planning method in the pre-training scenario. The parameter update process of the outer-loop network is shown in formula (10): (10) Among them, is the parameter of the outer loop network after the -th update, is the learning rate of the outer loop network; S25: Repeat steps S22 to S24, set the parameters of the inner-loop network to be the same as those of the outer-loop network, and perform inner-loop network update, experience collection, and outer-loop network update.
3. A multi-robot dynamic task planning method based on meta-reinforcement learning according to claim 2, characterized in that: In step S4, in the task planning scenario where the algorithm is expected to be applied, the outer loop network is trained again using the deep reinforcement learning method. After repeated training, the outer loop network should be able to obtain the optimal task planning solution in this scenario. When evaluating the task planning solution using formula (4), the value can reach the highest; When the task scenario changes, train and update the pre-trained outer-loop network again to obtain a task planning algorithm suitable for the new environment.
Citation Information
Patent Citations
Open-scene-oriented multi-robot autonomous coordinated search and rescue method
CN110587606A
Efficient adaption of robot control policy for new task using meta-learning based on meta-imitation learning and meta-reinforcement learning
US20220105624A1