Delivery robot order allocation method based on deep reinforcement learning
Through the multi-task-near-end strategy optimization model of deep reinforcement learning, the problems of unreasonable allocation of distribution robot resources and insufficient path planning in hospitals are solved, and task completion time is shortened, timeout rate is reduced and resource utilization is improved, and the efficiency of hospital logistics distribution is improved.
Patent Information
- Application Number
- CN202510268253.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional distribution robot scheduling method has unreasonable resource allocation in hospitals, and the path planning cannot adapt to dynamically changing people and logistics situations in real time, resulting in inefficient delivery.
Using a multi-task-near-end strategy optimization model based on deep reinforcement learning, we optimize task allocation and path planning by building a Markov decision-making model and reward function, and combining environmental prediction tasks, we realize robot task load balancing and efficient resource utilization.
It significantly reduces the average completion time of the delivery task, reduces the timeout rate, improves logistics distribution efficiency, and optimizes the utilization rate of robot resources, avoiding the excessive or idle state of robot tasks.
Smart Images

Figure CN120297843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular, to an order allocation method for a distribution robot based on deep reinforcement learning. Background Art
[0002] In the unmanned intelligent logistics system established by a hospital, distribution robots are responsible for the transportation tasks of various materials such as drugs, consumables, blood products, and specimens among almost all departments such as outpatient clinics, wards, and operating rooms.
[0003] Traditional distribution robot scheduling methods are mostly rule-based or simple algorithms, and there are many limitations when facing complex and changeable hospital logistics scenarios. Such limitations are specifically manifested in that it is difficult to comprehensively consider multiple factors such as the urgency of tasks, the load capacity of robots, and the driving distance in task allocation, resulting in unreasonable resource allocation, some robots being overloaded, and some being idle.
[0004] Meanwhile, in terms of path planning, traditional methods often cannot adapt to the dynamically changing human flow and logistics situation in the hospital in real time. For example, in the event of a sudden major operation that leads to a sharp increase in the flow of people and materials, the path cannot be adjusted in time to avoid congestion, resulting in a significant reduction in distribution efficiency. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to overcome the technical defects of unreasonable resource allocation in the prior art and the inability of path planning to adapt to the dynamically changing human flow and logistics situation in the hospital in real time. To overcome the above defects of the prior art, the present invention provides an order allocation method for a distribution robot based on deep reinforcement learning.
[0006] An order allocation method for a distribution robot based on deep reinforcement learning provided by the present invention includes the following steps: S1: Clean and normalize the distribution task data, robot status data, and environmental data extracted from the hospital to obtain preprocessed data; S2: Construct a Markov decision model, and the state of the Markov decision model includes task information, robot information, and environmental information, and the actions of the Markov decision model include the assigned tasks and adjusted tasks of the robot; S3: Based on the Markov decision model, construct a multi-task - proximal policy optimization model with the model architecture designed to predict the main task and the auxiliary task, where the main task is the assigned task and adjusted task of the robot, and the auxiliary task is the environmental prediction task; S4: Use the preprocessed data to train the multi-task - proximal policy optimization model based on the training scheme of deep reinforcement learning to obtain a prediction model; S5: Obtain the environmental data of the hospital in real time, and input the environmental data into the prediction model to obtain the corresponding primary task and auxiliary task.
[0007] The order allocation method for delivery robots based on deep reinforcement learning disclosed in the present invention, aiming at the technical problems of the present invention, adopts a multi-task proximal policy optimization model (MT-PPO) to dynamically optimize the task scheduling strategy, reducing the average completion time of delivery tasks and the overtime rate, and significantly improving the logistics delivery efficiency. Moreover, by setting up a Markov decision model to construct an intelligent task balancing strategy for predicting the primary task and auxiliary task of the model architecture, the task load of the robot is more balanced, effectively avoiding the situation that some robots are overloaded while other robots are idle, and improving the resource utilization rate of the robot.
[0008] In a possible implementation manner, the step S1 includes the following steps: S11: Extract the delivery task information from the hospital logistics system to obtain the delivery task data; the delivery task information includes the task starting place, destination, task urgency, type and quantity of required delivery materials. S12: Collect the real-time status data of the delivery robot to obtain the robot status data; the real-time status data includes the position, current load, remaining power, and driving speed of the robot. S13: Collect the environmental information in the hospital to obtain the environmental data; the environmental information includes the population density, logistics congestion, and traffic conditions at different times in each area of the hospital. S14: Clean the delivery task data, the robot status data, and the environmental data to remove outliers and duplicate data, and obtain the cleaned data. S15: Normalize the cleaned data so that data with different features are on the same scale, and obtain the preprocessed data. This solution cleans the extracted delivery task data, robot status data, and environmental data, can remove outliers and duplicate data, and normalizes the cleaned data, enabling data with different features to be on the same scale, thus facilitating subsequent model processing.
[0009] In a possible implementation manner, the step S2 includes the following steps: S21: Construct a state including task information, robot information, and environmental information to obtain the state of the Markov decision model. S22: Construct task options including assigned tasks and adjusted tasks to obtain the actions of the Markov decision model. S23: Based on the results obtained in step S21 and step S22, construct a reward function that integrates factors such as the completion status of the fusion task, resource utilization, environmental adaptation, and task balanced allocation to obtain the Markov decision model; The Markov process constructed by this solution can provide a prerequisite guarantee for the parameter optimization and policy planning of the model, and can fully consider the environmental conditions of the hospital, improve the accuracy of the robot's task allocation, and avoid resource waste.
[0010] In a possible implementation, the task information includes the task urgency level, task type code, and task start and end points.
[0011] In a possible implementation, the robot information includes the current position coordinates, load capacity, battery level, and current task information of the robot.
[0012] In a possible implementation, the environmental information includes the congestion level and special events in each area of the hospital, and the special events include department renovation and surgery.
[0013] In a possible implementation, the expression of the reward function is as follows: , where, R represents the reward function value; () represents the indicator function, whose function value is 1 when the condition in the parentheses is satisfied and 0 when it is not satisfied; represents the time for the task to be completed ahead of schedule, and is 0 if not completed ahead of schedule; represents the urgency level, which is divided into four levels: 1, 2, 3, and 4, with 1 being the most urgent; represents the current number of tasks of the robot; represents the maximum number of tasks of the robot; represents the remaining battery level; represents that if the task is completed ahead of schedule, a reward is given according to the task urgency level and the time ahead of schedule The higher the urgency level and the longer the time ahead of schedule, the higher the reward; represents that if the task times out, a penalty is given according to the task urgency level; It represents that if the current load of the robot does not exceed 80% of the load capacity, a reward will be given; It represents that if the remaining power of the robot after the task is completed is in the range of 30% to 70%, a reward will be given; The reward function of this scheme comprehensively considers factors such as task completion, resource utilization, environmental adaptation, and task balanced allocation, and can guide the delivery robot to make optimal decisions in different scenarios.
[0014] In a possible implementation manner, the multi-task - proximal policy optimization model constructed in the step S3 includes: An attention layer module, which is composed of multiple attention layer units connected in series. Each of the attention layer units contains a multi-head self-attention layer and a feed-forward layer; A main task module, which communicates with the feed-forward layer at the end of the attention layer module and is used for predicting the main task allocation; the main task module includes a fully connected layer and a Softmax layer arranged in sequence along the running direction; An auxiliary task module, which communicates with the feed-forward layer at the end of the attention layer module and is used for predicting the auxiliary task allocation; the auxiliary task module includes a decoder; The above model is optimized based on the Markov decision model established in the present invention by proposing an innovative reinforcement learning algorithm strategy, and the model architecture is designed as a main - auxiliary task (dual-task architecture). Among them, the main task is task allocation, which can determine which robot to assign new tasks and tasks that have not been picked up yet. At the same time, a task of environmental prediction is added as an auxiliary task to enable the policy network to deeply understand the meaning of the environment and potential changes.
[0015] In a possible implementation manner, the step S4 includes the following steps: S41: Initialize the model parameters of the multi-task - proximal policy optimization model; the model parameters include the number of samples, batch size, network update interval, learning rate, discount factor of the advantage function estimation parameter, decay coefficient of the advantage function estimation parameter, and clipping coefficient; S42: Regard the robot as an agent to simulate its execution of the robot's allocation task and adjustment task in the simulation environment, and use the reward function and the preprocessed data to explore different scheduling strategies through the reinforcement learning agent exploration mechanism to collect interaction data, and at the same time adopt the experience replay mechanism to construct a sample pool for improving training stability; the interaction data includes robot state, action, reward, and next state; S43: Optimize the scheduling strategy according to the interaction data collected in the step S42 to obtain an optimization result; S44: Update the model parameters using the gradient ascent method based on the optimization result to obtain the prediction model. Description of the Drawings
[0016] Figure 1 The flowchart of a method for allocating orders of a delivery robot based on deep reinforcement learning disclosed in an embodiment of the present application; Figure 2 The schematic structural diagram of a multi-task proximal policy optimization model disclosed in an embodiment of the present application. Detailed Embodiments
[0017] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can adjust it according to needs to adapt to specific application scenarios.
[0018] The present application will be further described in detail below in conjunction with the drawings and specific embodiments.
[0019] See Figures 1 - 2 , an embodiment of the present application discloses a method for allocating orders of a delivery robot based on deep reinforcement learning, Figure 1 is the flowchart of this method, and this method includes the following steps: S1: Clean and normalize the delivery task data, robot state data, and environmental data extracted from the hospital to obtain preprocessed data.
[0020] Please continue to see Figure 1 , in this embodiment, step S1 includes the following steps: S11: Extract delivery task information from the hospital logistics system to obtain delivery task data; the delivery task information includes the task starting place, destination, task urgency, type and quantity of required delivery materials; S12: Collect the real-time state data of the delivery robot to obtain robot state data; the real-time state data includes the position of the robot, the current load, the remaining power, and the driving speed; S13: Collect the environmental information in the hospital to obtain environmental data; the environmental information includes the population density, logistics congestion, and traffic conditions at different times in each area of the hospital; S14: Clean the delivery task data, robot state data, and environmental data to remove outliers and duplicate data to obtain cleaned data; S15: Normalize the cleaned data so that data with different features are on the same scale to obtain preprocessed data.
[0021] S2: Construct a Markov decision model, where the states of the Markov decision model include task information, robot information, and environmental information, and the actions of the Markov decision model include task assignment and task adjustment for the robot.
[0022] In this embodiment, step S2 includes the following steps: S21: Construct states including task information, robot information, and environmental information to obtain the states of the Markov decision model; among them, the task information includes the task urgency level, task type encoding (one-hot encoding is used in this embodiment), and task start and end points. The robot information includes the current position coordinates of the robot, load capacity (loaded and maximum load), battery percentage and endurance time, and current task information. The environmental information includes the congestion level of each area in the hospital and special events, and the special events include department renovation and surgery.
[0023] S22: Construct task options including task assignment and task adjustment to obtain the actions of the Markov decision model; among them, task assignment means comprehensively considering the states of each delivery robot and environmental information to decide which robot will undertake the task. Task adjustment means reassigning the currently assigned but not yet picked-up tasks and the currently newly added tasks.
[0024] S23: Construct a reward function that integrates factors such as task completion, resource utilization, environmental adaptation, and task balanced distribution to obtain the Markov decision model. In this embodiment, the expression of the reward function is as follows:
[0025] , where R represents the reward function value; () represents the indicator function, and its function value is 1 when the condition in the parentheses is satisfied and 0 when it is not satisfied; represents the time of early task completion, which is 0 if not completed in advance; represents the urgency level, which is divided into four levels: 1, 2, 3, and 4. Level 1 is the most urgent, and levels 2, 3, and 4 change in the order of gradually decreasing urgency; the economic level can be divided according to the importance of time; represents the current number of tasks of the robot; represents the maximum number of tasks of the robot; represents the remaining battery power It represents that if the task is completed ahead of schedule, rewards will be given according to the urgency of the task and the time of early completion The higher the urgency and the longer the time of early completion, the higher the reward; It represents that if the task times out, penalties will be given according to the urgency of the task; It represents that if the current load of the robot does not exceed 80% of the load capacity, rewards will be given; It represents that if the remaining power of the robot after the task is completed is in the range of 30% to 70% (note: including the endpoints), rewards will be given.
[0026] S3: Based on the Markov decision model, the model architecture is designed as a multi-task proximal policy optimization model for predicting the main task and the auxiliary task. Among them, the main task is the assignment task and adjustment task of the robot, and the auxiliary task is the environmental prediction task.
[0027] See Figure 2 , the multi-task proximal policy optimization model constructed in step S3 of this embodiment includes an attention layer module, a main task module and an auxiliary task module. The attention layer module is composed of multiple attention layer units connected in series. Each attention layer unit contains a multi-head self-attention layer and a feed-forward layer; the main task module communicates with the feed-forward layer at the end of the attention layer module for predicting the main task assignment; the main task module includes a fully connected layer and a Softmax layer arranged in sequence along the running direction; the auxiliary task module communicates with the feed-forward layer at the end of the attention layer module for predicting the auxiliary task assignment; the auxiliary task module includes a decoder (Decoder in the figure).
[0028] S4: Based on the training scheme of deep reinforcement learning, the multi-task proximal policy optimization model is trained using the preprocessed data to obtain a prediction model.
[0029] In this embodiment, step S4 includes the following steps: S41: Initialize the model parameters of the multi-task proximal policy optimization model; the model parameters include the number of samples, batch size, network update interval, learning rate, discount factor of the advantage function estimation parameter, decay coefficient of the advantage function estimation parameter and clipping coefficient. The initialization results of these parameters in this embodiment are listed as follows:
[0030] Parameter Description Q Sampling number, set to 30000 Batch Batch size, set to 128 D Network update interval, set to 10 lr Learning rate, set to 1e - 4 γ Discount factor of the advantage function estimation parameter, set to 0.88 λ Decay coefficient of the advantage function estimation parameter, set to 0.9 ε Clipping coefficient, set to 0.2 Proceed to the next step on this basis.
[0031] S42: Treat the robot as an agent to simulate its execution of the assigned tasks and adjustment tasks of the robot in a simulated environment, and use the reinforcement learning agent exploration mechanism to explore different scheduling strategies by means of a reward function and preprocessed data to collect interaction data. At the same time, adopt an Experience Replay mechanism to construct a sample pool for improving training stability; the interaction data includes robot state, action, reward, and next state.
[0032] S43: Optimize the scheduling strategy according to the interaction data collected in step S42 to obtain an optimization result.
[0033] S44: Update the model parameters based on the optimization result by using the gradient ascent method (the Adam optimizer is used in this embodiment) to obtain a prediction model.
[0034] S5: Obtain the environmental data of the hospital in real time, and input the environmental data into the prediction model to obtain corresponding primary tasks and auxiliary tasks.
[0035] The following table shows the comparison between the model constructed by the method of this embodiment and the traditional allocation model. Model Average task completion time (ATCT, s) Task timeout rate (OTR, %) Robot average load rate (RAL, %) Robot idle rate (RIR, %) Baseline 1: Traditional rule - based scheduling (fixed - priority - based) 215.3 18.2 55.6 34.1 Baseline 2: BERT task allocation (BERT - Scheduler) 178.6 12.7 63.2 26.5 Baseline 3: Transformer - based reinforcement learning (TRPO - based) 154.2 9.5 71.5 18.7 Baseline 4: Large model GPT - 4 task 147.8 7.9 73.4 15.8 Our method: MT - PPO reinforcement learning (Ours) 132.5 5.6 79.3 11.2 It can be found from the above table that the model constructed by the method in this embodiment has relatively prominent technical advantages in terms of average task completion time, task timeout rate, average robot load rate, and robot idle rate.
[0036] The order allocation method for a delivery robot based on deep reinforcement learning disclosed in this embodiment, aiming at the technical problems of the present invention, adopts a multi-task - proximal policy optimization model (MT-PPO) to dynamically optimize the task scheduling strategy, reducing the average completion time of the delivery task and the timeout rate, and significantly improving the logistics distribution efficiency. By setting a Markov decision model to construct an intelligent task balancing strategy for predicting primary tasks and auxiliary tasks of the model architecture, the task load of the robot is more balanced, effectively avoiding the situation where some robots have too heavy tasks while other robots are idle, and improving the utilization rate of robot resources.
[0037] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present application.
[0038] In the description of the present application, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0039] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A method for allocating orders of a delivery robot based on deep reinforcement learning, characterized in that It includes the following steps: S1: Clean and normalize the distribution task data, robot status data, and environmental data extracted from the hospital to obtain preprocessed data; S2: Construct a Markov decision model, and the states of the Markov decision model include task information, robot information, and environmental information, and the actions of the Markov decision model include the assigned tasks and adjusted tasks of the robot; S3: Based on the Markov decision model, construct a multi-task - proximal policy optimization model with the model architecture designed to predict the main task and auxiliary task, where the main task is the assigned task and adjusted task of the robot, and the auxiliary task is the environmental prediction task; S4: Use the preprocessed data to train the multi-task - proximal policy optimization model based on the training scheme of deep reinforcement learning to obtain a prediction model; S5: Real-time obtain the environmental data of the hospital, and input the environmental data into the prediction model to obtain the corresponding main task and auxiliary task.
2. The order allocation method of the delivery robot based on deep reinforcement learning according to claim 1, wherein The step S1 includes the following steps: S11: Extract distribution task information from the hospital logistics system to obtain the distribution task data; the distribution task information includes the task starting point, destination, task urgency, type and quantity of required distribution materials; S12: Collect the real-time status data of the distribution robot to obtain the robot status data; the real-time status data includes the position of the robot, current load, remaining power, and driving speed; S13: Collect the environmental information in the hospital to obtain the environmental data; the environmental information includes the population flow density, logistics congestion, and traffic conditions at different times in each area of the hospital; S14: Clean the distribution task data, the robot status data, and the environmental data to remove outliers and duplicate data to obtain cleaned data; S15: Normalize the cleaned data so that data with different features are on the same scale to obtain the preprocessed data.
3. The order allocation method for a delivery robot based on deep reinforcement learning according to claim 2, wherein The step S2 includes the following steps: S21: Construct a state including task information, robot information, and environmental information to obtain the state of the Markov decision model; S22: Construct task options including assigned tasks and adjusted tasks to obtain the actions of the Markov decision model; S23: Based on the results obtained in the step S21 and the step S22, construct a reward function that integrates factors such as task completion, resource utilization, environmental adaptation, and task balanced allocation to obtain the Markov decision model.
4. The order allocation method for a delivery robot based on deep reinforcement learning according to claim 3, wherein The task information includes the task urgency level, task type code, and task start and end points.
5. The order allocation method of the delivery robot based on deep reinforcement learning according to claim 4, characterized in that The robot information includes the current position coordinates, load capacity, power, and current task information of the robot.
6. The order allocation method for a delivery robot based on deep reinforcement learning according to claim 5, characterized in that, The environmental information includes the congestion level in each area of the hospital and special events, and the special events include department renovation and surgery.
7. The order allocation method for a delivery robot based on deep reinforcement learning according to any one of claims 3-6, characterized in that The expression of the reward function is as follows: , In the formula, R represents the reward function value; () represents the indicator function, whose function value is 1 when the condition inside the parentheses is satisfied and 0 when it is not satisfied; Represents the time when the task is completed ahead of schedule, and is 0 if not completed ahead of schedule; Indicates the urgency level, which is divided into four levels: 1, 2, 3, and 4, with 1 being the most urgent; Represents the current number of tasks of the robot; Represents the maximum number of tasks for the robot; Represents the remaining battery level; Indicates that if the task is completed ahead of schedule, according to the urgency of the task and the time of early completion rewards will be given. The higher the urgency and the longer the early completion time, the higher the reward; It represents that if the task times out, a penalty will be imposed according to the urgency of the task; It represents that if the current load of the robot does not exceed 80% of the load capacity, a reward will be given; It represents that a reward will be given if the remaining battery power of the robot is in the range of 30% to 70% after the task is completed.
8. The order allocation method for a delivery robot based on deep reinforcement learning according to claim 7, wherein The multi-task - proximal policy optimization model constructed in the step S3 includes: An attention layer module, which is composed of multiple attention layer units connected in series, and each attention layer unit contains a multi-head self-attention layer and a feed-forward layer; The main task module communicates with the feed-forward layer at the end of the attention layer module and is used for predicting the main task assignment; the main task module includes a fully-connected layer and a Softmax layer arranged in sequence along the running direction; The auxiliary task module communicates with the feed-forward layer at the end of the attention layer module and is used for predicting the auxiliary task assignment; the auxiliary task module includes a decoder.
9. The order allocation method of the delivery robot based on deep reinforcement learning according to claim 8, characterized in that The step S4 includes the following steps: S41: Initialize the model parameters of the multi-task proximal policy optimization model; the model parameters include the number of samples, batch size, network update interval, learning rate, discount factor of the advantage function estimation parameter, decay coefficient of the advantage function estimation parameter, and clipping coefficient; S42: Regard the robot as an agent to simulate its execution of the robot's assignment task and adjustment task in the simulation environment, and use the reward function and the preprocessed data to explore different scheduling strategies through the reinforcement learning agent exploration mechanism to collect interaction data, and at the same time adopt the experience replay mechanism to construct a sample pool for improving training stability; the interaction data includes robot state, action, reward, and next state; S43: Optimize the scheduling strategy according to the interaction data collected in the step S42 to obtain an optimization result; S44: Update the model parameters using the gradient ascent method based on the optimization result to obtain the prediction model.