Quantitative perception distributed deep reinforcement learning method for multi-robot dynamic scheduling system
Through the quantitatively perceptual distributed deep reinforcement learning method, the problems of slow convergence and low deployment efficiency in the high-dimensional state space in multi-robot systems are solved, and efficient task scheduling and collaborative scheduling efficiency are improved.
Patent Information
- Application Number
- CN202510573636.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-27
AI Technical Summary
In multi-robot systems, traditional deep reinforcement learning methods converge slowly in high-dimensional state spaces, low deployment efficiency, and difficult to implement efficient and highly adaptable scheduling strategies, especially in high-real-time environments.
A quantitative perception distributed deep reinforcement learning method for multi-robot dynamic scheduling systems is proposed. By pre-processing task information, establishing distributed environment simulation models, and applying a distributed deep reinforcement learning algorithm that integrates quantitative perception training, the optimal real-time resource scheduling solution is generated.
It realizes a scheduling and deployment solution that quickly obtains high returns in high-dimensional global state space, optimizes the coordinated scheduling efficiency of robot clusters in cargo loading and unloading operations, and improves deployment efficiency and model convergence speed.
Smart Images

Figure CN120218360A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot cluster task scheduling optimization, and in particular to a quantization-aware distributed deep reinforcement learning method for a multi-robot dynamic scheduling system. Background Art
[0002] With the rapid development of intelligent manufacturing technology, smart ports have become an important platform for large-scale automated scheduling systems. Especially in container handling tasks, the efficient scheduling of multi-robot systems has become a research hotspot due to the timeliness requirements.
[0003] Existing research on optimizing multi-objective scheduling strategies can be mainly divided into two categories: exact methods and approximate methods. Exact methods obtain the optimal solution by systematically traversing the solution space. For example, Saberikia et al. proposed a mixed integer linear programming scheme for optimizing task allocation strategies in cloud service systems; Kim and Park used the branch and bound algorithm to solve the quay crane scheduling problem in port terminals. Approximate methods perform iterative optimization through heuristic or bio-inspired algorithms to obtain high-quality approximate solutions. Mokhtari et al. optimized the robot cooperation problem using an adaptive particle swarm optimization algorithm, significantly improving the task completion rate in warehouse picking and distribution scenarios; Chen et al. developed a multi-farm discrete artificial bee colony method for optimizing the task allocation process of multiple weeding robots. However, with the expansion of the task scale, the complexity of the solution space increases exponentially, which poses a major challenge to the application of exact methods and approximate methods in real-time task scheduling scenarios.
[0004] In recent years, deep reinforcement learning has gradually become a new trend in port scheduling research due to its ability to interact with dynamic environments and iteratively optimize strategies. For example, Yu et al. proposed a multi-agent reinforcement learning framework based on Deep Q-Network (DQN), significantly improving the efficiency of task scheduling optimization strategies in human-robot collaboration systems; Li et al. optimized the scheduling efficiency of the bulk cargo loading process through an improved DRL method (double DQN). However, with the expansion of the task scale, traditional DRL faces the dual challenges of slow convergence in high-dimensional state spaces and low deployment efficiency. Especially in high-real-time environments, there is still a lack of systematic research on how to develop efficient and adaptable scheduling strategies for multi-robot systems to optimize task allocation and collaboration.
[0005] Therefore, a quantization-aware distributed deep reinforcement learning method for a multi-robot dynamic scheduling system is provided to solve the above problems. Summary of the Invention
[0006] To solve the above problems, the present invention provides a quantization-aware distributed deep reinforcement learning method for a multi-robot dynamic scheduling system, which realizes the optimization of task scheduling in the scenario of a multi-port multi-robot system, can quickly obtain a scheduling deployment plan with higher benefits when facing a high-dimensional global state space, and optimizes the collaborative scheduling efficiency of the robot cluster in cargo handling operations.
[0007] To achieve the above object, the present invention provides a quantization-aware distributed deep reinforcement learning method for a multi-robot dynamic scheduling system, including the following steps:
[0008] S1: Preprocess the container handling tasks to be planned. The preprocessing includes discretizing the scheduling time, quantifying the demand priority, and recording the demand constraint information.
[0009] S2: Based on the handling tasks in S1 and the compatibility information between the container type and the robot mechanical characteristics, establish a distributed environment simulation model for the scheduling optimization problem. The specific process includes the definition of decision variables, constraint conditions, and the establishment of the objective function.
[0010] S3: Apply a distributed deep reinforcement learning algorithm integrated with quantization awareness training to solve the model to obtain an optimal real-time resource scheduling plan. Specifically, it includes the generation of the initial task resource planning environment model, the generation of the optimal planning plan using the deep reinforcement learning algorithm, and the performance evaluation of the planning plan. Take the port agents making independent decisions as the agents in the algorithm, and the robot numbers required to execute the container tasks to be planned as the action space of the agents. Continuously interact with the environment model within the scheduling window time for training and learning, and finally output the planning plan with the maximum task benefit obtained during training as the final optimal planning plan.
[0011] S4: Output and deploy the executable optimal real-time resource scheduling plan.
[0012] Preferably, in S1, the discrete scheduling time evenly divides the total scheduling duration into discrete time intervals, indicating the number of time units consumed by the robots under scheduling to complete the movement between ports and the execution of handling requirements at each decision step.
[0013] The quantification of demand priority is to quantify the priority order of demands, convert the priority order into different integers, and the higher the value, the higher the importance, urgency, or potential value of the demand.
[0014] The recording of demand constraint information includes the port number and container number generating the constraint. The port number determines the distance information of the destination during robot scheduling, and the container number determines the compatibility of the robot mechanical structure. There is heterogeneity in the robot mechanical structure, and the types of cargo materials suitable for handling are different.
[0015] Preferably, in S2, since all ports operate in a distributed control manner in the entire scheduling system, there is no direct information exchange decision variable between ports, and the decision variable is independently determined for each port. In the i-th decision step, the decision variable o n,m,i is defined as the robot number r n selected by port P m . Its value being 0 indicates the abandonment of the current task, and the expression is:
[0016]
[0017] In the formula, d n,i represents the task demand generated by port P n in the i-th decision step;
[0018] The definition of the constraint conditions is as follows. In the i-th decision step, the corresponding allocation status of robot r m is represented as a vector s m,u of length K, where the k-th component is represented by a boolean vector, describing whether robot r m is idle during the corresponding time period:
[0019] s m,i [k] = {0, 1}
[0020] The total time required to execute the demand includes the execution time and the transfer time between ports:
[0021] rt m,n,i,1 - rt m,n,i,0 = l m,n,i + dt n,i,2
[0022] In the formula, rt m,n,i,1 represents the actual end time for the robot to complete demand d n,i ; rt m,n,i,0 represents the actual start time for the robot to complete demand d n,i ; dt n,i,2 represents the time span required to execute demand d n,i ;
[0023] The actual task execution time period must be within the task executable time region, and the periods when the robot is in a non-working state due to planned maintenance, charging, or other factors must be considered:
[0024]
[0025] In the formula, dt n,i,0 represents the absolute generation time of demand d n,i ; dt n,i,1 represents the absolute generation time of demand d n,iThe latest deadline; T off The set of time periods indicating the non - working state of the robot;
[0026] To achieve the optimization goals of maximizing the time utilization rate of the robot and minimizing the motion energy consumption of the robot, an objective function is needed to evaluate the quality of the current decision - making process. The optimization goal of the robot scheduling problem is expressed as:
[0027]
[0028] In the formula, O n represents the total number of tasks completed at the port; P n represents the cumulative value of task priorities; R n represents the energy resources consumed by the robot's motion; ω p,n and ω o,n respectively represent the index weights of task priorities and resource constraints relative to the number of tasks completed; α e represents the energy consumption coefficient; p n,i represents the priority of demand d n,i of.
[0029] Preferably, in S3, it specifically includes the following steps:
[0030] S31: According to the task information collected and recorded in step S1 and the time of the scheduling window, form the task matrix to be scheduled in this window, and generate an environment model suitable for the deep reinforcement learning algorithm based on this task matrix;
[0031] S32: In view of the characteristics that the port independently manages multiple container units and the task demands are dynamically generated, during the port decision - making process, evaluate the time and resource constraints through a deep Q - network, and determine the selection of the robot according to the matching of the robot's mechanical structure and the time limit imposed by the task;
[0032] S33: The profit of the obtained final solution and the total profit of the tasks issued in the environment model are obtained through the following operation:
[0033]
[0034] S34: Analyze the obtained total score. The higher the total score, the better the performance of the obtained planning solution.
[0035] Preferably, in the task scheduling solution of S4, in addition to including attribute information such as the port container number information and task priorities of each task, it also includes task assignment information; specifically, it is manifested as the robot to which the task is assigned, the start time of executing the task, and the task termination time.
[0036] Preferably, in S32, the distributed Q-learning network extracts key features from the high-dimensional global state space, and focuses on learning batches of experiences with significant temporal difference errors in iterative learning to optimize the decision-making strategy. When making decisions, it combines the teammate cooperation model with the greedy MaxNextQ method. The network identifies and preferentially selects actions with higher potential Q-value growth, balancing the relationship between exploration and exploitation. When the selected task does not meet the constraints described above, the state remains unchanged and the selected action is not recorded;
[0037] After training the full-precision baseline model, the impact of quantization error is simulated by adding pseudo-quantization nodes, and the quantization-aware training method is used to improve the deployment efficiency and reduce the quantization error, accelerating the model deployment process.
[0038] Therefore, the present invention adopts the above-mentioned quantization-aware distributed deep reinforcement learning method for the multi-robot dynamic scheduling system, and has the following beneficial effects:
[0039] (1) By integrating the MaxNextQ strategy and the ε-greedy algorithm, the present invention balances the exploration and exploitation of potential optimal decisions in high-dimensional combinatorial optimization problems.
[0040] (2) The present invention introduces a fine-tuned quantization-aware training method to accelerate model convergence and improve deployment efficiency.
[0041] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of a dynamic multi-task multi-robot scenario in the present invention;
[0043] Figure 2 It is a schematic diagram of the distributed task scheduling process in the present invention;
[0044] Figure 3 It is a structural diagram of the distributed deep reinforcement learning algorithm based on quantization-aware training in an embodiment of the present invention;
[0045] Figure 4 It is the change situation of the distributed training results of the agents in an embodiment of the present invention;
[0046] Figure 5 It is a training result diagram of the quantization-aware distributed deep reinforcement learning algorithm in an embodiment of the present invention;
[0047] Figure 6 It is a performance comparison diagram of the quantization-aware DDRL algorithm and the centralized DRL algorithm in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.
[0049] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.
[0050] The terms "including" or "comprising" and the like used in the present invention mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. The orientation or positional relationship indicated by terms such as "inside", "outside", "above", "below", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly. In the present invention, unless otherwise clearly defined and limited, terms such as "attached" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be directly connected, or indirectly connected through an intermediate medium. It can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0051] Embodiment
[0052] A quantization-aware distributed deep reinforcement learning method for a multi-robot dynamic scheduling system includes the following steps:
[0053] S1: Preprocess the container loading and unloading tasks to be planned. The preprocessing includes scheduling time discretization, demand priority quantization, and demand constraint information recording. In S1, the scheduling time discretization evenly divides the total scheduling duration into discrete time intervals, indicating the number of time units consumed by the robots under scheduling to complete the movement between ports and the execution of loading and unloading demands at each decision step.
[0054] The demand priority quantization quantifies the priority order of demands and converts the priority order into different integers. The higher the value, the higher the importance, urgency, or potential value of the demand.
[0055] The demand constraint information recording includes the port number and container number that generate the constraint. The port number determines the distance information of the destination during robot scheduling, and the container number determines the compatibility of the robot's mechanical structure. There is heterogeneity in the robot's mechanical structure, and the types of cargo materials suitable for handling are different.
[0056] For example, robots equipped with suction cups or grippers are more suitable for handling flexible materials, while high-load and high-precision robotic arms are more suitable for rigid materials. In addition, time constraint information including the absolute generation time and the latest deadline of the demand, and the number of time units required for the robot to complete the demand is also included. The executable time of the demand can be expressed as a time interval starting from the absolute generation time and ending at the latest deadline, and the interval length is often greater than or equal to the execution consumption time of the demand, making the task allocation more flexible and increasing the solution complexity of the task scheduling at the same time.
[0057] S2: Based on the handling tasks in S1 and the compatibility information between the container type and the robotic characteristics of the robot, establish a distributed environment simulation model for the scheduling optimization problem. The specific process includes the definition of decision variables, constraint conditions, and the establishment of the objective function; in S2, since all ports operate in a distributed control manner in the entire scheduling system and there is no direct information exchange decision variable between ports, the decision variables are independently determined for each port. In the i-th decision step, the decision variable o n,m,i is defined as the robot number r n selected by port P m , and its value of 0 indicates giving up the current task. The expression is:
[0058]
[0059] where d n,i represents the task demand generated by port P n in the i-th decision step;
[0060] The definition of the constraint conditions is as follows. In the i-th decision step, the corresponding allocation status of robot r m is represented as a vector s m,i of length K, where the k-th component is represented by a boolean vector, describing whether robot r m is idle during the corresponding time period:
[0061] s m,i [k] = {0, 1}
[0062] The total time required to execute the demand includes the execution time and the transfer time between ports:
[0063] rt m,n,i,1 -rt m,n,i,0 = l m,n,i + dt n,i,2
[0064] where rt m,n,i,1 represents the actual end time for the robot to complete demand d n,i ; rt m,n,i,0 represents the actual end time for the robot to complete demand dn,i The actual start time; dt n,i,2 Indicates the execution requirement d n,i The required time span;
[0065] The actual execution time period of the task must be within the executable time region of the task. The time periods when the robot is in a non - working state due to planned maintenance, charging, or other factors need to be considered:
[0066]
[0067] In the formula, dt n,i,0 Indicates the absolute generation time of requirement d n,i ; dt n,i,1 Indicates the requirement d n,i The latest deadline; T off Indicates the set of time periods when the robot is in a non - working state;
[0068] To achieve the optimization goals of maximizing the robot's time utilization rate and minimizing the robot's motion energy consumption, an objective function is needed to evaluate the quality of the current decision - making process. The optimization goal of the robot scheduling problem is expressed as:
[0069]
[0070] In the formula, O n Indicates the total number of tasks completed at the port; P n Indicates the cumulative value of task priorities; R n Indicates the energy resources consumed by the robot's motion; ω p,n And ω o,n Respectively represent the index weights of task priority and resource constraint relative to the number of completed tasks; α e Indicates the energy consumption coefficient; p n,i Indicates the requirement d n,i The priority of.
[0071] S3: Apply the distributed deep reinforcement learning algorithm with fused quantization - aware training to solve the model to obtain the optimal real - time resource scheduling scheme, which specifically includes the generation of the initial task - resource planning environment model, the generation of the optimal planning scheme using the deep reinforcement learning algorithm, and the performance evaluation of the planning scheme; Use each independent - decision - making port agent as the agent in the algorithm, and the robot numbers required to plan the execution of container task requirements as the action space of the agent. Continuously interact with the environment model within the scheduling window time for training and learning, and finally output the planning scheme with the maximum task revenue obtained during training as the final optimal planning scheme;
[0072] In S3, it specifically includes the following steps:
[0073] S31: Based on the task information collected and recorded in step S1 and the time of the scheduling window, form the task matrix to be scheduled in this window, and generate an environment model adapted to the deep reinforcement learning algorithm based on this task matrix;
[0074] S32: In view of the characteristics that the port independently manages multiple container units and the task requirements are dynamically generated, during the port decision-making process, evaluate the time and resource constraints through a deep Q-network, and determine the selection of the robot according to the matching of the robot's mechanical structure and the time limit imposed by the task;
[0075] In S32, the distributed Q-learning network extracts key features from the high-dimensional global state space, and focuses on learning the experience batches with significant temporal difference errors in iterative learning to optimize the decision-making strategy. When making a decision, combine the teammate cooperation model with the greedy MaxNextQ method. The network identifies and preferentially selects actions with higher potential Q-value growth, balancing the relationship between exploration and exploitation. When the selected task does not meet the constraints described above, the state remains unchanged and the selected action is not recorded;
[0076] After training to obtain a full-precision benchmark model, simulate the impact of quantization errors by adding pseudo-quantization nodes, use the quantization-aware training method to improve the deployment efficiency and reduce quantization errors, and accelerate the model deployment process.
[0077] S33: The benefit of the obtained final solution is calculated with the total benefit of the tasks published in the environment model as follows:
[0078]
[0079] S34: Analyze the obtained total score. The higher the total score, the better the performance of the obtained planning solution.
[0080] S4: Output and deploy an executable optimal real-time resource scheduling solution.
[0081] In the task scheduling solution of S4, in addition to including attribute information such as the port container numbers and task priorities of each task, it also includes task assignment information; specifically, it is manifested as the robot to which the task is assigned, the start time of executing the task, and the task termination time.
[0082] Example 1
[0083] Such as Figure 1As shown, the automated control model is constructed based on the multi-agent system framework, where each agent is responsible for managing the container equipment within its fixed port area. In the context of large-scale automated industrial facilities, the number of available robots is much larger than the number of port agents, enabling multiple robots to operate simultaneously within the same port, handling various types of containers for loading and unloading operations. Meanwhile, the generation of loading and unloading demands is highly flexible, and future demands are inherently unpredictable. Due to the differences in importance and time limits among various demands, they have different priorities and completion deadlines. In this model, each port agent operates in a decentralized and collaborative manner, generating loading and unloading tasks from local containers distributively and making decisions on deploying idle and applicable robots distributively, thus realizing the dynamic scheduling of loading and unloading operations in a multi-port and multi-robot environment.
[0084] Taking into comprehensive consideration the intrinsic attributes of tasks and the schedulability of robots, each port can decide whether to execute the current task. If so, it further determines the robot number for scheduling. The process of task generation and assignment is as Figure 2 shown. The learning objective of the port agent is achieved by integrating the above variables in the following three stages: observing the unified state space before decision-making, independently selecting robots during the decision-making process, and evaluating the results based on a common reward function after decision-making. Since multiple port agents in the same scenario make decisions independently during decision-making, neither relying on past states, previous decisions, nor referring to the actions of other agents, but only depending on the observation of the current state space, the learning process of DDRL can be formalized as a finite Markov decision process. Thus, the scheduling problem can be effectively abstracted as an interaction process between the agent and the environment through DQN.
[0085] During the decision-making process, to design an effective action selection strategy, this study combines the greedy algorithm with the MaxNextQ strategy. The MaxNextQ strategy uses lists L1, L2, …, L N to store the set of actions corresponding to the maximum Q value for each port in the current state; the list L contains various action combinations with increased Q values in the next decision step. The Q-table is iteratively updated using the Bellman equation and approximated through DQN to ensure robust learning and adaptation. Finally, the decision-making process of the port probabilistically selects between the greedy exploration algorithm and the MaxNextQ algorithm. This method ensures sufficient exploration in the initial stage and tends to select actions that can increase the Q value, thus effectively achieving the optimization of the final goal.
[0086] To accelerate the rapid deployment of the DDRL model in a multi-robot scenario, this algorithm integrates two strategies: PER and QAT. This integrated method can not only maintain high accuracy but also improve the convergence speed of DDRL while reducing computational consumption. The PER strategy dynamically updates the priority of each experience by evaluating the temporal difference (TD) error and selects the experiences for training based on these priority metrics, enabling the model to focus on experiences with significant impacts and accelerating the learning process. The QAT strategy simulates quantization errors during the training phase and then fine-tunes the trained model to maintain the robust performance of the quantized model.
[0087] With this strategy, a standard accuracy model is trained by training the baseline DDRL model with 32-bit floating-point precision (FP32). For each port P n , a distributed experience pool E n is established to store the experiences generated during the interaction between the agent and the environment. These experiences include states, actions, rewards, subsequent states, and a flag indicating the end of the episode. The key of the PER strategy is to assign a priority to each experience. When the deep neural network samples a batch of experiences from E n , the sampling probability is determined by these priorities. Experiences with larger TD errors are more important for the current strategy and are thus assigned higher priorities. When the agents at each port sample from the experience pool in a distributed manner and then update the network parameters, the TD error fluctuates, and at this time, the priorities of the experiences need to be recalibrated. In the backpropagation stage of the Q-learning network, the network parameters are continuously optimized by the gradient descent method.
[0088] After training the FP32 baseline model, a QAT model is generated by inserting pseudo-quantization nodes. Figure 3 Shows the comparison of the network architectures before and after quantization. In the baseline model, temporal features are first extracted through convolutional layers, and then decisions are made by fully connected layers to estimate the Q value in DDRL; after introducing pseudo-quantization nodes, the weights before each convolutional layer and fully connected layer will be quantized, and the activation values after each activation layer will also undergo quantization and dequantization processes. These modifications aim to simulate a low-bit integer operation environment.
[0089] Although the pseudo - quantization nodes are introduced, the loss function of the QAT model is still calculated based on FP32. Therefore, the model needs to simulate the quantization error during backpropagation and fine - tune the weights and activation values to optimize the parameters in an iterative training manner, thus reducing the accuracy loss caused by quantization. After fine - tuning, a small - scale dataset is used for calibration. This step determines the quantization parameters of each layer by evaluating the statistical characteristics of the parameter data distribution of each layer. Finally, the QAT model is quantized using the quantization parameters S and Z, transformed into an INT8 model, and the quantized INT8 model is deployed in the test framework to reduce the storage requirements and computational complexity.
[0090] Actual port implementation cases
[0091] To verify the feasibility of this solution, a simulation experiment on the distributed scheduling problem of a multi - robot system in a port environment was carried out using the proposed quantization - aware DDRL algorithm. In the initial stage, two matrices were generated using MATLAB: the feasible demand set matrix and the robot - container timing constraint matrix The third dimension of matrix M f corresponds to the following task attributes: container type, task priority, earliest start time, deadline, and task execution duration. To ensure the diversity of the scheduling scenarios, all task parameters were generated by a distributed random method. Subsequently, the generated data was pre - processed to eliminate tasks that became infeasible due to constraint conflicts. The third dimension of matrix M c represents the feasible time periods for robots to perform container operations in the port. This set was constructed by fully considering the mechanical characteristics of the robots and key factors such as charging and maintenance, thus ensuring the feasibility of task scheduling and the stability of the overall system operation.
[0092] Under the PyTorch framework, a deep neural network was designed and implemented as Figure 3As shown, it is used to evaluate the value of state-action pairs and adopts a prioritized experience replay strategy to optimize the decision-making process. This network takes the robot allocation state as input. After shaping the input data, it sequentially passes through three convolutional layers: the first layer (conv1) uses 32 filters with a stride of 4; the second layer (conv2) uses 64 filters with a stride of 2; the third layer (conv3) uses 64 filters with a stride of 1. These convolutional layers gradually extract hierarchical features and reduce the data dimension through downsampling, thereby improving the computational efficiency. The extracted features are then input into two fully connected layers: the first layer (fc1) contains 512 neurons, which maps the features to a high-dimensional space to enhance the model's representation ability; the second layer (fc2) outputs a scalar representing the evaluation value of the action. To enhance the model's non-linear expression ability and alleviate the vanishing gradient problem, the rectified linear unit (ReLU) is used as the activation function after each convolutional layer and fully connected layer.
[0093] In the simulation experiment, the scheduling process is set to 20 minutes and discretized into 240 time steps, with each step lasting 5 seconds. Under this configuration, a scheduling system is constructed, including 10 ports and 50 robots. Each port can autonomously generate tasks and their related attributes throughout the scheduling cycle. The detailed configuration of the optimization model is shown in Table 1.
[0094] Table 1
[0095]
[0096] As Figure 4 shown, we plotted the iteration curves of the number of tasks executed by each agent and the corresponding scores during the distributed training process. It can be seen from the figure that in the initial few iterations, the number of tasks executed by each agent is relatively low; as the training progresses, the number of tasks executed by each agent gradually increases, and the curve oscillation gradually becomes stable.
[0097] To reduce the impact brought by the randomness of the algorithm, under the same experimental conditions, the DDRL algorithm was independently run 10 times, and its optimization performance was evaluated. The experimental results are as Figure 5 shown. The figure shows the changing trends of the cumulative reward and the total number of tasks. In the first 53 iterations, the learning curve generally shows an upward trend with a small amount of fluctuation, reflecting that the model is in the exploration stage; while between the 53rd and 60th iterations, the results gradually stabilize, indicating that the algorithm has entered the convergence stage and begins to optimize using the learned strategy. Finally, the experimental results prove that the DDRL algorithm achieves stable convergence after obtaining the optimal decision-making strategy, thus realizing efficient task allocation.
[0098] Subsequently, we introduced a metric for the work intensity, that is, the ratio R of robots to ports PRBased on the standard ratio of 50:10, we further examined scenarios with different numbers of robots and ports to simulate the algorithm performance under three different working intensities. To verify the effectiveness of the DDRL algorithm, we conducted a comparative analysis with the traditional DRL algorithm under the same constraints and task input conditions, and recorded the performance of the two algorithms in 10 repeated experiments. The evaluation metrics included the average task completion rate, the average total score, and the time required to complete the tasks.
[0099] Figure 6 The iterative training and optimization results under different workloads are shown. Initially, the trajectories of the two algorithms were relatively similar in the three scenarios; however, after the 40th iteration, the DDRL algorithm demonstrated stronger learning ability, with significant improvements in both the number of tasks and the score. Subsequently, we modified the constraints and task input for the test set deployment under different workload scenarios, and the results are shown in Table 2.
[0100] Table 2
[0101]
[0102]
[0103] The results indicate that in the three scenarios of standard, insufficient robots, and insufficient ports, the average task completion rate increased by 4.6%, 6.8%, and 6.2% respectively, the average total score increased by 5.75%, 6.32%, and 7.05% respectively, and the time required to complete the tasks decreased by 22.95%, 15.09%, and 23.37% respectively. These results fully demonstrate that the quantization-aware DDRL algorithm has stronger exploration ability in the scheduling optimization problem. In addition, the distributed policy helps to consider the preferences of different robot types among ports, accelerates model deployment, and improves the environmental perception efficiency and the efficiency of large-scale task scheduling. Under the many constraints described above, it can be concluded that the scheduling scheme obtained by the applied algorithm has better performance and meets the design expectations.
[0104] Therefore, the present invention adopts the quantization-aware distributed deep reinforcement learning method for the multi-robot dynamic scheduling system, which integrates the MaxNextQ strategy and the ε-greedy algorithm to balance the exploration and exploitation of potential optimal decisions in high-dimensional combinatorial optimization problems, and at the same time introduces a fine-tuned quantization-aware training method to accelerate model convergence and improve deployment efficiency.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A quantized perceptual distributed deep reinforcement learning method for multi-robot dynamic scheduling system, characterized by: The following steps are involved: S1: pre-processing the planned container loading and unloading tasks, including scheduling time discretization, demand priority quantification and demand constraint information recording; S2: Based on the loading and unloading tasks and the compatible information of container type and robot mechanical characteristics in S1, a distributed environment simulation model for the scheduling optimization problem is established. The specific process includes the definition of decision variables, constraints and the establishment of the objective function. S3: Apply the distributed deep reinforcement learning algorithm integrated with quantitative perception training to solve the model and obtain the optimal real-time resource scheduling solution, including the generation of the initial task resource planning environment model, the application of the deep reinforcement learning algorithm to generate the optimal planning solution, and the performance evaluation of the planning solution; the port intelligent agents that make independent decisions are used as agents in the algorithm, and the robot numbers required to plan the execution of container tasks are used as the action space of the agents. They are trained and learned by continuous interaction with the environmental model within the scheduling window time, and finally the planning solution with the maximum task benefit obtained in the training is output as the final optimal planning solution; S4: Output and deploy an executable optimal real-time resource scheduling solution.
2. The quantitative perception distributed deep reinforcement learning method for a multi-robot dynamic scheduling system according to claim 1, characterized in that: In S1, the scheduling time discretization evenly divides the total scheduling time into discrete time intervals, which represents the number of time units consumed by the scheduled robots to complete the movement between ports and the execution of loading and unloading requirements at each decision step; Demand priority quantification is to quantify the priority of requirements and convert the priority into different integers. The higher the value, the higher the importance, urgency or potential value of the requirement. The demand constraint information record includes the port number and container number that generate the constraint. The port number determines the distance information of the destination when the robot is scheduled, and the container number determines the compatibility of the robot's mechanical structure. The robot's mechanical structure is heterogeneous, and different types of cargo materials are suitable for handling.
3. The quantitative perception distributed deep reinforcement learning method for a multi-robot dynamic scheduling system according to claim 2, characterized in that: In S2, since all ports operate in a distributed control mode in the entire scheduling system, there is no direct information exchange decision variables between ports. The decision variables are determined independently for each port. In the i-th decision step, the decision variable o n,m,i Defined as port P n The robot number you selected m , whose value is 0 means giving up the current task, and the expression is: Where, d n,i Indicates that port P in the i-th decision step n Generated task requirements; The constraints are defined as follows: In the i-th decision step, the robot r m The corresponding allocation state is represented as a vector s of length K m,i , where the kth component is represented by a Boolean vector describing the robot r m Whether it is idle during the corresponding time period: s m,i [k]={0,1} The total time required to execute a request includes execution time and port transfer time: rt m,n,i,1 -rt m,n,i,0 =l m,n,i +dt n,i,2 In the formula, rt m,n,i,1 Indicates that the robot completes the requirement d n,i The actual end time of rt m,n,i,0 Indicates that the robot completes the requirement d n,i The actual start time of dt n,i,2 Indicates execution requirement d n,i The time span required; The actual execution time of the task must be within the task executable time area, and the period when the robot is in a non-working state due to planned maintenance, charging or other factors must be considered: Where, dt n,i,0 Indicates the demand n,i The absolute generation time of dt n,i,1 Indicates the demand n,i The latest deadline; T off A set of time periods representing the robot's non-working state; In order to achieve the optimization goal of maximizing the robot's time utilization and minimizing the robot's motion energy consumption, an objective function is needed to evaluate the pros and cons of the current decision-making process. The optimization goal of the robot scheduling problem is expressed as: In the formula, O n Indicates the total number of tasks completed by the port; P n Represents the cumulative value of task priority; R n Represents the energy resources consumed by the robot's motion; ω p,n and ω o,n Respectively represent the indicator weights of task priority and resource constraint relative to the number of completed tasks; α e represents the energy consumption coefficient; p n,i Indicates the demand n,i priority.
4. The quantitative perception distributed deep reinforcement learning method for a multi-robot dynamic scheduling system according to claim 3, characterized in that: In S3, the following steps are specifically included: S31: Based on the task information collected and recorded in step S1 and the time of the scheduling window, a task matrix to be scheduled in this window is formed, and an environment model adapted to the deep reinforcement learning algorithm is generated based on the task matrix; S32: In view of the fact that ports independently manage multiple container units and task requirements are dynamically generated, in the port decision-making process, time and resource constraints are evaluated through a deep Q-network, and the robot selection is determined based on the matching of the robot's mechanical structure and the time limit imposed by the task; S33: The benefit of the final solution and the total benefit of the tasks published in the environment model are calculated as follows: S34: Analyze the total score obtained. The higher the total score, the better the performance of the obtained planning scheme.
5. The quantitative perception distributed deep reinforcement learning method for a multi-robot dynamic scheduling system according to claim 4, characterized in that: The task scheduling scheme of S4 includes not only the attribute information of each task, such as the port container number information and the task priority, but also the task allocation information, which is specifically manifested as the robot to which the task is assigned, the start time of the task and the end time of the task.
6. The quantitative perception distributed deep reinforcement learning method for a multi-robot dynamic scheduling system according to claim 5, characterized in that: In S32, the distributed Q-learning network extracts key features from the high-dimensional global state space and focuses on learning experience batches with significant temporal difference errors in iterative learning to optimize the decision-making strategy. When making decisions, the teammate collaboration model and the greedy MaxNextQ method are combined. The network identifies and prioritizes actions with higher potential Q-value growth, balancing the relationship between exploration and utilization. When the selected task does not meet the constraints described above, the state remains unupdated and the selected action is not recorded. After training to obtain a full-precision benchmark model, pseudo-quantization nodes are added to simulate the impact of quantization errors, and quantization-aware training methods are used to improve deployment efficiency and reduce quantization errors, thereby accelerating the model deployment process.
Citation Information
Cited By
Coal bulk cargo loading and unloading efficiency optimization method and system based on machine learning
CN120806776A
Multi-modal large model driven robot cluster collaborative awareness and decision-making method
CN121959441A
Distributed combat mission scheduling method and system for unmanned equipment
CN121961162A