An optimization method for joint sorting of small-piece sorters in a logistics transfer yard
The integration of digital twin technology and MAPPO algorithm in logistics hubs optimizes sorting machine operations, addressing inefficiencies by enhancing throughput, accuracy, and resource utilization through adaptive strategies.
Patent Information
- Application Number
- CN202510467523.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-15
AI Technical Summary
Traditional logistics sorting systems are difficult to take into account high productivity, high accuracy and resource utilization at the same time, and lack dynamic joint serial parallel sorting strategies, so they cannot efficiently use various sorting machines to work together.
Digital twin technology is used to combine multi-agent reinforcement learning (MAPPO algorithm) with the Markov decision-making process (MDP) framework to build serial and parallel sorting strategies, and efficient coordinated scheduling of equipment is achieved through multi-agent near-end strategy optimization algorithms, and the overall efficiency of the sorting system is optimized.
It significantly improves the overall throughput capability and scheduling response speed of the logistics sorting system, realizes dynamic collaborative optimization between key performance indicators, and provides the theoretical basis of the smart logistics system.
Smart Images

Figure CN119990710B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of logistics sorting, and particularly to a method for optimizing the combined sorting of small-piece sorters in a logistics transfer yard. Background Art
[0002] With the continuous expansion of the scale of logistics transfer yards and the increasing diversification of types of automated equipment, there are significant differences in the sorting performance of different models of sorters. This difference leads to a low effective utilization rate of various sorters in the actual logistics sorting process.
[0003] At present, traditional scheduling schemes lack dynamic combined serial and parallel sorting strategies, and traditional methods are difficult to simultaneously consider high production capacity, high accuracy, and resource utilization rate. Moreover, there has not yet been a dynamic sorting strategy that can efficiently utilize the collaborative operation of various types of sorters. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for optimizing the combined sorting of small-piece sorters in a logistics transfer yard. By using digital twin technology, high-precision simulation and real-time feedback of the actual system are realized, and combined with the joint scheduling and adaptive optimization of the MAPPO algorithm in the MDP framework in multi-agent reinforcement learning, the response ability of the system to complex dynamic changes is effectively improved, and the overall efficiency of the logistics sorting system is significantly improved.
[0005] The purpose of the present invention is achieved through the following technical solutions: A method for optimizing the combined sorting of small-piece sorters in a logistics transfer yard, comprising the following steps:
[0006] The sorting methods are divided into serial sorting and parallel sorting, and sorting strategies are respectively constructed for serial sorting and parallel sorting; the sorting equipment under each sorting strategy includes a large-scale automated cross-belt sorter, a small-scale automated cross-belt sorter, and a manual sorting cabinet;
[0007] According to the number of packages within a sorting shift, a sorting strategy is selected from serial sorting and parallel sorting to perform package sorting;
[0008] Under the selected sorting strategy, a Markov decision process model is constructed for the logistics sorting scheduling problem and solved based on the multi-agent proximal policy optimization algorithm.
[0009] The beneficial effects of the present invention are as follows: The present invention constructs a scheduling architecture for a logistics sorting system supported by a multi-objective performance evaluation system and a digital twin platform. MDP models are respectively established for serial and parallel sorting scenarios, and the MAPPO multi-agent reinforcement learning algorithm is introduced to achieve dynamic collaborative optimization. This solution fully considers the balance among throughput, sorting accuracy, and labor cost between the pre-sorting and fine-sorting stages. Through real-time monitoring and data feedback of different equipment states, efficient collaborative scheduling of automated equipment and manual sorting resources is realized, effectively improving the overall throughput capacity and scheduling response speed of the system. Therefore, the present invention provides a new theoretical basis and technical method for the comprehensive scheduling and intelligent decision-making of logistics sorting systems, realizes the dynamic collaborative optimization among key performance indicators, and lays a solid theoretical and practical foundation for the development and application of subsequent intelligent logistics systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0011] The technical solution of the present invention will be further described in detail below with reference to the drawings, but the protection scope of the present invention is not limited to the following description.
[0012] Considering that traditional sorting scheduling methods are difficult to simultaneously take into account production capacity, accuracy, and resource utilization, and cannot respond in a timely manner to dynamic changes within a complex system, digital twin technology can achieve high-precision simulation and data feedback of the actual system, and multi-agent reinforcement learning (MARL) provides effective decision support for multi-device collaborative scheduling. The present invention first designs combined serial and parallel sorting strategies based on the sorting logics of different sorters, and with the help of the real-time feedback system state provided by the digital twin platform, uses the MAPPO algorithm to achieve combined sorting scheduling and adaptive optimization within the MDP framework, providing a new theoretical basis and technical path for this field. Specifically:
[0013] In a logistics transfer yard, considering that the problem we are targeting is the sorting of small-piece packages, common sorting equipment includes: large-scale automated cross-belt sorters, small-scale automated cross-belt sorters, and manual sorting cabinets. Among them, automated cross-belt sorters have the advantages of high-speed continuous transmission, automatic identification, and rapid classification, and are mainly used for preliminary high-volume sorting; while manual sorting cabinets are based on manual sorting and are suitable for meticulous and refined selection tasks.
[0014] As Figure 1 shown, a combined sorting optimization method for small-piece sorters in a logistics transfer yard includes the following steps:
[0015] The sorting methods are divided into serial sorting and parallel sorting, and sorting strategies are constructed for serial sorting and parallel sorting respectively; the sorting equipment under each sorting strategy includes large automated cross-belt sorters, small automated cross-belt sorters, and manual sorting cabinets;
[0016] Regarding the characteristics and sorting capabilities of the above different sorting equipment, we first design a combined serial and parallel sorting strategy:
[0017] I. Serial sorting strategy
[0018] In serial sorting, the logistics sorting is designed as multiple consecutive stages, including:
[0019] 1. Pre-sorting stage:
[0020] In practice, since the number of different package flows is often much larger than the sorting grid number of the automated cross-belt sorter, the clustering algorithm is first used to preliminarily classify the packages so that the packages with close destinations are grouped into a "mixed-sorted package" set. Considering that the processing capacity of the large cross-belt sorter is limited by its physical structure, we define:
[0021] a. State variable:
[0022] represents the state of the pre-sorting stage, including the package arrival rate, clustering result, and the queue length of the automated cross-belt sorter, where \(t\) represents the time.
[0023] b. State variable:
[0024] represents the preliminary clustering and task assignment decision for the package flow in the pre-sorting stage.
[0025] c. Pre-sorting throughput function:
[0026] The function represents the number of packages that can be processed at time t and is expressed as:
[0027]
[0028] where is the maximum number of packages that the large cross-belt sorter can process; is the number of packages with the package flow assigned under this large cross-belt sorter after clustering; is the conversion coefficient, reflecting the relationship between the current number of clustered packages and the actual processing capacity of the equipment.
[0029] 2. Fine sorting stage:
[0030] The pre-sorted mixed packages enter the fine sorting stage composed of a small automated cross-belt sorter and a manual sorting cabinet. This stage not only requires further subdivision and verification (to correct mis-sorting and over-circling), but also needs to consider the additional costs brought by manual sorting:
[0031] a. State variables:
[0032] To represent the state of the fine sorting stage, including the number of mis-sorted packages, over-circling situation, and load information of the manual sorting cabinet;
[0033] b. State variables:
[0034] Represents the task allocation and scheduling decision in the fine sorting stage, determining how many packages enter the manual secondary confirmation and the sorting strategy.
[0035] c. Fine sorting effective throughput function:
[0036]
[0037] Among them, Represents the maximum processing capacity of the sorting cabinet; Represents the number of packages that can be effectively processed based on the current scheduling decision; Is the conversion coefficient.
[0038] d. Manual sorting cost function:
[0039] To reflect the additional costs brought by the consumption of human resources, a cost function is introduced:
[0040]
[0041] Among them, Measures the amount of labor to be shared due to manual sorting at time t ; Is the fixed cost, Is the unit labor cost.
[0042] 3. Overall objective function of the serial system
[0043] The overall system requires that while ensuring the high-speed digestion of packages in pre-sorting, the final sorting accuracy is improved through fine sorting, and the control of labor costs is taken into account during this process. Therefore, the optimization objective of the serial sorting system is defined as maximizing the cumulative utility within a given time period, and its objective function can be written as
[0044]
[0045]
[0046] Among them, is the discount factor, is the trade-off coefficient of labor cost, and the constraint condition ensures that the number of packages processed in the fine sorting stage does not exceed the number of packages output in the pre-sorting stage, so as to maintain process consistency.
[0047] II. Parallel sorting strategy
[0048] In the parallel sorting strategy, all sorting devices (including large and small automated cross-belt sorters) are deployed as independent sorting units, and real-time dynamic task allocation is achieved through a preprocessing diversion system (such as using six-sided sweep bar code recognition). At the same time, parallel fine sorting modules are set up to handle mis-sorted and over-circled packages, and the labor sorting cost is taken into account.
[0049] 1. Automated sorting stage:
[0050] The preprocessing diversion system evenly distributes the packages arriving at the transfer yard according to the real-time traffic and the proportion of equipment bays. Denote the bays of large automated equipment and small automated equipment as and , and define the distribution ratio:
[0051]
[0052] a. State variable:
[0053] represents the real-time state of the automated equipment, including equipment load, task queue, and real-time processing rate.
[0054] b. Decision variable:
[0055] represents the dynamic task diversion decision based on the real-time state, ensuring balanced distribution according to the above ratio.
[0056] c. Automated sorting throughput function:
[0057]
[0058] Among them, and are the processing capabilities of large and small automated sorters respectively, which implies the role of decision in actual distribution.
[0059] 2. Fine sorting stage:
[0060] Although automated sorting can quickly process a large volume of packages, it may still produce mis-sorted and over-circled packages. Therefore, subsequent manual sorting cabinets are needed to correct the remaining error packages.
[0061] a. State variables:
[0062] Indicate the status of the fine sorting section, including the number of mis-sorted packages and the manual workload.
[0063] b. Decision variables:
[0064] Indicate the scheduling decision for the manual sorting cabinet, mainly determining the inflow of error packages to be processed and the manual sorting operation mode.
[0065] c. Fine sorting throughput function:
[0066]
[0067] Among them, Indicates the maximum processing capacity of the sorting cabinet; Indicates the number of packages that can be effectively processed based on the current scheduling decision; Is a conversion coefficient.
[0068] 3. Overall objective function of the parallel system
[0069] Integrating the automated sorting stage and the fine sorting stage, and introducing a penalty for manual sorting costs, the overall objective function of the system can be written as:
[0070]
[0071]
[0072] Among them, Is the penalty coefficient for labor costs, and the constraint condition ensures that the number of packages processed in the fine correction stage does not exceed the number of error packages generated in the automated stage.
[0073] When describing the state transition, decision variables and objective function, the above mathematical model not only considers the maximization of system throughput, but also adds penalties for resource consumption and labor costs, thus providing a solid theoretical basis for realizing the dynamic optimization under serial and parallel sorting strategies.
[0074] Select a sorting strategy from serial sorting and parallel sorting to sort packages according to the package volume within a sorting shift;
[0075] When the package volume within a sorting shift is less than the set threshold, select the serial sorting strategy to sort packages;
[0076] When the package volume within a sorting shift is not less than the set threshold, select the parallel sorting strategy to sort packages.
[0077] Under the selected sorting strategy (serial sorting or parallel sorting strategy), a Markov decision process model is constructed for the logistics sorting and scheduling problem, and it is solved based on the multi-agent proximal policy optimization algorithm.
[0078] The following presents the detailed Markov decision process (MDP) model constructed for the logistics sorting and scheduling problem under serial and parallel sorting strategies, and the complete solution process and mathematical formula description based on the multi-agent proximal policy optimization (MAPPO) algorithm. All the following content is developed at discrete time steps t = 0, 1, …, T −1, with the discount factor γ ∈ (0,1].
[0079] 1. MDP Modeling and MAPPO Solution under Serial Sorting Strategy
[0080] 1.1 MDP Modeling:
[0081] In the serial sorting system, the system consists of two consecutive stages: the pre-sorting stage and the fine-sorting stage. We define the state, action, transition, and reward functions of the entire system as follows.
[0082] a. State Space:
[0083] Let the system state at time t be
[0084]
[0085] where, represents the state of the pre-sorting stage, such as the parcel arrival rate, clustering results, the queue length of the automated cross-belt sorter, etc.;
[0086] represents the state of the fine-sorting stage, including information such as the number of mis-sorted parcels, over-loop situations, the load of the manual sorting cabinet, etc.
[0087] b. Action Space:
[0088] At each time t , the joint decision made by the system, i.e., the action, is
[0089]
[0090] where, represents the preliminary clustering and task assignment decision for the parcel flow direction in the pre-sorting stage; represents the task assignment and scheduling decision in the fine-sorting stage, determining how many parcels enter the manual secondary confirmation and the sorting strategy.
[0091] c. State Transition Function:
[0092] The state transition is affected by multiple factors such as the current queue, parcel arrival, equipment response, and manual operations. The overall transfer dynamics are written as
[0093]
[0094] d. Reward function:
[0095] In a serial system, at time t the reward function can be defined as
[0096]
[0097] The expected total return is
[0098]
[0099] 1.2 Solving MAPPO in Serial Policies
[0100] Although the serial system can be regarded as the centralized scheduling of the overall system, in actual operation, the two stages of pre-sorting and fine-sorting are often implemented by different "agents" or sub-modules. Here, a multi-agent reinforcement learning framework can be adopted. The system is divided into two agents, each agent is responsible for the scheduling of its respective stage, but shares the global state feedback. Let the policy of the agent in the pre-sorting stage be , and the policy of the agent in the fine-sorting stage be , where and are their respective local observations, and the action is .
[0101] a. Policy representation:
[0102] The policy of each agent i is parameterized as
[0103]
[0104] b. Importance sampling ratio:
[0105] For agent i at time t, the probability ratio is
[0106]
[0107] c. Advantage function estimation:
[0108] Using Generalized Advantage Estimation (GAE), for agent i, it is defined as
[0109]
[0110] where, , is the smoothing parameter, is the state value estimated by the centralized Critic network.
[0111] d. Shearing strategy objective:
[0112] MAPPO introduces the clipping objective of PPO to update the policy. For each agent i, its clipping objective function is
[0113]
[0114] where is the clipping threshold.
[0115] e. Value function update:
[0116] The value network adopts the mean squared error loss
[0117]
[0118] where the target value can be obtained through multi-step GAE.
[0119] f. Total loss function:
[0120] Combining policy update, value evaluation and entropy regularization term, the total loss of the system is written as:
[0121]
[0122] where are the weights of the value loss and the entropy regularization term respectively, and the entropy term is used to encourage policy exploration.
[0123] g. MAPPO process:
[0124] Data collection: First, in the simulation environment constructed by the digital twin platform, use the current policy and the global state feedback to collect the state , the corresponding local observation , the action , the reward and the next state .
[0125] Advantage evaluation: For the collected trajectory data, use the GAE method to calculate the advantage of each agent.
[0126] Policy update: Calculate the probability ratio of each agent and perform gradient ascent according to the clipping objective function to update the policy parameters ; At the same time, update the parameters of the centralized Critic network .
[0127] Iterative convergence: Repeat the data collection and parameter update process until convergence to obtain the optimal serial joint scheduling strategy.
[0128] 2. MDP Modeling and MAPPO Solving under the Parallel Sorting Strategy
[0129] 2.1 MDP Modeling
[0130] In the parallel strategy, all automated sorting devices participate in collaborative scheduling as multiple independent agents. At the same time, there is still an intervention from the manual sorting module to correct sorting errors. The system state is divided into an automated stage and a fine correction stage:
[0131] a. State Space
[0132] Let the global state be
[0133]
[0134] where describes the real-time state of automated devices (large and small automated cross-belt sorters), including device load, task queue, real-time processing rate, etc.; describes the state of the fine sorting part (manual sorting cabinet), such as the number of mis-sorted packages and the manual workload.
[0135] b. Action Space
[0136] Let the decision variable in the automated sorting stage be , representing the dynamic task diversion decision based on the real-time state, ensuring proportional and balanced distribution, that is, distributed according to the cross-belt sorter grid proportion allocation,
[0137] The automated sorting throughput function is:
[0138]
[0139] where and are the processing capabilities of large and small automated devices respectively.
[0140] The scheduling decision in the fine sorting stage is denoted as
[0141]
[0142] Its corresponding fine sorting throughput function is
[0143]
[0144] Meanwhile, the cost function generated by manual sorting still uses
[0145]
[0146] c. State transition
[0147]
[0148] d. Reward function
[0149] The reward function combines the two stages of automated sorting and fine sorting and is defined as
[0150]
[0151] Among them, is the penalty coefficient of the labor cost.
[0152] The overall optimization goal is
[0153]
[0154] That is, to solve the optimal joint strategy , corresponding to large-scale automation, small-scale automation, and manual scheduling agents respectively.
[0155] 2.2 Solving MAPPO in parallel policies
[0156] In a parallel system, each device (or device group) acts as an agent, and its local decision is based on local observations to determine the local action .
[0157] a. Policy representation
[0158] The policy of each agent i is parameterized as
[0159]
[0160] Among them, the action represents decisions such as the acceptance and diversion of sorted packages determined by the corresponding device.
[0161] b. Importance sampling ratio
[0162] For the probability ratio of agent i at time t is
[0163]
[0164] c. Advantage function estimation
[0165] Similarly, the generalized advantage estimation (GAE) is adopted and defined for agent i as
[0166]
[0167] Among them,
[0168]
[0169] d. Shearing target
[0170] The policy update target of each agent adopts the shearing policy objective function:
[0171]
[0172] e. Value function update
[0173] The centralized Critic network estimates the global state value function, and its mean square error loss is
[0174]
[0175] f. Total loss function
[0176]
[0177] g. Algorithm process
[0178] Utilize the current policy Interact with the global state feedback in the simulation environment constructed by digital twin to collect trajectory data , and calculate the advantage based on the trajectories sampled by each agent using the GAE method . Calculate the gradient according to their respective shearing objective functions and update the parameters using the Adam optimizer ; meanwhile, update the parameters of the centralized value network . Repeat the process of data collection and parameter update until the overall policy converges to obtain the optimal joint scheduling policy.
[0179] The above is the preferred implementation manner of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the technology or knowledge in related fields. And the changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall all be within the protection scope of the appended claims of the present invention.
Claims
1. A method for optimizing the combined sorting of small-piece sorting machines in a logistics transfer yard, characterized in that: It includes the following steps: The sorting methods are divided into serial sorting and parallel sorting. Sorting strategies are constructed respectively for serial sorting and parallel sorting; The sorting equipment under each sorting strategy includes a large automated cross-belt sorter, a small automated cross-belt sorter, and a manual sorting cabinet; The sorting strategy of the serial sorting includes: (1). Pre-sorting stage: This stage is carried out in the large cross-belt sorter. The clustering algorithm is used to preliminarily classify the packages, so that the packages with close destinations are grouped into a mixed-sorting package set. Define the parameters of this stage: State variable , representing the status of the pre-sorting stage, including the parcel arrival rate, clustering results, and the queue length of the automated cross-belt sorting machine; Decision variable , representing the initial clustering and task assignment decisions for the parcel flow direction in the pre-sorting stage; Pre-sorting throughput function , reflecting the relationship between the number of packages after current clustering and the actual processing capacity of the equipment; (2). Fine-sorting stage: The mixed-sorting packages after pre-sorting enter the fine-sorting stage composed of a combination of a small automated cross-belt sorter and a manual sorting cabinet. This stage requires further subdivision and verification, and the additional cost brought by manual sorting is considered. Define the parameters of this stage: State variable , representing the state of the fine sorting stage, including the number of misclassified packages, the situation of exceeding the circle, and the load information of the manual sorting cabinet; Decision variable , representing the task assignment and scheduling decision in the fine sorting stage, determining how many packages enter the manual secondary confirmation and the sorting strategy; Fine-sorting effective throughput function: Among them, represents the maximum processing capacity of the sorting cabinet; represents the number of packages effectively processed based on the current scheduling decision; is the conversion coefficient; Artificial sorting cost function , take the product of the labor amount to be shared in artificial sorting and the unit labor cost, and add the fixed total; (3). Overall objective function of the serial system: The optimization objective of the serial sorting system is defined as maximizing the cumulative utility within a given time period, and its objective function is: wherein, is the discount factor, is the trade-off coefficient of the labor cost, and the constraint condition ensures that the number of packages processed in the fine sorting stage does not exceed the number of packages output in the pre-sorting stage; The sorting strategy of the parallel sorting includes: In the parallel sorting strategy, all sorting equipment is deployed as independent sorting units, and real-time dynamic task allocation is realized through a preprocessing diversion system. At the same time, parallel fine-sorting modules are set up to handle mis-sorted and over-circled packages, considering the manual sorting cost, including: (1). Automated sorting stage: The preprocessing diversion system evenly distributes the parcels arriving at the transfer yard according to the real-time traffic and the proportion of the device compartments; denote the compartments of large automated devices and small automated devices as and , and define the distribution ratio: Define the following parameters: State variable Indicates the real-time status of the automation device, including device load, task queue, and real-time processing rate; Decision variable Represents the dynamic task diversion decision based on the real-time status to ensure balanced allocation according to the above ratio; Automated sorting throughput function: ; wherein, and are the processing capabilities of large and small automated sorting machines, respectively; (2). Fine-sorting stage: For mis-sorted and over-circled packages, the residual error packages are corrected through the manual sorting cabinet. Define the following parameters: Status variable Indicates the status of the fine sorting section, including the number of mis-sorted packages and the manual workload; Decision variable Represents the scheduling decision for the manual sorting cabinet, that is, the scheduling decision in the fine sorting stage, which determines the inflow of error packages to be processed and the manual sorting operation method; Fine sorting throughput function: Among them, represents the maximum processing capacity of the sorting cabinet; represents the number of packages that can be effectively processed based on the current scheduling decision; is the conversion coefficient; (3). Overall objective function of the parallel system: Combining the automated sorting stage and the fine-sorting stage, and introducing the penalty of manual sorting cost, the overall objective function of the system is written as: Among them, is the penalty coefficient of labor cost, and the constraint condition ensures that the number of packages processed in the fine correction stage does not exceed the number of error packages generated in the automation stage; According to the package volume within the sorting shift, select a sorting strategy from serial sorting and parallel sorting to sort the packages; Under the selected sorting strategy, construct a Markov decision process model for the logistics sorting and scheduling problem, and solve it based on the multi-agent proximal policy optimization algorithm.
2. A combined sorting optimization method for small-piece sorters in a logistics transfer yard according to claim 1, wherein: Selecting a sorting strategy from serial sorting and parallel sorting to sort the packages according to the package volume within the sorting shift includes: When the package volume within the sorting shift is less than the set threshold, select the serial sorting strategy to sort the packages; When the package volume within the sorting shift is not less than the set threshold, select the parallel sorting strategy to sort the packages.
3. An optimization method for combined sorting of small-piece sorting machines in a logistics transfer yard according to claim 1, characterized in that: When the selected sorting strategy is the serial sorting strategy, construct a Markov decision process model for the logistics sorting and scheduling problem, and solve it based on the multi-agent proximal policy optimization algorithm, including: A1. The constructed Markov decision process model: In the serial sorting system, the system consists of two consecutive stages: the pre-sorting stage and the fine-sorting stage. Define the state, action, transition, and reward functions of the whole system as follows: A101 State space: Let the system state at time t be Among them, represents the status of the pre-sorting stage, including the parcel arrival rate, clustering results, and the queue length of the automated cross-belt sorting machine; represents the status of the fine-sorting stage, including the number of mis-sorted parcels, over-circulation situations, and the load information of the manual sorting cabinets; A102 Action space: At each moment t , the joint decision taken by the system, i.e., the action is: Among them, represents the preliminary clustering and task assignment decision for the parcel flow direction in the pre-sorting stage; The task assignment and scheduling decision in the fine-sorting stage determines how many parcels enter the manual secondary confirmation and the sorting strategy; A103 State transition function: The state transition is affected by multiple factors such as current queuing, package arrival, equipment response, and manual operation. The overall transfer dynamics is written as: A104 Reward Function: In the serial system, at time t the reward function is defined as The expected total return is Among them, represents the expectation under the policy. A2. Define the relevant parameters for solving based on the multi-agent proximal policy optimization algorithm; The system is divided into two agents, each agent is responsible for the scheduling of its respective stage, but shares the global state feedback; let the policy of the pre-sorting stage agent be , and the policy of the fine-sorting stage agent be , where and are their respective local observations, and the action is ; A201 Policy Representation: The policy parameters of each agent i are parameterized as A202 Importance Sampling Ratio: For agent i at time t, the probability ratio is A203 Advantage Function Estimation: Using Generalized Advantage Estimation (GAE), define the advantage function for agent i: Among them, , is the smoothing parameter, is the state value estimated by the centralized Critic network; A204 Clipped Policy Objective: For each agent i, its clipped objective function is Among them, is the shear threshold; represents the expectation at time t; A205 Value Function Update: The value network adopts mean squared error loss: Among them, the target value can be obtained through multiple steps of GAE; A206 Total Loss Function: Combining policy update, value evaluation, and entropy regularization term, the total system loss is written as: Among them, are the weights of the value loss and the entropy regularization term, respectively, and the entropy term is used to encourage policy exploration.
4. A method for optimizing the combined sorting of small-item sorters in a logistics transfer yard according to claim 3, characterized in that: When the selected sorting strategy is the serial sorting strategy, the specific process of solving based on the multi-agent proximal policy optimization algorithm includes: In the simulation environment built using the digital twin platform, the current policy and global state feedback are used to collect the state , corresponding to the local observation , action , reward and the next state ; For the collected trajectory data, use the GAE method to calculate the advantage of each agent ; Calculate the probability ratio of each agent and perform gradient ascent according to the clipping objective function to update the policy parameters ; simultaneously update the parameters of the centralized Critic network ; Repeat the data collection and parameter update process until convergence to obtain the optimal serial joint scheduling strategy.
5. The combined sorting optimization method for small-piece sorters in a logistics transfer yard according to claim 1, wherein: When the selected sorting strategy is the parallel sorting strategy, construct a Markov decision process model for the logistics sorting and scheduling problem and solve it based on the multi-agent proximal policy optimization algorithm, including: B1. Constructed Markov Decision Process Model: All automated sorting devices in the parallel strategy participate in collaborative scheduling as multiple independent agents. At the same time, there is still an intervention by the manual sorting module to correct sorting errors; the system state is divided into the automated stage and the fine correction stage: B101 State Space: Let the global state be Among them, Describe the real-time status of the automated equipment, including equipment load, task queue, and real-time processing rate; Describe the status of the fine sorting part, including the number of mis-sorted packages and the manual workload; B102 Action Space Set decision variables for the automated sorting stage , representing the dynamic task diversion decision based on the real-time status, ensuring balanced distribution according to the ratio, that is, according to the ratio of the compartments of the cross-belt sorter allocate The automated sorting throughput function is: Among them, and are the processing capabilities of large and small automated devices respectively; The scheduling decision in the fine sorting stage is denoted as , and its corresponding fine sorting throughput function is At the same time, the cost function generated by manual sorting still adopts B103 State Transition B104 Reward Function The reward function combines the automated sorting and fine sorting stages and is defined as: Among them, is the penalty coefficient of labor cost; The overall optimization objective is That is, to solve the optimal joint strategy , corresponding to large-scale automation, small-scale automation, and manual scheduling agents respectively; among them, represents the strategy under the expectation; B2. Define the relevant parameters for solving based on the multi-agent proximal policy optimization algorithm; In a parallel system, each device serves as an agent, and its local decision is based on local observations to determine local actions ; B201 Policy Representation The policy of each agent i is parameterized as , where the action represents the decision on accepting and diverting sorted packages determined by the corresponding device; B202 Importance Sampling Ratio For agent i at time t, the probability ratio is B203 Advantage Function Estimation Using Generalized Advantage Estimation (GAE), the advantage function estimation for agent i is defined as where B204 Clipped Objective The policy update objective of each agent adopts the clipped policy objective function: Denote the expectation at time t; B205 Value Function Update The centralized Critic network estimates the global state value function, and its mean squared error loss is B206 Total Loss Function 。 6. A method for optimizing the combined sorting of small-piece sorting machines in a logistics transfer yard according to claim 5, characterized in that: When the selected sorting strategy is the parallel sorting strategy, the specific process of solving based on the multi-agent proximal policy optimization algorithm includes: Using the current policy Interact with the global state feedback in the simulation environment of digital twin construction to collect trajectory data and calculate the advantages using the GAE method based on the sampled trajectories of each agent ; Calculate the gradients according to their respective clipping objective functions and update the parameters using the Adam optimizer ; At the same time, update the parameters of the centralized value network Repeat the data collection and parameter update process until the overall policy converges to obtain the optimal joint scheduling policy
Citation Information
Patent Citations
Intelligent warehouse goods sorting system
CN119313082A