Combined sorting optimization method for small piece sorting machine in logistics transfer field

Through digital twin technology and multi-agent reinforcement learning algorithm, joint scheduling and adaptive optimization are achieved in the logistics sorting system, solving the shortcomings of the existing system in terms of taking into account production capacity, accuracy and resource utilization, and significantly improving system efficiency and response capabilities.

CN119990710AActive Publication Date: 2025-05-13THE CHINESE UNIV OF HONG KONG (SHENZHEN)

Patent Information

Application Number
CN202510467523.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing logistics sorting system has shortcomings in taking into account high productivity, high accuracy and resource utilization, and it is difficult to cope with complex dynamic changes, resulting in a low effective use rate of sorting machines.

Method used

Using digital twin technology combined with the MAPPO algorithm in multi-agent reinforcement learning (MARL), the joint scheduling and adaptive optimization of the logistics sorting system are realized under the framework of Markov decision-making process (MDP), and dynamic collaborative optimization of serial and parallel sorting strategies.

Benefits of technology

It significantly improves the overall efficiency of the logistics sorting system, improves the system's response ability to complex dynamic changes, realizes efficient coordinated scheduling of automation equipment and manual sorting resources, and improves throughput capability and scheduling response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990710A_ABST
    Figure CN119990710A_ABST
Patent Text Reader

Abstract

The invention discloses a combined sorting optimization method for a small piece sorting machine in a logistics transfer field, which comprises the following steps: sorting modes are divided into serial sorting and parallel sorting, and sorting strategies are respectively constructed for the serial sorting and the parallel sorting; selecting one sorting strategy from serial sorting and parallel sorting to sort the parcels according to the parcel quantity in the sorting shift; and under the selected sorting strategy, constructing a Markov decision process model for a logistics sorting scheduling problem, and solving based on a multi-agent near-end strategy optimization algorithm. According to the method, high-precision simulation and real-time feedback of an actual system are achieved through the digital twinborn technology, combined scheduling and self-adaptive optimization of the MAPPO algorithm in multi-agent reinforcement learning under an MDP framework are combined, the response capacity of the system to complex dynamic changes is effectively improved, and the overall efficiency of the logistics sorting system is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of logistics sorting, and in particular to a combined sorting optimization method for small-parts sorting machines in a logistics transfer yard. Background Art

[0002] With the continuous expansion of the scale of logistics transfer sites and the increasing diversification of automated equipment, there are large differences in sorting performance among different types of sorting machines. This difference leads to a low effective utilization rate of various types of sorting machines in the actual logistics sorting process.

[0003] At present, traditional scheduling schemes lack dynamic joint serial and parallel sorting strategies, and traditional methods are difficult to take into account high production capacity, high accuracy and resource utilization at the same time. At present, there is no dynamic sorting strategy that can efficiently utilize various types of sorting machines to work together. Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a joint sorting optimization method for small-piece sorting machines in logistics transfer yards. Digital twin technology is used to achieve high-precision simulation and real-time feedback of the actual system, and combined with the joint scheduling and adaptive optimization of the MAPPO algorithm in multi-agent reinforcement learning under the MDP framework, which effectively improves the system's response ability to complex dynamic changes and significantly improves the overall performance of the logistics sorting system.

[0005] The object of the present invention is achieved through the following technical solutions: A method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station, comprising the following steps: The sorting methods are divided into serial sorting and parallel sorting, and sorting strategies are constructed for serial sorting and parallel sorting respectively; the sorting equipment under each sorting strategy includes large-scale automatic cross-belt sorting machines, small-scale automatic cross-belt sorting machines and manual sorting cabinets; According to the parcel volume in the sorting shift, a sorting strategy is selected from serial sorting and parallel sorting to sort the parcels; Under the selected sorting strategy, a Markov decision process model is constructed for the logistics sorting scheduling problem, and it is solved based on the multi-agent proximal strategy optimization algorithm.

[0006] The beneficial effects of the present invention are as follows: the present invention constructs a logistics sorting system scheduling architecture based on a multi-objective performance evaluation system and supported by a digital twin platform, establishes MDP models for serial and parallel sorting scenarios respectively, and introduces a MAPPO multi-agent reinforcement learning algorithm to achieve dynamic collaborative optimization; the scheme fully considers the balance between throughput, sorting accuracy and labor costs between the two stages of pre-sorting and fine sorting, and realizes efficient collaborative scheduling of automated equipment and manual sorting resources through real-time monitoring and data feedback of different equipment status, effectively improving the overall throughput capacity and scheduling response speed of the system; thus, the present invention provides a new theoretical basis and technical method for comprehensive scheduling and intelligent decision-making of logistics sorting systems, realizes dynamic collaborative optimization between key performance indicators, and lays a solid theoretical and practical foundation for the subsequent development and application of smart logistics systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0008] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings, but the protection scope of the present invention is not limited to the following.

[0009] Considering that traditional sorting and scheduling methods are difficult to take into account production capacity, accuracy and resource utilization at the same time, and cannot respond to dynamic changes in complex systems in a timely manner, digital twin technology can achieve high-precision simulation and data feedback of actual systems, and multi-agent reinforcement learning (MARL) provides effective decision support for multi-device collaborative scheduling; the present invention first designs joint serial and parallel sorting strategies based on the sorting logic of different sorting machines, and with the help of the real-time feedback system status provided by the digital twin platform, the MAPPO algorithm is used to realize joint sorting scheduling and adaptive optimization under the MDP framework, providing a new theoretical basis and technical path for this field, specifically: In the logistics transfer yard, considering that the problem we are targeting is the logistics sorting of small parcels, common sorting equipment includes: large-scale automated cross-belt sorters, small-scale automated cross-belt sorters and manual sorting cabinets. Among them, the automated cross-belt sorting machine has the advantages of high-speed continuous transmission, automatic identification and rapid classification, and is mainly used for preliminary large-volume sorting; while the manual sorting cabinet is based on manual sorting and is suitable for detailed and refined selection tasks.

[0010] like Figure 1 As shown, a method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station includes the following steps: The sorting methods are divided into serial sorting and parallel sorting, and sorting strategies are constructed for serial sorting and parallel sorting respectively; the sorting equipment under each sorting strategy includes large-scale automatic cross-belt sorting machines, small-scale automatic cross-belt sorting machines and manual sorting cabinets; In view of the characteristics and sorting capabilities of the above different sorting equipment, we first designed a combined serial and parallel sorting strategy: 1. Serial sorting strategy In serial sorting, logistics sorting is designed into multiple consecutive stages, including: 1. Pre-sorting stage: In practice, since the number of different parcel flows is often much larger than the number of sorting slots of an automated cross-belt sorter, a clustering algorithm is first used to preliminarily classify the parcels so that parcels with similar destinations are grouped into a “mixed parcel” set. Considering that the processing capacity of a large cross-belt sorter is limited by its physical structure, we define: a. State variables: Indicates time t The sorting status of large-scale automated cross-belt sorters (including package arrival rate, clustering, current equipment occupancy, etc.).

[0011] b. State variables: It represents the sorting grid decision for the parcel flow distribution, and determines the route and priority of each mixed parcel set.

[0012] c. Pre-sorting throughput function: Use function Indicates at time t The number of packages that can be processed is expressed as: in, The maximum number of packages that can be processed by a large cross-belt sorter; The number of parcels allocated to the large cross-belt sorter for the parcel flow after clustering; is the conversion coefficient, which reflects the relationship between the number of packages after current clustering and the actual processing capacity of the device.

[0013] 2. Fine sorting stage: After pre-sorting, mixed parcels enter the fine sorting stage, which is composed of a small automatic cross-belt sorter and a manual sorting cabinet. This stage not only requires further segmentation and verification (to correct misclassification and over-circle phenomena), but also needs to consider the additional costs brought by manual sorting: a. State variables: For the moment tThe status of the fine sorting stage (including the number of missorted packages, the load of the manual sorting cabinet, etc.).

[0014] b. State variables: It represents the task allocation and scheduling decision in the fine sorting stage, which determines how many packages enter manual secondary confirmation and the sorting strategy.

[0015] c. Fine sorting effective throughput function: in, Indicates the maximum processing capacity of the sorting cabinet; represents the number of packages that can be effectively processed based on the current scheduling decision; is the conversion factor.

[0016] d. Manual sorting cost function: In order to reflect the additional cost caused by the consumption of human resources, the cost function is introduced: in, Measure at the moment t The amount of labor that needs to be shared due to manual sorting; is a fixed cost, The unit labor cost.

[0017] 3. Overall objective function of the serial system The overall system requires that while ensuring the high-speed digestion of pre-sorting packages, the final sorting accuracy rate should be improved through fine sorting, and the labor cost control should be taken into account in the process. Therefore, the optimization goal of the serial sorting system is defined as maximizing the cumulative utility within a given time period, and its objective function can be written as in, is the discount factor, The constraint is a trade-off factor for labor cost, and ensures that the number of parcels processed in the fine sorting stage cannot exceed the number of parcels output in the pre-sorting stage, thereby maintaining process consistency.

[0018] 2. Parallel sorting strategy In the parallel sorting strategy, all sorting equipment (including large and small automated cross-belt sorters) are deployed as independent sorting units, and real-time dynamic task allocation is achieved through a pre-processing diversion system (for example, using six-sided barcode recognition). At the same time, a parallel fine sorting module is set up to handle missorted and over-circled packages, and the manual sorting costs are taken into account.

[0019] 1. Automated sorting stage: The pre-processing and diversion system distributes the parcels arriving at the transfer site in a balanced manner according to the real-time flow and equipment grid ratio. The grids of large-scale automation equipment and small-scale automation equipment are and , define the allocation ratio: a. State variables: Indicates the real-time status of each device in the automated sorting stage, including processing rate, load level and error rate.

[0020] b. Decision variables: It represents the dynamic task diversion decision based on real-time status to ensure balanced distribution according to the above ratio.

[0021] c. Automated sorting throughput function: in, and The processing capabilities of large and small automated sorting machines are respectively, and the decision Role in actual allocation.

[0022] 2. Fine sorting stage: Although automated sorting can quickly process a large volume of parcels, it may still result in misclassification or out-of-circle parcels, so subsequent manual sorting cabinets are required to correct the remaining erroneous parcels.

[0023] a. State variables: Indicates the status of the fine sorting stage, including the number of missorted packages, the working status of the manual sorting cabinet, etc.

[0024] b. Decision variables: It represents the scheduling decision for the manual sorting cabinet, which mainly determines the inflow of error packages to be processed and the manual sorting operation mode.

[0025] c. Fine sorting throughput function: in, Indicates the maximum processing capacity of the sorting cabinet; represents the number of packages that can be effectively processed based on the current scheduling decision; is the conversion factor.

[0026] 3. Overall objective function of parallel system Combining the automated sorting stage and the fine sorting stage, and introducing the manual sorting cost penalty, the overall objective function of the system can be written as: in, is the penalty coefficient for labor cost, and the constraint ensures that the number of packages processed in the fine correction stage cannot exceed the number of error packages generated in the automation stage.

[0027] The above mathematical model not only takes into account the maximization of system throughput when describing state transitions, decision variables and objective functions, but also adds penalties for resource consumption and labor costs, thus providing a solid theoretical basis for achieving dynamic optimization under serial and parallel sorting strategies.

[0028] According to the parcel volume in the sorting shift, a sorting strategy is selected from serial sorting and parallel sorting to sort the parcels; When the number of parcels in a sorting shift is less than the set threshold, the serial sorting strategy is selected to sort the parcels; When the number of parcels in a sorting shift is not less than the set threshold, a parallel sorting strategy is selected to sort the parcels.

[0029] Under the selected sorting strategy (serial sorting or parallel sorting strategy), a Markov decision process model is constructed for the logistics sorting scheduling problem, and it is solved based on the multi-agent proximal strategy optimization algorithm.

[0030] The following is a detailed Markov decision process (MDP) model for the logistics sorting scheduling problem under serial and parallel sorting strategies, and a complete solution process and mathematical formula description based on the multi-agent proximal policy optimization (MAPPO) algorithm. The following content is described in discrete time steps. t = 0, 1, …, T −1, the discount factor is γ ∈ (0,1].

[0031] 1. MDP modeling and MAPPO solution under serial sorting strategy 1.1 MDP Modeling: In the serial sorting system, the system consists of two consecutive stages: the pre-sorting stage and the fine sorting stage. We define the state, action, transition and reward function of the entire system as follows.

[0032] a. State space: Order Moment t The system status is in, Indicates the status of the pre-sorting stage, such as package arrival rate, clustering results, queue length of the automated cross-belt sorter, etc. Indicates the status of the fine sorting stage, including the number of missorted packages, over-circle conditions, load of manual sorting cabinets, etc.

[0033] b. Action space: At every moment t , the joint decision taken by the system is in, Represents the preliminary clustering and task allocation decisions for package flow in the pre-sorting stage; Represents the task scheduling decision in the fine sorting stage, which determines how many packages enter the manual sorting cabinet and how to allocate correction tasks.

[0034] c. State transfer function: State transition is affected by multiple factors such as the current queue, package arrival, device response, and manual operation. Overall transition dynamic writing d. Reward function: In a serial system, at time t The reward function can be defined as The expected total return is 1.2 MAPPO solution for serial strategy Although the serial system can be regarded as the centralized scheduling of the whole system, in actual operation, the pre-sorting and fine sorting stages are often implemented by different "agents" or submodules. Here, a multi-agent reinforcement learning framework can be used to divide the system into two agents, each of which is responsible for the scheduling of its own stage, but shares the global state feedback. Suppose the strategy of the agent in the pre-sorting stage is , the strategy of the agent in the fine sorting stage is ,in and are their respective local observations, and the overall action is .

[0035] a. Strategy representation: The strategy of each agent i is parameterized as b. Importance sampling ratio: The probability ratio for agent i at time t is c. Advantage function estimation: Using the generalized advantage estimation (GAE), for agent i, we define in, , is the smoothing parameter, is the state value estimated by the centralized Critic network.

[0036] d. Cutting strategy objectives: MAPPO introduces the clipping objective of PPO to update the strategy. For each agent i, its objective function is in, is the clipping threshold.

[0037] e. Value function update: The value network uses mean square error loss Among them, the target value It can be obtained through multi-step GAE.

[0038] f. Total loss function: Combining strategy update, value assessment and entropy regularization, the total system loss is written as: in, are the weights of the value loss and entropy regularization term, respectively, and the entropy term Used to encourage strategic exploration.

[0039] g.MAPPO process: Data collection: First, use the current strategy in the simulation environment built by the digital twin platform Feedback with global status, collection status , corresponding to the local observation ,action ,award and the next state .

[0040] Advantage evaluation: Based on the collected trajectory data, the GAE method is used to calculate the advantage of each agent. .

[0041] Policy update: Calculate the probability ratio for each agent And according to the cutting objective function Perform gradient ascent and update strategy parameters ; At the same time, update the centralized Critic network parameters .

[0042] Iterative convergence: Repeat the data collection and parameter update process until convergence to obtain the optimal serial joint scheduling strategy.

[0043] 2. MDP modeling and MAPPO solution under parallel sorting strategy 2.1 MDP Modeling In the parallel strategy, all automated sorting equipment participates in collaborative scheduling as multiple independent intelligent agents, while manual sorting modules intervene to correct sorting errors. The system state is divided into the automation stage and the fine correction stage: a. State space Let the global state be in, Describe the real-time status of automated equipment (large and small automated cross-belt sorters), including equipment load, task queue, real-time processing rate, etc. Describes the status of the fine sorting part (manual sorting cabinet), such as the number of missorted packages and manual workload.

[0044] b. Action Space The scheduling decision in the automated sorting phase is recorded as To achieve proportional distribution of tasks, that is, according to the proportion of the cross-belt sorter slots Allocation, the actual allocation relationship is: in, and Respectively, the processing capabilities of large and small automation equipment.

[0045] The scheduling decision in the fine sorting stage is recorded as The corresponding effective throughput function is At the same time, the cost function generated by manual sorting still uses c. State transfer d. Reward function The reward function integrates the two stages of automatic sorting and fine sorting and is defined as in, is the penalty coefficient for labor cost.

[0046] The overall optimization goal is That is, solving the optimal joint strategy , corresponding to large-scale automation, small-scale automation and manual scheduling agents respectively.

[0047] 2.2 MAPPO in parallel strategy solution In a parallel system, each device (or device group) acts as an intelligent agent, and its local decision-making is based on local observations. To determine local actions . a. Policy representation The strategy of each agent i is parameterized as The action Indicates the decisions on acceptance and diversion of sorting packages determined by the corresponding equipment.

[0048] b. Importance sampling ratio The probability ratio for agent i at time t is c. Advantage function estimation We also use the generalized advantage estimation (GAE) to define agent i as in, d. Cut target The strategy update target of each agent adopts the cutting strategy objective function: e. Value function update The centralized critic network estimates the global state value function, and its mean square error loss is f. Total loss function g. Algorithm flow Leverage current strategies Interact with the global state feedback in the simulation environment constructed by the digital twin and collect trajectory data }, and use the GAE method to calculate the advantage based on the sampled trajectories of each agent . According to their respective cutting objective functions Calculate the gradient and update the parameters using the Adam optimizer ; At the same time, update the centralized value network parameters The data collection and parameter updating process is repeated until the overall strategy converges and the optimal joint scheduling strategy is obtained.

[0049] The above is a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art do not depart from the spirit and scope of the present invention, and should be within the scope of protection of the claims attached to the present invention.

Claims

1. A method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station, characterized in that: The following steps are involved: The sorting methods are divided into serial sorting and parallel sorting, and sorting strategies are constructed for serial sorting and parallel sorting respectively; the sorting equipment under each sorting strategy includes large-scale automatic cross-belt sorting machines, small-scale automatic cross-belt sorting machines and manual sorting cabinets; According to the parcel volume in the sorting shift, a sorting strategy is selected from serial sorting and parallel sorting to sort the parcels; Under the selected sorting strategy, a Markov decision process model is constructed for the logistics sorting scheduling problem, and it is solved based on the multi-agent proximal strategy optimization algorithm.

2. According to claim 1, a combined sorting optimization method for small-parts sorting machines in a logistics transfer station is characterized by: The sorting strategy of the serial sorting includes:

1. Pre-sorting stage: This stage is carried out in a large cross-belt sorter. The clustering algorithm is used to preliminarily classify the packages so that the packages with similar destinations are grouped into a mixed package set. The parameters of this stage are defined as follows: State variables Indicates time t The status of sorting by large-scale automated cross-belt sorter; Decision variables It represents the sorting grid decision for the parcel flow distribution, and determines the route and priority of each mixed parcel set; Pre-sorting throughput function , which reflects the relationship between the number of packages after current clustering and the actual processing capacity of the device; (II) Fine sorting stage: After pre-sorting, the mixed parcels enter the fine sorting stage, which is composed of a small automatic cross-belt sorter and a manual sorting cabinet. This stage requires further segmentation and verification, and the additional cost of manual sorting is considered. The parameters of this stage are defined as follows: State variables , indicating the time t The status of the next fine sorting stage; Decision variables , represents the task allocation and scheduling decision in the fine sorting stage, which determines how many packages enter manual secondary confirmation and the sorting strategy; Fine sorting effective throughput function: in, Indicates the maximum processing capacity of the sorting cabinet; Indicates the number of packages effectively processed based on the current dispatch decision; is the conversion factor; Manual sorting cost function , take the product of the labor amount required for manual sorting and the unit labor cost, and add the fixed sum; (III) Overall objective function of the serial system: The optimization goal of the serial sorting system is defined as maximizing the cumulative utility within a given time period, and its objective function is: in, is the discount factor, is the trade-off coefficient for labor cost, and the constraint ensures that the number of parcels processed in the fine sorting stage cannot exceed the number of parcels output in the pre-sorting stage.

3. The method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station according to claim 1 is characterized in that: The sorting strategy of the parallel sorting includes: In the parallel sorting strategy, all sorting equipment is deployed as independent sorting units, and real-time dynamic task allocation is achieved through the pre-processing diversion system. At the same time, a parallel fine sorting module is set up to handle missorted and over-circle packages, taking into account the manual sorting costs, including: (I) Automated sorting stage: The pre-processing and diversion system distributes the parcels arriving at the transfer site evenly according to the real-time flow and equipment grid ratio; the grids of large-scale automation equipment and small-scale automation equipment are and , define the allocation ratio: Define the following parameters: State variables Indicates the real-time status of each device in the automated sorting stage, including processing rate, load level and error rate; Decision variables It represents the dynamic task diversion decision based on real-time status to ensure balanced distribution according to the above ratio; Automated sorting throughput function: ; in, and The processing capacity of large and small automated sorting machines respectively; (II) Fine sorting stage: For packages that are misclassified or out of order, the remaining erroneous packages are corrected through manual sorting cabinets, and the following parameters are defined: State variables Indicates the status of the fine sorting stage, including the number of missorted packages and the working status of the manual sorting cabinet; Decision variables It represents the scheduling decision for the manual sorting cabinet, which determines the inflow of error packages to be processed and the manual sorting operation mode; Fine sorting throughput function: in, Indicates the maximum processing capacity of the sorting cabinet; represents the number of packages that can be effectively processed based on the current scheduling decision; is the conversion factor; (III) Overall objective function of parallel system: Combining the automated sorting stage and the fine sorting stage, and introducing the manual sorting cost penalty, the overall objective function of the system is written as: in, is the penalty coefficient for labor cost, and the constraint ensures that the number of packages processed in the fine correction stage cannot exceed the number of error packages generated in the automation stage.

4. The method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station according to claim 1 is characterized in that: According to the parcel volume within the sorting shift, a sorting strategy is selected from serial sorting and parallel sorting to sort the parcels, including: When the number of parcels in a sorting shift is less than the set threshold, the serial sorting strategy is selected to sort the parcels; When the number of parcels in a sorting shift is not less than the set threshold, a parallel sorting strategy is selected to sort the parcels.

5. The method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station according to claim 2 is characterized in that: When the selected sorting strategy is the serial sorting strategy, a Markov decision process model is constructed for the logistics sorting scheduling problem, and the solution is based on the multi-agent proximal strategy optimization algorithm, including: A1. Constructed Markov decision process model: In the serial sorting system, the system consists of two consecutive stages: the pre-sorting stage and the fine sorting stage. The state, action, transfer and reward function of the entire system are defined as follows: A101 State Space: Order Moment t The system status is in, Indicates the status of the pre-sorting stage, including the package arrival rate, clustering results, and queue length of the automated cross-belt sorter; Indicates the status of the fine sorting stage, including the number of missorted packages, over-circle conditions, and load information of the manual sorting cabinet; A102 Action Space: At every moment t , the joint decision taken by the system is: in, Represents the preliminary clustering and task allocation decisions for package flow in the pre-sorting stage; Represents the task scheduling decision in the fine sorting stage, which determines how many packages enter the manual sorting cabinet and how to allocate the correction tasks; A103 state transfer function: The state transition is affected by multiple factors such as the current queue, package arrival, device response, and manual operation. The overall transition dynamics are written as: A104 reward function: In a serial system, at time t The reward function is defined as The expected total return is in, Representation strategy expectations; A2. Define the relevant parameters for solving the multi-agent proximal strategy optimization algorithm; The system is divided into two agents, each of which is responsible for the scheduling of its own stage, but shares the global state feedback; the strategy of the agent in the pre-sorting stage is , the strategy of the agent in the fine sorting stage is ,in and are their respective local observations, and the overall action is ; A201 Strategy says: The strategy of each agent i is parameterized as A202 Importance Sampling Ratio: The probability ratio for agent i at time t is A203 advantage function estimation: Using the generalized advantage estimate GAE, the advantage function is defined for agent i: in, , is the smoothing parameter, is the state value estimated by the centralized Critic network; A204 shearing strategy objectives: For each agent i, its objective function is in, is the clipping threshold; represents the expectation at time t; A205 value function update: The value network uses mean square error loss: Among them, the target value It can be obtained through multi-step GAE; A206 total loss function: Combining strategy update, value assessment and entropy regularization, the total system loss is written as: in, are the weights of the value loss and entropy regularization term, respectively, and the entropy term Used to encourage strategic exploration.

6. The method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station according to claim 5 is characterized by: When the selected sorting strategy is a serial sorting strategy, the specific process of solving the problem based on the multi-agent proximal strategy optimization algorithm includes: In the simulation environment built by the digital twin platform, the current strategy is used Feedback with global status, collection status , corresponding to the local observation ,action ,award and the next state ; For the collected trajectory data, the GAE method is used to calculate the advantage of each agent ; Calculate the probability ratio for each agent And according to the cutting objective function Perform gradient ascent and update strategy parameters ; At the same time, update the centralized Critic network parameters ; The data collection and parameter updating process is repeated until convergence to obtain the optimal serial joint scheduling strategy.

7. The method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station according to claim 3 is characterized by: When the selected sorting strategy is a parallel sorting strategy, a Markov decision process model is constructed for the logistics sorting scheduling problem, and the solution is based on the multi-agent proximal strategy optimization algorithm, including: B1. Constructed Markov decision process model: In the parallel strategy, all automated sorting equipment participates in collaborative scheduling as multiple independent intelligent agents. At the same time, there is still a manual sorting module involved to correct the sorting error; the system state is divided into the automation stage and the fine correction stage: B101 state space: Let the global state be in, Describe the real-time status of automation equipment, including equipment load, task queue, and real-time processing rate; Describe the status of the fine sorting part, including the number of missorted packages and the manual workload; B102 Action Space The scheduling decision in the automated sorting phase is recorded as , used to realize the proportional distribution of tasks, that is, according to the proportion of the cross-belt sorter slots Allocation, the actual allocation relationship is: in, and The processing capacity of large and small automation equipment respectively; The scheduling decision in the fine sorting stage is recorded as , and its corresponding effective throughput function is At the same time, the cost function generated by manual sorting still uses B103 state transfer B104 Reward Function The reward function integrates the two stages of automated sorting and fine sorting and is defined as: in, is the penalty coefficient of labor cost; The overall optimization goal is That is, solving the optimal joint strategy , corresponding to large-scale automation, small-scale automation and manual scheduling agents respectively; among them, Representation strategy expectations; B2. Define the relevant parameters for solving the multi-agent proximal strategy optimization algorithm; In a parallel system, each device acts as an intelligent agent, and its local decision-making is based on local observations. To determine local actions ; B201 Strategy Representation The strategy of each agent i is parameterized as , where the action Indicates the sorting package acceptance and diversion decision determined by the corresponding equipment; B202 Importance Sampling Ratio The probability ratio for agent i at time t is B203 Advantage Function Estimation Using generalized advantage estimation GAE, the advantage function estimate for agent i is defined as in, B204 Cutting Target The strategy update target of each agent adopts the cutting strategy objective function: represents the expectation at time t; B205 Value Function Update The centralized critic network estimates the global state value function, and its mean square error loss is B206 total loss function 。 8. The method for optimizing the combined sorting of small-part sorting machines in a logistics transfer station according to claim 7, characterized in that: When the selected sorting strategy is a parallel sorting strategy, the specific process of solving the problem based on the multi-agent proximal strategy optimization algorithm includes: Leverage current strategies Interact with the global state feedback in the simulation environment constructed by the digital twin and collect trajectory data }, and use the GAE method to calculate the advantage based on the sampled trajectories of each agent ; According to their respective cutting objective functions Calculate the gradient and update the parameters using the Adam optimizer ; At the same time, update the centralized value network parameters , repeat the data collection and parameter updating process until the overall strategy converges and the optimal joint scheduling strategy is obtained.

Citation Information

Patent Citations

  • Movable sorting station, dynamic adjusting method and device thereof, and storage medium

    CN109848050A

  • Logistics sorting method fusing combinatorial optimization and reinforcement learning accelerated convergence

    CN117556959A

  • Intelligent warehouse goods sorting system

    CN119313082A

Cited By

  • Secondary sorting method and system for plane sorting

    CN121119874A

  • Man-machine collaborative dynamic scheduling method fused with double-layer optimization mechanism

    CN121526196A