Method, device, equipment and product for cooperation among multiple storage robots

CN122074067APending Publication Date: 2026-05-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2024-09-20
Publication Date
2026-05-22

Smart Images

  • Figure CN122074067A_ABST
    Figure CN122074067A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device, equipment and a product for cooperation among a plurality of storage robots. The method includes acquiring task information (112) by the picking robot (102). The method further includes acquiring, by the picking robot, a plurality of position information of the plurality of transport robots (104, 106, 108). The method further includes determining, by the picking robot, a target transport robot from the plurality of transport robots based on the plurality of location information. The method further includes sending, by the picking robot, transport task information to the target transport robot (120). In addition, the method includes handing over, by the picking robot, the picked target item (114) to the target transport robot at a handing-over location (122). In this way, the sorting robot can transfer the articles to the conveying robot immediately after completing the grabbing task, the articles do not need to be conveyed to a farther target position by themselves, in this way, the next sorting task can be started more quickly, and therefore the utilization rate of the sorting assembly can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device, and product for collaboration among multiple warehouse robots TECHNICAL FIELD

[0001] The present disclosure relates to the field of robotics, and more specifically to a method, apparatus, device, and product for collaboration among multiple warehouse robots. BACKGROUND

[0002] Picking robots and transport robots are two core robots in modern warehouse systems. Picking robots are capable of utilizing advanced vision systems (e.g., cameras, lidar, etc.) to quickly and accurately identify information such as the location, shape, size, and barcode of an item. For example, a vision system trained by a deep learning algorithm can accurately distinguish between different types of items in a complex warehouse environment, even if the items have similar packaging colors and patterns. Picking robots can be equipped with various types of end effectors, such as robotic arms, grippers, suction cups, etc., to accommodate the grasping needs of items of different shapes and materials.

[0003] Transport robots are typically equipped with various navigation technologies such as laser navigation, vision navigation, or magnetic navigation, and are capable of autonomously planning paths in a warehouse environment, avoiding obstacles, and safely and efficiently transporting items to designated locations. Transport robots can be designed with different load specifications to meet the transportation needs of various items, enabling the transportation of everything from small parts to large pallets of goods.

[0004] SUMMARY

[0005] In a first aspect of embodiments of the present disclosure, a method for collaboration among multiple warehouse robots is provided, wherein the multiple warehouse robots include a picking robot and multiple transport robots. The method includes obtaining, by the picking robot, task information, the task information including target item information and a target location to which the target item is to be transported. The method further includes obtaining, by the picking robot, multiple location information of the multiple transport robots. The method further includes determining, by the picking robot, a target transport robot from the multiple transport robots based on the multiple location information. The method further includes sending, by the picking robot, transport task information to the target transport robot, the transport task information including the target item information, a handover location, and the target location. In addition, the method further includes handing over, by the picking robot, the picked target item to the target transport robot at the handover location.

[0006] In a second aspect of the embodiments of the present disclosure, an apparatus for cooperation among multiple warehousing robots is provided. The apparatus includes a task information obtaining module configured to obtain, by a picking robot, task information, the task information including target item information and a target location to which a target item is transported. The apparatus further includes a location information obtaining module configured to obtain, by the picking robot, multiple location information of multiple transport robots. The apparatus further includes a transport robot determining module configured to determine, by the picking robot, a target transport robot from the multiple transport robots based on the multiple location information. The apparatus further includes a transport task sending module configured to send, by the picking robot, transport task information to the target transport robot, the transport task information including the target item information, a handover location, and the target location. In addition, the apparatus further includes a target item handover module configured to hand over, by the picking robot, the picked target item to the target transport robot at the handover location.

[0007] In a third aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device storing one or more programs, when the one or more programs are executed by the one or more processors, cause the one or more processors to implement a method for cooperation among multiple warehousing robots, wherein the multiple warehousing robots include a picking robot and multiple transport robots. The method includes obtaining, by the picking robot, task information, the task information including target item information and a target location to which a target item is transported. The method further includes obtaining, by the picking robot, multiple location information of the multiple transport robots. The method further includes determining, by the picking robot, a target transport robot from the multiple transport robots based on the multiple location information. The method further includes sending, by the picking robot, transport task information to the target transport robot, the transport task information including the target item information, a handover location, and the target location. In addition, the method further includes handing over, by the picking robot, the picked target item to the target transport robot at the handover location.

[0008] In a fourth aspect of embodiments of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer readable medium and includes machine executable instructions that, when executed, cause a machine to implement a method for collaboration among a plurality of warehouse robots, where the plurality of warehouse robots includes a picking robot and a plurality of transport robots. The method includes obtaining, by the picking robot, task information, the task information including target item information and a target location to which the target item is to be transported. The method further includes obtaining, by the picking robot, a plurality of location information of the plurality of transport robots. The method further includes determining, by the picking robot, a target transport robot from the plurality of transport robots based on the plurality of location information. The method further includes sending, by the picking robot, transport task information to the target transport robot, the transport task information including the target item information, a handoff location, and the target location. In addition, the method further includes handing off, by the picking robot, the picked target item to the target transport robot at the handoff location.

[0009] The summary is provided to introduce a selection of concepts, in a simplified form, that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent when described in conjunction with the following detailed description in conjunction with the accompanying drawings. In the drawings, like or similar elements are referred to with like or similar reference numerals, in which:

[0011] FIG. 1 illustrates a schematic diagram of an example environment in which a plurality of embodiments of the present disclosure can be implemented;

[0012] FIG. 2 illustrates a flowchart of a method for collaboration among a plurality of warehouse robots, according to some embodiments of the present disclosure;

[0013] FIG. 3 illustrates a flowchart of an example process for determining a target transport robot from a plurality of transport robots, according to some embodiments of the present disclosure;

[0014] FIG. 4 illustrates a schematic diagram of an example of determining a target transport robot from a plurality of transport robots, according to some embodiments of the present disclosure;

[0015] FIG. 5 illustrates a schematic diagram of an example of a handoff location determination model, according to some embodiments of the present disclosure;

[0016] FIG. 6 illustrates a flowchart of an example process for training a handoff location determination model, according to some embodiments of the present disclosure;

[0017] FIG. 7 shows a block diagram of an apparatus for cooperation of multiple warehouse robots, according to some embodiments of the present disclosure;

[0018] FIG. 8 shows a flowchart of a method for determining a handoff location of a picking robot and a transport robot, according to some embodiments of the present disclosure; and

[0019] FIG. 9 shows a block diagram of a device capable of implementing a number of embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] It can be understood that all user-related data involved in the technical solution should be obtained and used after the user's authorization. This means that in the technical solution, if the user's personal information needs to be used, the user's explicit consent and authorization are required before obtaining these data, otherwise the relevant data collection and use will not be carried out. It should also be understood that in the implementation of the technical solution, relevant laws and regulations should be strictly followed in the process of data collection, use and storage, and necessary technical and measures should be taken to protect the user's data security and ensure the safe use of data.

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood as open-ended, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects unless explicitly stated otherwise. The following can also include other explicit and implicit definitions.

[0023] In existing warehouse systems, some robots have both picking and transporting capabilities, which are responsible for picking items from shelves and transporting them to designated locations (e.g., packing areas or sorting areas). However, such multi-functional robots can only perform tasks one by one. For example, after completing the picking of an item, the robot needs to go to the designated location for storage alone and then return to continue the next picking task. In this process, a large amount of time is consumed in the round trip, so that the picking components (e.g., mechanical arms, grippers, etc.) of the robots cannot be fully utilized. In addition, for large-scale warehouse businesses, such robots are difficult to cope with high-concurrency task requirements. During the peak period of orders, robots may cause goods to pile up and prolong the processing time of orders due to the inability to timely handle a large number of picking and transporting tasks, thereby reducing user experience.

[0024] To this end, embodiments of the present disclosure provide a multi-robot collaboration scheme. Multi-robot collaboration refers to the process in which multiple robots work together to complete a specific task in a warehouse space. Such collaboration can occur between robots of the same type or a combination of different types of robots. The key to multi-robot collaboration is that they can effectively communicate, coordinate actions, allocate tasks, and adapt to changes in uncertain environments.

[0025] In the scheme provided by the present disclosure, a picking robot and multiple transporting robots are included. The picking robot can obtain task information, which includes target item information and a target location to which the target item is transported. The picking robot can also obtain multiple location information of the multiple transporting robots. The picking robot can determine a target transporting robot from the multiple transporting robots based on the multiple location information. Then, the picking robot can send transportation task information to the target transporting robot, the transportation task information including the target item information, a handover location, and the target location. Then, the picking robot can hand over the picked target item to the target transporting robot at the handover location.

[0026] In this way, the picking and transporting functions of the robots can be decoupled, so that the picking robot and the transporting robot are responsible for different tasks, i.e., the picking robot can focus on picking items from shelves, and the transporting robot can focus on transporting items. Through collaboration, the picking robot can immediately hand over the item to the transporting robot after completing the picking task, without personally transporting the item to a distant target location, which can enable the next picking task to be started faster, thereby improving the utilization rate of the picking components.

[0027] FIG. 1 illustrates a schematic diagram of an example environment 100 in which various embodiments of the present disclosure can be implemented. As shown in FIG. 1, in the environment 100 includes a picking robot 102, a transport robot 104, a transport robot 106, a transport robot 108, and a control center 110. It should be understood that for brevity, only three transport robots are shown in FIG. 1, but it is not intended to limit the number of transport robots, and any number of transport robots can be included in other implementations. In the environment 100, the picking robot 102 has both picking capability and transport capability. The control center 110 can send task information 112 to the picking robot 102, where the task information 112 can include information associated with a target item 114 (e.g., item identification, item name, item category, corresponding shelf, etc.) and a target location 116 to which the target item 114 needs to be transported.

[0028] In related art, after receiving the task information 112, the picking robot 102 can move to a shelf 118 where the target item 114 is located to pick the target item 114, and then transport it to the target location 116. However, since the target location 116 can be far away from the shelf 118, the picking robot 102 needs to spend more time on transporting the item rather than picking the item, which reduces the utilization of the picking components (e.g., mechanical arm, gripper, etc.) of the picking robot 102, and thus the picking capability of the picking robot 102 cannot be fully utilized. In the embodiments provided by the present disclosure, after receiving the task information 112, the picking robot 102 can autonomously collaborate with a transport robot to jointly complete the task of transporting the target item 14 to the target location 116.

[0029] As shown in FIG. 1, the picking robot 102 can obtain position information of a plurality of transport robots in the nearby area. For example, in the environment 100, there are the transport robot 104, the transport robot 106, and the transport robot 108 in the vicinity of the picking robot 102. For example, the nearby transport robots can be determined by a predetermined distance threshold. The picking robot 102 can obtain the position information of these transport robots. The position information can include, for example, current coordinates representing the exact position of the transport robot in the warehouse, which can be two-dimensional coordinates representing the specific position of the transport robot in the warehouse plane, or three-dimensional coordinates representing the position of the transport robot in some multi-layer warehouse or vertical stacking system. The position information can be obtained by the navigation system of the transport robot itself and shared in real time through a wireless network or the control center 110.

[0030] In the environment 100, the picking robot 102 can determine a target transport robot to collaborate with from the transport robots 104, 106, and 108 based on the position information of the transport robots. In some embodiments of the present disclosure, the target transport robot can be determined by considering the sum of distances from the picking robot 102 to each of the transport robots to reach the handoff location. In some embodiments, the target transport robot can be determined by considering the time difference (i.e., the waiting time) from the picking robot 102 to each of the transport robots to reach the handoff location. In some embodiments, the target transport robot can be determined by considering the sum of energy consumed from the picking robot 102 to each of the transport robots to reach the handoff location. In some embodiments, the target transport robot can be determined by considering the sum of distances, the time difference, and the sum of energy consumed.

[0031] In determining the target robot, the picking robot 102 can also determine a handoff location 122 at which the picking robot 102 and the target transport robot are to hand off the target item 114. In some embodiments, the handoff location 122 can be a pre-determined fixed location. In some embodiments, the picking robot 102 can determine a respective plurality of handoff locations based on the position information of the plurality of transport robots, and determine a suitable target transport robot based on the handoff locations, and then determine the handoff location determined for the target transport robot as the handoff location 122. For example, in the environment 100, the picking robot 102 can determine handoff locations to hand off with the transport robots 104, 106, and 108, respectively, where the handoff location corresponding to the transport robot 106 is the handoff location 122. The picking robot 102 can determine that the handoff location 122 is the most suitable handoff location, and thus determine the transport robot 106 as the target transport robot.

[0032] As shown in FIG. 1, after determining that the transport robot 106 is the target transport robot, the picking robot 102 can send the transport task information 120 to the transport robot 106. The transport task information 120 can include information associated with the target item 114 to let the transport robot 106 know the item it is to transport. The transport task information 120 can also include the handoff location 122 to let the transport robot 106 know that it needs to move to the handoff location 122 to hand off with the picking robot 102. The transport task information 120 can also include the target location 116 to let the transport robot 106 know the final location to which the item needs to be transported after handing off the item.

[0033] In some embodiments, the picking robot 102 can also determine a target time to reach the handoff location 122. The target time can then be included in the transport task information 120 sent to the target transport robot. This can enable the transport robot 106 to better plan its own tasks and reduce the waiting time of the transport robot 106 at the handoff location 122.

[0034] In some embodiments, the picking robot 102 can communicate with the transport robot 106 one or more rounds to determine that both parties have received and confirmed the transport task information 120. The picking robot 102 can then transport the picked target item 114 to the handoff location 122 and hand off the target item 114 to the transport robot 106 at the handoff location 122. The picking robot 102 can confirm the target transport robot and send the transport task information 120 before moving to the shelf 118 to pick the target item 114, or can confirm the target transport robot and send the transport task information 120 after having picked the target item 114. After the handoff is completed, the picking robot 102 can continue to pick the next item, while the transport robot 106 can transport the target item 114 from the handoff location 122 to the target location 116.

[0035] In this way, the moving distance of the picking robot 102 can be reduced, and the time for the picking robot 102 to perform picking actions can be increased, so that the picking capacity of the picking robot 102 can be fully utilized, and the operation efficiency of the warehouse system can be improved. In addition, the transport robot can be specially equipped with higher transport capacity (e.g., higher moving speed, or higher terrain compatibility), so that the operation efficiency of the warehouse system can be further improved while saving costs.

[0036] FIG. 2 illustrates a flowchart of a method 200 for cooperation of multiple warehouse robots according to some embodiments of the present disclosure. The method 200 can be performed by a picking robot, e.g., the picking robot 102 in FIG. 1. As shown in FIG. 2, at block 202, the picking robot can obtain task information including target item information and a target location to which the target item is to be transported. For example, in the environment 100 as shown in FIG. 1, the picking robot 102 can receive the task information 112 from the control center 110, where the task information 112 can include information associated with the target item 114 (e.g., item identification, item name, item category, corresponding shelf, etc.) and the target location 116 to which the target item 114 is to be transported.

[0037] At block 204, the picking robot can obtain a plurality of position information of a plurality of transport robots. For example, in the environment 100 as shown in FIG. 1, there are the transport robot 104, the transport robot 106, and the transport robot 108 in the vicinity of the picking robot 102. The picking robot 102 can obtain the position information of these transport robots. The position information can include, for example, the current coordinates representing the exact position of the transport robot in the warehouse.

[0038] At block 206, the picking robot can determine a target transport robot from the plurality of transport robots based on the plurality of position information. For example, in the environment 100 as shown in FIG. 1, the picking robot 102 can determine a target transport robot to cooperate with from the transport robots 104, 106, and 108 based on the position information of the transport robots. In some embodiments of the present disclosure, the picking robot 102 can determine the target transport robot by considering the moving distance, time difference, energy consumption, or a combination thereof between the picking robot 102 and the target transport robot to reach the handoff location.

[0039] At block 208, the picking robot can send transport task information to the target transport robot, the transport task information including target item information, a handoff location, and a target location. For example, in the environment 100 as shown in FIG. 1, the picking robot 102 can send the transport task information 120 to the transport robot 106. The transport task information 120 can include information associated with the target item 114 to make the transport robot 106 aware of the item it is going to transport. The transport task information 120 can also include the handoff location 122 to make the transport robot 106 aware of the need to move to the handoff location 122 to hand off with the picking robot 102. The transport task information 120 can also include the target location 116 to make the transport robot 106 aware of the final location the item needs to be transported to after handing off the item.

[0040] At block 210, the picking robot can hand off the picked target item to the target transport robot at the handoff location. For example, in the environment 100 as shown in FIG. 1, the picking robot 102 can transport the picked target item 114 to the handoff location 122 and hand off the target item 114 to the transport robot 106 at the handoff location 122. After the handoff is completed, the picking robot 102 can continue to pick the next item while the transport robot 106 can transport the target item 114 from the handoff location 122 to the target location 116.

[0041] In this way, the picking robot is able to hand off the item to the transport robot immediately after completing the grasping task without personally transporting the item to a remote target location, which enables faster start of the next picking task and thus improves the utilization of the picking assembly.

[0042] FIG. 3 illustrates a flowchart of an example process 300 for determining a target transport robot from a plurality of transport robots, according to some embodiments of the present disclosure. As shown in FIG. 3, at block 302, the picking robot can determine a plurality of candidate handoff locations corresponding to the plurality of transport robots based on a plurality of location information of the plurality of transport robots. The picking robot can determine a plurality of handoff locations that are most suitable for each transport robot according to a current location of each transport robot and a target item location of the picking task. The picking robot can determine the respective handoff locations based on a working range of itself and locations reachable by the transport robots. The picking robot can also determine the respective handoff locations based on a shortest handoff path to reduce a total moving distance of the picking robot and the transport robots. In addition, the picking robot can also determine the respective handoff locations based on limitations of the warehouse layout, such as avoidance of narrow passages or busy traffic areas.

[0043] At block 304, the picking robot can determine a target transport robot based on the plurality of candidate handoff locations. In some embodiments, the picking robot can determine a plurality of time differences corresponding to the plurality of transport robots based on the plurality of candidate handoff locations, the time difference indicating a difference between a time for the picking robot to reach a respective candidate handoff location and a time for the transport robot to reach the respective candidate handoff location. The picking robot can determine a candidate handoff location having a smallest time difference among the plurality of time differences. Then, the picking robot can determine the candidate handoff location as a target handoff location and determine a transport robot corresponding to the candidate handoff location as a target transport robot. In this way, the waiting time of both parties can be reduced and the efficiency of completing the task can be improved.

[0044] In some embodiments, the picking robot can determine a plurality of moving distances based on the plurality of candidate handoff locations, the moving distance indicating a distance for the picking robot to move to a respective candidate handoff location. The picking robot can determine a candidate handoff location having a smallest moving distance among the plurality of moving distances. Then, the picking robot can determine the candidate handoff location as a target handoff location and determine a transport robot corresponding to the candidate handoff location as a target transport robot. In this way, the time for the picking robot to transport can be reduced and the utilization of the picking assembly can be improved.

[0045] In some embodiments, the picking robot can determine a plurality of total energy consumptions based on the plurality of candidate handoff locations, the total energy consumptions including energy consumptions of the picking robot and energy consumptions of the transport robots. The picking robot can determine a candidate handoff location having a smallest total energy consumption among the plurality of total energy consumptions. Then, the picking robot can determine the candidate handoff location as the target handoff location, and determine a transport robot corresponding to the candidate handoff location as the target transport robot. In determining the energy consumptions, the picking robot can determine the respective moving distances and the unit distance energy consumptions of the two parties, and determine the total energy consumptions based on the moving distances and the unit distance energy consumptions of the two parties. In this way, the total energy consumptions of the two parties can be reduced.

[0046] In some embodiments, the picking robot can obtain a weight for the sum of moving distances, a weight for the time difference, and a weight for the sum of energy consumptions. These weights can be predetermined based on the importance of the respective factors, and their sum is one. Then, the picking robot can determine a candidate handoff location having a smallest value by weighted summing the sum of moving distances, the time difference, and the sum of energy consumptions. Then, the picking robot can determine the candidate handoff location as the target handoff location, and determine a transport robot corresponding to the candidate handoff location as the target transport robot. In this way, various factors and their importance can be considered comprehensively, and the needs of the warehouse system can be better met.

[0047] FIG. 4 shows a schematic diagram of an example 400 of determining a target transport robot from a plurality of transport robots, according to some embodiments of the present disclosure. As shown in FIG. 4, the example 400 includes a picking robot 402, a transport robot 404, a transport robot 406, and a transport robot 408. The picking robot 402 can determine a candidate handoff location 414 for the transport robot 404, a candidate handoff location 416 for the transport robot 406, and a candidate handoff location 418 for the transport robot 408, respectively. Then, the picking robot 402 can determine a moving distance D1, a time T1, and an energy consumption E1 of itself to the candidate handoff location 414. The picking robot 402 can also determine a time T2 and an energy consumption E2 of the transport robot 404 to the candidate handoff location 414. Then, the picking robot 402 can determine a time difference of itself and the transport robot 404 to arrive at the candidate handoff location 414 (i.e., the absolute value of T1-T2), and a sum of energy consumptions (i.e., E1+E2).

[0048] Further, the picking robot 402 can determine its moving distance D2, time T3, and energy consumption E3 to the candidate handoff location 416. The picking robot 402 can also determine the time T4 and energy consumption E4 of the transport robot 406 to the candidate handoff location 416. Then, the picking robot 402 can determine the time difference (i.e., the absolute value of T3-T4) and the energy consumption sum (i.e., E3+E4) of the picking robot 402 and the transport robot 406 to reach the candidate handoff location 416.

[0049] In addition, the picking robot 402 can determine its moving distance D3, time T5, and energy consumption E5 to the candidate handoff location 418. The picking robot 402 can also determine the time T6 and energy consumption E6 of the transport robot 408 to the candidate handoff location 418. Then, the picking robot 402 can determine the time difference (i.e., the absolute value of T5-T6) and the energy consumption sum (i.e., E5+E6) of the picking robot 402 and the transport robot 408 to reach the candidate handoff location 418.

[0050] Then, the picking robot 402 can normalize the moving distance, time difference, and energy consumption of the picking robot, and determine the comprehensive score Score by the following equation (1):

[0051] Score = norm_D * w1 + norm_T * w2 + norm_E * w3 (1)

[0052] where norm_D represents the normalized moving distance of the picking robot, norm_T represents the normalized time difference, norm_E represents the normalized energy consumption, w1, w2, and w3 represent the corresponding weights (or importance) respectively, and the sum of w1, w2, and w3 is one.

[0053] After determining the comprehensive scores for the transport robots 404, 406, and 408 respectively, the transport robot with the smallest comprehensive score can be determined as the target transport robot, and the candidate handoff location corresponding thereto can be determined as the target handoff location. In this way, the moving distance, time difference, energy consumption of the picking robot 402 and the candidate transport robots to reach the corresponding candidate handoff location, and their importance can be considered comprehensively, so that the needs of the warehouse system can be better met.

[0054] In some embodiments, the plurality of transport robots includes a first transport robot, the plurality of candidate handoff locations includes a first candidate handoff location corresponding to the first transport robot, and in determining the plurality of candidate handoff locations corresponding to the plurality of transport robots based on the plurality of location information of the plurality of transport robots, the picking robot can determine a location of the picking robot and a location of the first transport robot, and then determine the first candidate handoff location based on the location of the picking robot, the location of the first transport robot, and the target location, using a reinforcement learning model (also referred to as a handoff location determination model herein).

[0055] In some embodiments, the picking robot can generate a state embedding based on the location of the picking robot, the location of the first transport robot, and the target location. The picking robot can generate a plurality of expected long-term rewards corresponding to the plurality of sub candidate handoff locations based on the state embedding. Then, the picking robot can determine the sub candidate handoff location with a maximum expected long-term reward in the plurality of expected long-term rewards as the first candidate handoff location.

[0056] In some embodiments, the picking robot can determine a first movement distance of the picking robot to the respective sub candidate handoff location, and determine a second movement distance of the first transport robot to the respective sub candidate handoff location. The picking robot can determine a movement distance sum of the first movement distance and the second movement distance. Then, the picking robot can generate the plurality of expected long-term rewards based on the movement distance sum.

[0057] In some embodiments, the picking robot can determine a first time for the picking robot to arrive at the respective sub candidate handoff location, and determine a second time for the first transport robot to arrive at the respective sub candidate handoff location. The picking robot can determine a time difference between the first time and the second time. Then, the picking robot can generate the plurality of expected long-term rewards based on the movement distance sum and the time difference.

[0058] FIG. 5 shows a schematic diagram of an example 500 of a handoff location determination model according to some embodiments of the present disclosure. The handoff location determination model shown in example 500 can be a deep Q network (DQN) based reinforcement learning model. Each transport robot can have a plurality of sub candidate handoff locations, and the handoff location determination model can determine a most suitable one from the plurality of sub candidate handoff locations as the candidate handoff location for the transport robot. For example, in example 400 shown in FIG. 4, transport robot 404 can have a plurality of sub candidate handoff locations, and the handoff location determination model can determine candidate handoff location 414 from the plurality of sub candidate handoff locations.

[0059] As shown in FIG. 5, example 500 includes an input layer 502, a hidden layer 504, and an output layer 506. It should be understood that, for brevity, although the hidden layer 504 is shown as a single module in FIG. 5, the hidden layer 504 can include multiple hidden layers. In example 500, the input layer 502 can receive a position 512 of a picking robot, a position 514 of a transport robot, and a target position 516. For example, the position 514 of the transport robot can be a position of the transport robot 404 in FIG. 4. The target position 516 is a final position to which a target item needs to be transported (e.g., the target position 116 in FIG. 1). The position 512 of the picking robot, the position 514 of the transport robot, and the target position 516 are also referred to as environment state information. The input layer 502 can generate a state embedding 518 based on the environment state information, which is a vector containing the state information of the environment in which the picking robot is located.

[0060] As shown in FIG. 5, the state embedding 518 can be input into the hidden layer 504. The hidden layer 504 can generate an activation value 520 based on the state embedding 518, which is generated by a plurality of nonlinear transformations on the state embedding 518. In the hidden layer 504, a linear transformation can be performed by performing a weight matrix multiplication and a bias addition on the state embedding 518. Then, a nonlinear transformation can be added to the result of the linear transformation using an activation function such as a rectified linear unit (ReLU).

[0061] As shown in FIG. 5, the activation value 520 can be input into the output layer 506. Each node of the output layer 506 corresponds to an expected long-term reward (i.e., a Q-value) of a sub candidate handoff position. For example, the output layer 506 can perform a linear transformation on the activation value 520 to output an expected long-term reward corresponding to each sub candidate handoff position. In example 500, three sub candidate handoff positions are included, i.e., sub candidate handoff positions 522, 524, and 526. The output layer 506 can generate an expected long-term reward 532 for the sub candidate handoff position 522, an expected long-term reward 534 for the sub candidate handoff position 524, and an expected long-term reward 536 for the sub candidate handoff position 526 based on the activation value 520. Then, the interaction position determination model can select the sub candidate handoff position with the largest expected long-term reward as the candidate handoff position for the transport robot, which can maximize the expected reward.

[0062] When computing the expected long-term reward 532, the movement distance of the picking robot to the sub-candidate handoff location 522 and the movement distance of the transport robot to the sub-candidate handoff location 522 can be determined, and the sum of the two can be used to determine the expected long-term reward 532. For example, the smaller the sum of the movement distances, the greater the value of the expected long-term reward 532. In this way, the determined candidate handoff location can have the smallest sum of movement distances with respect to the picking robot and the transport robot. In addition, the time for the picking robot to reach the sub-candidate handoff location 522 and the time for the transport robot to reach the sub-candidate handoff location 522 can also be determined, and the difference between the two can be used to determine the expected long-term reward 532. For example, the smaller the time difference, the shorter the waiting time, the greater the value of the expected long-term reward 532. In this way, the determined candidate handoff location can have the smallest waiting time with respect to the picking robot and the transport robot.

[0063] By using a reinforcement learning model to determine the handoff location, the picking robot can automatically find a better handoff location and strategy in a complex environment. Compared to traditional rule-based or algorithm-based optimization, the reinforcement learning model focuses on long-term returns. By selecting a handoff location, not only the short-term task completion effect is considered, but also how to reduce the total time and cost in future multiple tasks is learned, achieving global optimization.

[0064] FIG. 6 illustrates a flowchart of an example process 600 for training a handoff location determination model, according to some embodiments of the present disclosure. As shown in FIG. 6, at block 602, the process 600 can initialize the handoff location determination model. For example, the Q-network and the target network can be randomly initialized. The Q-network is used to estimate the expected long-term reward of each candidate handoff location in the current state. In the example 600, the weights and biases of the Q-network can be randomly initialized. The target network can have the same structure as the Q-network, and its parameters can be the same as the Q-network at initialization. The target network is used to calculate the target expected long-term reward, and its parameters are updated at a slower rate during training to prevent training instability. In addition, an experience replay pool can also be created to store past experiences, which will be randomly drawn for subsequent training to update the Q-network. In addition, the exploration rate ∈ can also be initialized to control the balance between exploration and exploitation when selecting a handoff location. A higher ∈ value can be set at the beginning to encourage exploration.

[0065] At block 604, the process 600 can generate training data by having the picking robot interact with the environment. The picking robot can perceive the environment to obtain a current state, which can include, for example, the picking robot’s position, the transporting robot’s position, and the target location, etc. The picking robot can then select a handoff location according to the current state and by the Q network, which can follow an ε-greedy policy. For example, the picking robot can randomly select a candidate handoff location with a probability of ε for exploration of new strategies. The picking robot can then select the candidate handoff location with the largest expected long-term reward output by the current Q network with a probability of 1 - ε. The picking robot can then move to the selected candidate handoff location and handoff with the transporting robot, causing the environment to feedback a new state and an immediate reward. For example, the environment can return an immediate reward based on the selection of the handoff location and whether the handoff is successful, with a positive reward for a successful handoff and a negative reward for an unsuccessful handoff. After the handoff is completed, the environment can transition to a new state.

[0066] At block 606, the process 600 can store the experience of the current interaction into an experience replay pool. The experience replay pool is used to store the history of interactions of the picking robot with the environment, from which experience can be sampled to update the network in subsequent training.

[0067] At block 608, the process 600 can perform an update operation of the network after each determination of a handoff location. The process 600 can randomly sample a batch of experience from the experience replay pool, which is interaction data from multiple time steps in the past. For each sampled experience, the target network can be used to compute a target expected long-term reward. A loss function, such as mean squared error, can then be used to compute the loss between the current estimate of the Q network and the target long-term reward, i.e., the difference between the estimate of the current Q network and the target expected long-term reward. Backpropagation and gradient descent algorithms can then be used to minimize the loss function to update the parameters of the Q network. As training progresses, the Q network can gradually learn how to estimate the correct expected long-term reward for each state-handoff location pair. The parameters of the Q network can be copied to the target network every fixed number of steps. The target network can be updated at a lower frequency, which can improve the stability of training and avoid oscillation of the expected long-term reward.

[0068] At block 610, the process 600 can gradually decrease the exploration rate until a training end condition is met. As training progresses, the exploration rate ε can be gradually decreased so that the model can rely more on the learned strategies rather than random exploration. A higher exploration rate is needed initially so that the robot can sufficiently explore the warehouse environment and find possible optimal strategies. Over time, ε can decay to a lower value, at which point the Q network’s estimate is relied on more for selecting handoff locations.

[0069] FIG. 7 illustrates a block diagram of an apparatus 700 for cooperation of multiple warehousing robots, according to some embodiments of the present disclosure. As shown in FIG. 7, the apparatus 700 includes a task information obtaining module 702 configured to obtain, by a picking robot, task information, the task information including target item information and a target location to which the target item is transported. The apparatus 700 further includes a location information obtaining module 704 configured to obtain, by the picking robot, multiple location information of multiple transport robots. The apparatus 700 further includes a transport robot determining module 706 configured to determine, by the picking robot, a target transport robot from the multiple transport robots based on the multiple location information. The apparatus 700 further includes a transport task sending module 708 configured to send, by the picking robot, transport task information to the target transport robot, the transport task information including the target item information, a handover location, and the target location. In addition, the apparatus 700 further includes a target item handover module 710 configured to hand over, by the picking robot, the picked target item to the target transport robot at the handover location.

[0070] In some embodiments, the transport robot determining module 706 includes: a candidate handover location determining module configured to determine multiple candidate handover locations corresponding to the multiple transport robots based on the multiple location information of the multiple transport robots; and a candidate handover location using module configured to determine the target transport robot based on the multiple candidate handover locations.

[0071] In some embodiments, the candidate handover location using module includes: a time difference determining module configured to determine multiple time differences corresponding to the multiple transport robots based on the multiple candidate handover locations, the time difference indicating a difference between a time for the picking robot to arrive at a corresponding candidate handover location and a time for the transport robot to arrive at the corresponding candidate handover location; a minimum time difference determining module configured to determine a candidate handover location having a minimum time difference in the multiple time differences; a first handover location determining module configured to determine the candidate handover location as the handover location; and a first target transport robot determining module configured to determine a transport robot corresponding to the candidate handover location as the target transport robot.

[0072] In some embodiments, the candidate handoff location usage module includes: a movement distance determination module configured to determine a plurality of movement distances based on the plurality of candidate handoff locations, the movement distance indicating a distance for the picking robot to move to the respective candidate handoff location; a minimum movement distance determination module configured to determine the candidate handoff location having a minimum movement distance of the plurality of movement distances; a second handoff location determination module configured to determine the candidate handoff location as the handoff location; and a second target transport robot determination module configured to determine the transport robot corresponding to the candidate handoff location as the target transport robot.

[0073] In some embodiments, the candidate handoff location usage module includes: an energy consumption determination module configured to determine a plurality of total energy consumptions based on the plurality of candidate handoff locations, the total energy consumption including an energy consumption of the picking robot and an energy consumption of the transport robot; a minimum energy consumption determination module configured to determine the candidate handoff location having a minimum total energy consumption of the plurality of total energy consumptions; a third handoff location determination module configured to determine the candidate handoff location as the handoff location; and a third target transport robot determination module configured to determine the transport robot corresponding to the candidate handoff location as the target transport robot.

[0074] In some embodiments, the plurality of transport robots includes a first transport robot, the plurality of candidate handoff locations includes a first candidate handoff location corresponding to the first transport robot, and the candidate handoff location determination module includes: a location determination module configured to determine a location of the picking robot and a location of the first transport robot; and a model usage module configured to determine the first candidate handoff location using a reinforcement learning model based on the location of the picking robot, the location of the first transport robot, and the target location.

[0075] In some embodiments, the model usage module includes: a state embedding generation module configured to generate a state embedding based on the location of the picking robot, the location of the first transport robot, and the target location; a reward generation module configured to generate a plurality of expected long-term rewards corresponding to a plurality of sub-candidate handoff locations based on the state embedding; and a maximum reward determination module configured to determine the sub-candidate handoff location having a maximum expected long-term reward of the plurality of expected long-term rewards as the first candidate handoff location.

[0076] In some embodiments, the reward generation module comprises: a first movement distance determination module configured to determine a first movement distance of the picking robot to the respective sub-candidate handover position; a second movement distance determination module configured to determine a second movement distance of the first transport robot to the respective sub-candidate handover position; a movement distance sum determination module configured to determine a movement distance sum of the first movement distance and the second movement distance; and a movement distance sum usage module configured to generate the plurality of expected long-term rewards based on the movement distance sum.

[0077] In some embodiments, the reward generation module comprises: a first time determination module configured to determine a first time for the picking robot to arrive at the respective sub-candidate handover position; a second time determination module configured to determine a second time for the first transport robot to arrive at the respective sub-candidate handover position; a time difference generation module configured to determine a time difference between the first time and the second time; and a time difference usage module configured to generate the plurality of expected long-term rewards based on the movement distance sum and the time difference.

[0078] In some embodiments, the transport task sending module comprises: a target time determination module configured to determine a target time for arriving at the handover position; and a target time sending module configured to send the transport task information to the target transport robot, wherein the transport task information further comprises the target time.

[0079] It can be understood that, with the device 700 of the present disclosure, at least one of the many advantages that can be achieved by the method or process as described above can be achieved. For example, through cooperation, the picking robot can hand over the item to the transport robot immediately after completing the picking task, without personally transporting the item to a distant target position, which can enable the next picking task to be started faster, thereby enabling the utilization rate of the picking assembly to be improved.

[0080] FIG. 8 shows a flowchart of a method 800 for determining a handover position of a picking robot and a transport robot, according to some embodiments of the present disclosure. The method 800 can be performed by a picking robot (e.g., the picking robot 102 in FIG. 1). As shown in FIG. 8, at block 802, the picking robot can obtain position information of the picking robot, position information of the transport robot, and a target position to which a target item to be picked by the picking robot is transported. At block 804, the picking robot can generate a state embedding based on the position information of the picking robot, the position information of the transport robot, and the target position. At block 806, the picking robot can generate a plurality of expected long-term rewards corresponding to a plurality of candidate handover positions based on the state embedding. At block 808, the picking robot can determine a candidate handover position having a maximum expected long-term reward in the plurality of expected long-term rewards as the handover position.

[0081] By utilizing the reinforcement learning model to determine the handover position, the picking robot can automatically find a better handover position and strategy in a complex environment. Compared with traditional rule-based or algorithm-based optimization, the reinforcement learning model focuses on long-term returns. By selecting a handover position, not only the short-term task completion effect is considered, but also how to reduce the total time and cost in future multiple tasks is learned to achieve global optimization.

[0082] In some embodiments, wherein generating the plurality of expected long-term rewards corresponding to the plurality of candidate handover positions based on the state embedding comprises: determining a first movement distance of the picking robot to a respective candidate handover position; determining a second movement distance of the transport robot to the respective candidate handover position; determining a movement distance sum of the first movement distance and the second movement distance; and generating the plurality of expected long-term rewards based on the movement distance sum. In this way, the determined candidate handover position can have a minimum movement distance sum with respect to the picking robot and the transport robot, thereby improving system operation efficiency.

[0083] In some embodiments, wherein generating the plurality of expected long-term rewards based on the movement distance sum comprises: determining a first time for the picking robot to reach the respective candidate handover position; determining a second time for the transport robot to reach the respective candidate handover position; determining a time difference between the first time and the second time; and generating the plurality of expected long-term rewards based on the movement distance sum and the time difference. In this way, the determined candidate handover position can have a minimum waiting time with respect to the picking robot and the transport robot.

[0084] FIG. 9 shows a block diagram of a device 900 that can implement a number of embodiments of the present disclosure. The device 900 may, for example, be a processing unit of a picking robot 102 as shown in FIG. 1. As shown in FIG. 9, the device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901 that can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or loaded into a random access memory (RAM) 903 from a storage unit 908. In the RAM 903, various programs and data required for operation of the device 900 can also be stored. The CPU / GPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904. Although not shown in FIG. 9, the device 900 can also include a coprocessor.

[0085] A number of the components in device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, a CD, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows device 900 to exchange information / data with other devices over a computer network, such as the Internet, and / or various telecommunication networks.

[0086] The various methods or processes described above can be performed by the CPU / GPU 901. For example, in some embodiments, the methods can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 908. In some embodiments, portions or all of the computer program can be loaded and / or installed onto device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded onto the RAM 903 and executed by the CPU / GPU 901, one or more steps or actions of the methods or processes described above can be performed.

[0087] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions embodied therewith.

[0088] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a

[0089] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0090] Computer readable program instructions for carrying out operations of the present disclosure can be assembly-level instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object-oriented programming languages and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0091] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including a manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0092] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0093] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of devices, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0094] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are within the scope of the disclosure. The selection of the terms to be used in the description is not intended to limit the scope of the embodiments described herein, but rather to best describe the principles of the embodiments in the context of the specific application.

Claims

1. A method for collaboration among a plurality of warehouse robots, the plurality of warehouse robots comprising a picking robot and a plurality of transport robots, the method comprising: obtaining, by the picking robot, task information, the task information comprising target item information and a target location to which a target item is to be transported; obtaining, by the picking robot, a plurality of location information of the plurality of transport robots; determining, by the picking robot, a target transport robot from the plurality of transport robots based on the plurality of location information; sending, by the picking robot, transport task information to the target transport robot, the transport task information comprising the target item information, a handover location, and the target location; and handing over, by the picking robot, the picked target item to the target transport robot at the handover location. 2.The method of claim 1, wherein determining, by the picking robot, the target transport robot from the plurality of transport robots based on the plurality of location information comprises: determining a plurality of candidate handover locations corresponding to the plurality of transport robots based on the plurality of location information of the plurality of transport robots; and determining the target transport robot based on the plurality of candidate handover locations. 3.The method of claim 2, wherein determining the target transport robot based on the plurality of candidate handover locations comprises: determining a plurality of time differences corresponding to the plurality of transport robots based on the plurality of candidate handover locations, a time difference indicating a difference between a time at which the picking robot arrives at a respective candidate handover location and a time at which a transport robot arrives at the respective candidate handover location; determining a candidate handover location having a smallest time difference among the plurality of time differences; determining the candidate handover location as the handover location; and determining a transport robot corresponding to the candidate handover location as the target transport robot. 4.The method of claim 2, wherein determining the target transport robot based on the plurality of candidate handover locations comprises: determining a plurality of movement distances based on the plurality of candidate handover locations, a movement distance indicating a distance at which the picking robot moves to a respective candidate handover location; determining a candidate handover location having a smallest movement distance among the plurality of movement distances; determining the candidate handover location as the handover location; and determining a transport robot corresponding to the candidate handover location as the target transport robot. 5.The method of claim 2, wherein determining the target transport robot based on the plurality of candidate handover locations comprises: determining a plurality of total energy consumptions based on the plurality of candidate handover locations, a total energy consumption comprising an energy consumption of the picking robot and an energy consumption of the transport robot; determining a candidate handover location having a smallest total energy consumption among the plurality of total energy consumptions; determining the candidate handover location as the handover location; and determining a transport robot corresponding to the candidate handover location as the target transport robot. ​ ​ ​ 6.The method of claim 2, wherein the plurality of transport robots comprises a first transport robot, the plurality of candidate handoff locations comprises a first candidate handoff location corresponding to the first transport robot, and determining the plurality of candidate handoff locations corresponding to the plurality of transport robots based on the plurality of location information of the plurality of transport robots comprises: determining a location of the picking robot and a location of the first transport robot; and determining the first candidate handoff location using a reinforcement learning model based on the location of the picking robot, the location of the first transport robot, and the target location. 7.The method of claim 6, wherein determining the first candidate handoff location using the reinforcement learning model based on the location of the picking robot, the location of the first transport robot, and the target location comprises: generating a state embedding based on the location of the picking robot, the location of the first transport robot, and the target location; generating a plurality of expected long-term rewards corresponding to a plurality of sub candidate handoff locations based on the state embedding; and determining a sub candidate handoff location having a maximum expected long-term reward in the plurality of expected long-term rewards as the first candidate handoff location. 8.The method of claim 7, wherein generating the plurality of expected long-term rewards corresponding to the plurality of sub candidate handoff locations based on the state embedding comprises: determining a first movement distance of the picking robot to a respective sub candidate handoff location; determining a second movement distance of the first transport robot to the respective sub candidate handoff location; determining a movement distance sum of the first movement distance and the second movement distance; and generating the plurality of expected long-term rewards based on the movement distance sum. 9.The method of claim 8, wherein generating the plurality of expected long-term rewards based on the movement distance sum comprises: determining a first time for the picking robot to reach the respective sub candidate handoff location; determining a second time for the first transport robot to reach the respective sub candidate handoff location; determining a time difference between the first time and the second time; and generating the plurality of expected long-term rewards based on the movement distance sum and the time difference. 10.The method of claim 1, wherein sending, by the picking robot, the transport task information to the target transport robot comprises: determining a target time to reach the handoff location; and sending the transport task information to the target transport robot, wherein the transport task information further comprises the target time. 11.A method for determining a handoff location of a picking robot and a transport robot, comprising: obtaining location information of the picking robot, location information of the transport robot, and a target location to which a target item to be picked by the picking robot is transported; generating a state embedding based on the location information of the picking robot, the location information of the transport robot, and the target location; and determining the handoff location using a reinforcement learning model based on the state embedding. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ generating, based on the state embedding, a plurality of expected long-term rewards corresponding to a plurality of candidate handoff locations; and determining, as the handoff location, a candidate handoff location having a largest expected long-term reward among the plurality of expected long-term rewards.

12. The method of claim 11, wherein generating, based on the state embedding, the plurality of expected long-term rewards corresponding to the plurality of candidate handoff locations comprises: determining a first movement distance of the picking robot to a respective candidate handoff location; determining a second movement distance of the transporting robot to the respective candidate handoff location; determining a movement distance sum of the first movement distance and the second movement distance; and generating the plurality of expected long-term rewards based on the movement distance sum.

13. The method of claim 12, wherein generating the plurality of expected long-term rewards based on the movement distance sum comprises: determining a first time for the picking robot to reach the respective candidate handoff location; determining a second time for the transporting robot to reach the respective candidate handoff location; determining a time difference between the first time and the second time; and generating the plurality of expected long-term rewards based on the movement distance sum and the time difference.

14. An apparatus for collaboration among a plurality of warehouse robots, the plurality of warehouse robots comprising a picking robot and a plurality of transporting robots, comprising: a task information obtaining module configured to obtain, by the picking robot, task information, the task information comprising target item information and a target location to which a target item is to be transported; a location information obtaining module configured to obtain, by the picking robot, a plurality of location information of the plurality of transporting robots; a transporting robot determining module configured to determine, by the picking robot, a target transporting robot from the plurality of transporting robots based on the plurality of location information; a transportation task sending module configured to send, by the picking robot, transportation task information to the target transporting robot, the transportation task information comprising the target item information, a handoff location, and the target location; and a target item handoff module configured to handoff, by the picking robot, the picked target item to the target transporting robot at the handoff location.

15. An electronic device, comprising: a processor; and a memory coupled with the processor, the memory having instructions stored therein that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1-13.

16. A computer program product tangibly stored on a non-transitory computer- readable medium and comprising machine-executable instructions that, when executed, cause a machine to implement the method of any one of claims 1-13. ​ ​