Method, device, apparatus and product for picking articles

CN122074057APending Publication Date: 2026-05-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YOUZHUJU NETWORK TECH CO LTD
Filing Date
2024-09-20
Publication Date
2026-05-22

Smart Images

  • Figure CN122074057A_ABST
    Figure CN122074057A_ABST
Patent Text Reader

Abstract

A method, apparatus, device and product for sorting articles, the method comprising: in response to a sorting task received by a sorting robot, generating sorting task information based on the sorting task, the sorting task information comprising the number of articles to be sorted, an article image and a specified storage location; determining a moving path of the picking robot based on a predefined warehouse map and a specified storage position; in response to the fact that the picking robot moves to the designated storage position through the moving path, the action model determines an action sequence to be executed by the arm of the picking robot based on the current state of the arm of the picking robot and the picking task information; and based on the determined action sequence, picking the articles needing to be picked by the picking robot. The sorting robot not only can directly navigate to the position of the article, but also can accurately sort out the article, so that the user experience in the aspect of warehouse logistics is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and product for picking items TECHNICAL FIELD

[0001] The present disclosure relates to the field of robotics, and more specifically to a method, device, equipment and product for picking items. BACKGROUND

[0002] A picking robot is an intelligent device integrating multiple technologies such as robotics, sensor technology, computer vision and artificial intelligence, which can pick goods in warehouses, logistics centers and other places.

[0003] Generally, a picking robot can be equipped with various types of end effectors, such as mechanical arms, grippers, suction cups, etc., to adapt to the grasping needs of items of different shapes and materials. The emergence of picking robots has greatly improved the efficiency and quality of logistics operations, and has become an important part of the intelligent transformation of logistics and warehousing.

[0004] SUMMARY

[0005] In a first aspect of embodiments of the present disclosure, a method for picking items is provided. The method comprises, in response to a picking robot receiving a picking task, generating picking task information based on the picking task, wherein the picking task information comprises a number of items required to be picked, images of the items, and a designated storage location. The method further comprises determining a movement path of the picking robot based on a predefined warehouse map and the designated storage location. The method further comprises, in response to the picking robot moving to the designated storage location through the movement path, determining, by an action model, a sequence of actions to be performed by an arm of the picking robot based on a current state of the arm and the picking task information. In addition, the method further comprises picking, by the picking robot, the items required to be picked based on the determined sequence of actions.

[0006] In a second aspect of embodiments of the present disclosure, a device for picking items is provided. A picking task information generation module is configured to, in response to a picking robot receiving a picking task, generate picking task information based on the picking task, wherein the picking task information comprises a number of items required to be picked, images of the items, and a designated storage location. The device further comprises a movement path determination module configured to determine a movement path of the picking robot based on a predefined warehouse map and the designated storage location. The device further comprises an action sequence determination module configured to, in response to the picking robot moving to the designated storage location through the movement path, determine, by an action model, a sequence of actions to be performed by an arm of the picking robot based on a current state of the arm and the picking task information. In addition, the device further comprises an item picking module configured to pick, by the picking robot, the items required to be picked based on the determined sequence of actions.

[0007] In a third aspect of embodiments of the present disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device storing one or more programs, when executed by the one or more processors, cause the one or more processors to implement a method for picking an item. The method includes, in response to a picking robot receiving a picking task, generating picking task information based on the picking task, wherein the picking task information includes a quantity of items required to be picked, item images, and a designated storage location. The method further includes determining a movement path of the picking robot based on a predefined warehouse map and the designated storage location. The method further includes, in response to the picking robot moving to the designated storage location via the movement path, determining, by an action model, a sequence of actions to be performed by an arm of the picking robot based on a current state of the arm of the picking robot and the picking task information. In addition, the method further includes picking, by the picking robot, the items required to be picked based on the determined sequence of actions.

[0008] In a fourth aspect of embodiments of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer readable medium and includes machine executable instructions that, when executed, cause a machine to implement a method for picking an item. The method includes, in response to a picking robot receiving a picking task, generating picking task information based on the picking task, wherein the picking task information includes a quantity of items required to be picked, item images, and a designated storage location. The method further includes determining a movement path of the picking robot based on a predefined warehouse map and the designated storage location. The method further includes, in response to the picking robot moving to the designated storage location via the movement path, determining, by an action model, a sequence of actions to be performed by an arm of the picking robot based on a current state of the arm of the picking robot and the picking task information. In addition, the method further includes picking, by the picking robot, the items required to be picked based on the determined sequence of actions.

[0009] The summary is provided to introduce a selection of concepts, in a simplified form, that are further described below in the DETAILED DESCRIPTION. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent upon reading the following detailed description in conjunction with the accompanying drawings, in which like references refer to like elements. In the drawings:

[0011] FIG. 1 illustrates a schematic diagram of an example environment in which multiple embodiments of the present disclosure can be implemented;

[0012] FIG. 2 shows a flowchart of a method for picking items, according to some embodiments of the present disclosure;

[0013] FIG. 3 shows a schematic diagram of an example process for a picking robot to perform a picking task, according to some embodiments of the present disclosure;

[0014] FIG. 4 shows a schematic diagram of an example of an action model, according to some embodiments of the present disclosure;

[0015] FIG. 5 shows a schematic diagram of an example for training an action model, according to some embodiments of the present disclosure;

[0016] FIG. 6 shows a schematic diagram of an example process for fine-tuning an action model, according to some embodiments of the present disclosure;

[0017] FIG. 7 shows a schematic diagram of an example process for further fine-tuning an action model, according to some embodiments of the present disclosure;

[0018] FIG. 8 shows a block diagram of an apparatus for picking items, according to some embodiments of the present disclosure; and

[0019] FIG. 9 shows a block diagram of a device capable of implementing a number of embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] It can be understood that all user-related data involved in the present technical solution should be obtained and used after the user's authorization. This means that in the present technical solution, if the user's personal information needs to be used, the user's explicit consent and authorization are required before obtaining these data, otherwise the relevant data collection and use will not be carried out. It should also be understood that in the implementation of the present technical solution, relevant laws and regulations should be strictly followed in the process of data collection, use and storage, and necessary technical and measures should be taken to protect the user's data security and ensure the safe use of data.

[0021] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] In the description of embodiments of the disclosure, the term "includes" and its similar terms are understood to be open-ended, i.e., "includes but is not limited to". The term "based on" is understood to be "based, at least in part, on". The term "one embodiment" or "the embodiment" is understood to be "at least one embodiment". The terms "first", "second", and the like can refer to different or identical objects unless explicitly stated otherwise. Other explicit and implicit definitions can also be included below.

[0023] As described above, the picking robot can assist in picking up items. Generally, warehouse items need to be strictly placed according to specific requirements and standards to achieve efficient use of space and convenient management. However, in reality, many items can be cleverly stacked or arranged due to their unique material, shape, size, and other special properties, thereby saving warehouse space. However, in the related art, picking robots are mainly good at picking items that are placed in a regular manner. For items of different types, shapes, or properties, especially when they are arranged in an irregular manner, it is difficult for picking robots to effectively and accurately pick the target items. In addition, existing picking robots are mainly designed for picking items within a fixed area. However, when the logistics task volume in the warehouse surges, they cannot directly provide assistance, which can cause a bottleneck in the warehouse logistics process, because the picked items need additional human resources or equipment for transportation, thereby increasing the complexity and cost of the overall operation.

[0024] To this end, embodiments of the present disclosure provide a scheme for picking items. When the picking robot receives a picking task, it generates detailed task information including the number of items, images, and storage locations according to the task requirements, and plans an optimal moving path to the specified location using a predefined warehouse map. Upon arrival, the robot combines the current state of the arm and the task information to plan and execute a series of precise action sequences through an action model, thereby efficiently picking the required items.

[0025] In this way, the picking robot can not only directly navigate to the location of the items, but also accurately pick the items that meet the picking task. This reduces logistics costs and improves the user experience of managing warehouse logistics.

[0026] FIG. 1 illustrates a schematic diagram of an example environment 100 in which various embodiments of the present disclosure can be implemented. As shown in FIG. 1, before a picking job starts, a picking robot 120 can receive picking tasks 110 from a warehouse management system (or control center). These picking tasks 110 can be confirmed by a user and sent to the warehouse management system, which then dispatches the picking tasks 110 to the picking robot 120 to guide the picking robot 120 to complete subsequent picking tasks. In some embodiments, the outbound area of the picking robot can be a packing area of items, and the picking area where the picking robot goes to pick items can be a warehouse area.

[0027] With continued reference to FIG. 1, after the picking tasks 110 are dispatched to the picking robot 120, the picking robot 120 can first analyze the picking tasks 110 to determine the type and quantity of items that need to be picked, which can ensure the accuracy of subsequent picking by the picking robot 120. In some embodiments, the location of the items that need to be picked can also be determined based on the picking tasks 110. This information related to the picking tasks can be referred to as picking task information 112. In some embodiments, the picking task information 112 can be described as "item labeled 3 on the second layer of the S1 shelf in the second column and third row of picking area A," which can have image information of the target item attached thereto.

[0028] In some embodiments, the picking task information 112 can also include an image of the item to be picked, and the image of the item to be picked can also be parsed into a textual form of a feature description. In some embodiments, the picking task information can also include other types of feature descriptions, which can be a two-dimensional code or a bar code of the item to be picked. With the aid of this image information or feature description information, the picking robot 120 can quickly and accurately identify the target item through a visual recognition system.

[0029] Continuing to refer to FIG. 1, the picking robot 120 can determine a movement path to reach the storage location of the item to be picked based on its built-in warehouse map and the storage location of the item to be picked, such as the location of the target shelf 130. In some embodiments, the storage location information can be represented by coordinates on the warehouse map, a shelf number, a layer number, and a bin number. The picking robot 120 can use the storage location information in conjunction with its own positioning system and navigation algorithm to plan an optimal movement path to quickly and accurately reach the location of the item to be picked to perform the picking task. In some embodiments, the picking robot 120 can typically pre-store a warehouse map in its system, which indicates information such as nodes, areas, passages, and obstacles of the warehouse. In some embodiments, factors such as distance, time, and energy consumption can also be considered when planning the movement path, so as to ensure that the picking robot 120 can efficiently and safely reach the picking location.

[0030] Continuing to refer to FIG. 1, when the picking robot 120 reaches the target shelf 130, it can determine the action sequence 126 to be performed by the arm 122 of the picking robot 120 and the opening and closing state of the gripper 124 based on the picking task information and the current state of the picking robot arm (including the picking robot arm 122 and the gripper 124 configured at the end of the picking robot arm). In some embodiments, the gripper 124 can also be a suction cup, a mechanical gripper, a magnetic gripper, a multi-axis gripper, etc. As can be seen, the items in the target shelf 130 are not neatly arranged, and these items are also irregular, as shown by the items 132 and the target item 134. Alternatively, when the picking robot 120 reaches the target shelf 130, it can take a picture of the target shelf, so as to identify the specific location of the target item from the shelf through image recognition technology.

[0031] In some embodiments, the picking robot 120 can evaluate the current state of the arm 122 for better performance of the picking task, which can include the position of the arm 122 (i.e., its coordinates in three-dimensional space), the pose of the arm 122 (i.e., the angle of the arm 122), and the load capacity of the arm 122 (i.e., whether the arm 122 can currently bear the weight of the item to be picked).

[0032] In some embodiments, in order to be able to more accurately grasp the item to be picked, an action model can be invoked to calculate the action sequence 126 that the arm 122 of the picking robot 120 and the gripper 124 need to take when performing the picking task 110. In some embodiments, this action model can output the action sequence 126 that the robot arm 122 should perform according to the current picking task information (e.g., the target item 134 labeled as 3 on the second layer of the S1 shelf in the second column and third row of the picking area A) and the current state of the picking robot arm. When the action sequence 126 is generated, the picking robot can perform the specific picking task according to the action sequence 126.

[0033] With continued reference to FIG. 1, after the picking robot 120 picks the target item 134, the picking robot 120 can place the target item 134 on a carrying platform (not shown) carried by the picking robot 120. Then, the picking robot 120 can return to the original starting point according to the original movement path. In some embodiments, the picking robot 120 can also real-time re-plan the route for returning.

[0034] In this way, the picking robot 120 not only can directly locate and navigate to the location of the item, but also can accurately pick the item that matches the picking task, thereby reducing the logistics cost and improving the experience of managing warehouse logistics.

[0035] FIG. 2 illustrates a flowchart of a method 200 for picking an item, according to some embodiments of the present disclosure. The method 200 can be performed by a picking robot, such as the picking robot 120 in FIG. 1. As shown in FIG. 2, in response to the picking robot receiving a picking task, picking task information is generated based on the picking task, where the picking task information includes the quantity of items to be picked, images of the items, and the designated warehouse location, at block 202. For example, in the environment 100 as shown in FIG. 1, the picking robot 110 can receive the picking task 110 from a control center or a warehouse management system, and then can generate the picking task information 112 according to the received picking task 110, where the picking task information 112 can include information associated with the item to be picked (e.g., the type, quantity, and image-based feature description of the item, etc.) and the warehouse location information of the target item 134.

[0036] At block 204, a movement path of the picking robot is determined based on the predefined warehouse map and the designated storage location. For example, in the environment 100 as shown in FIG. 1, the picking robot 120 can determine a movement path to reach the storage location of the item to be picked according to its built-in warehouse map and the storage location of the item to be picked (e.g., the location of the target shelf 130). In some embodiments, the storage location information can be represented by coordinates on the warehouse map, a shelf number, a layer number, and a bin number. The picking robot 120 can use the storage location information in combination with its own positioning system and navigation algorithm to plan an optimal movement path to quickly and accurately reach the location of the item to be picked to perform the picking task. In some embodiments, the picking robot 120 can typically pre-store a warehouse map in its system, which maps the nodes, areas, passages, and obstacles of the warehouse. In some embodiments, factors such as example, time, and energy consumption can also be considered when planning the movement path, so as to ensure that the picking robot 120 can efficiently and safely reach the picking location.

[0037] At block 206, in response to the picking robot moving to the designated storage location via the movement path, an action sequence to be performed by the picking robot arm is determined based on a current state of the picking robot arm and the picking task information by the action model. For example, in the environment 100 as shown in FIG. 1, when the picking robot 120 reaches the target shelf 130, it can determine the action sequence 126 to be performed by the arm 122 of the picking robot 120 and the opening and closing state of the gripper 124 according to the picking task information and the current state of the picking robot arm (including the picking robot arm 122 and the gripper 124 configured at the end of the picking robot arm). In some embodiments, in order to more accurately grasp the item to be picked, the action model can be invoked to calculate the action sequence 126 to be taken by the arm 122 of the picking robot 120 and the gripper 124 when performing the picking task 110. In some embodiments, the action model can output the action sequence 126 to be performed by the picking robot arm 122 according to the current picking task information (e.g., the target item 134 to be picked) and the current state of the picking robot arm.

[0038] At block 208, the item to be picked is picked by the picking robot based on the determined action sequence. For example, in the environment 100 as shown in FIG. 1, when the action sequence 126 is generated, the picking robot can perform the specific picking task according to the action sequence 126, i.e., the target item 134 can be picked.

[0039] In this way, the picking robot can not only directly locate and navigate to the location of the item, but also accurately pick the item that matches the picking task. In this way, the logistics cost is reduced, thereby improving the experience of managing warehouse logistics.

[0040] FIG. 3 shows a schematic diagram of an example 300 of a flow of performing a picking task by a picking robot, according to some embodiments of the present disclosure. As shown in FIG. 3, at 310, a picking task is received and parsed. In some embodiments, the picking robot can receive picking tasks from a control center or a warehouse management system, which are submitted by a user and confirmed by the warehouse management system or the control center, and then issued to the picking robot for execution. In some embodiments, the picking robot's departure area can be a packing area, and the area where the picking robot goes to pick the item is a warehouse area, which is specially used for storing items.

[0041] In some embodiments, when the picking robot receives the picking task, it first parses the picking task to obtain picking task information, which can include the quantity, type, and image of the item to be picked (also referred to as the target item), the storage location, and other information, so as to ensure the accuracy of the picking robot when picking the target item. In some embodiments, the storage location information can be represented by coordinates on the warehouse map, shelf number, layer number, and bin number. For example, the location information of the target item 134 shown in FIG. 1 can be [Area A, second column, third row, S1 shelf, second layer, and label 3]. In some embodiments, the picking task information 112 can be described as "picking the item labeled 3 on the second layer of the S1 shelf in the third row of the second column of Area A", which can be attached with image information of the target item.

[0042] Referring to FIG. 3, at 320, a path to the area where the item to be picked is located is planned and navigated. In order to save transportation resources, the picking robot can proceed from the packing area to the warehouse area to pick the target item. In some embodiments, the picking robot usually pre-stores the map of the warehouse in its system, which standardizes the nodes, areas, channels, and obstacles of each warehouse storage, and other information. In some embodiments, in order to enhance the accuracy of navigation, the picking robot can use a laser scanner or a camera to identify specific markers or features in the environment and match them with the information in the map, thereby determining its own location.

[0043] In some embodiments, the picking robot can use the target item storage location information in combination with its own positioning system and navigation algorithm to plan an optimal movement path to quickly and accurately reach the location of the item to be picked for performing the picking task. In some embodiments, the navigation system can combine various techniques such as positioning techniques (e.g., GPS, laser positioning, visual positioning, etc.), path planning algorithms (e.g., A* algorithm, Dijkstra algorithm, RRT algorithm, etc.), and obstacle avoidance strategies, etc. to enable autonomous navigation of the picking robot in a complex warehouse environment. In some embodiments, factors such as distance, time, energy consumption, etc. can also be considered when planning the movement path to ensure that the picking robot can efficiently and safely reach the picking location.

[0044] With continued reference to FIG. 3, at 330, picking is performed according to the picking task. To ensure the stability and accuracy of picking, when the picking robot reaches the target location, it determines the action sequence to be performed by the picking robot arm according to the picking task information and the current state of the picking robot arm. In some embodiments, to more accurately grasp the target item to be picked, an action model can be invoked to calculate the action sequence to be taken by the picking robot arm and the gripper when performing the picking task.

[0045] The process of determining the action sequence using the action model will be described below in conjunction with FIG. 4. FIG. 4 shows a schematic diagram of an example of an action model 400 according to some embodiments of the present disclosure. As shown in FIG. 4, the action model includes at least a task information encoder 412, a state encoder 422, an image encoder 432, an action decoder 416, and an image decoder 426. In some embodiments, the task information encoder 412 can be a pre-trained encoder of text-image pairs, which can be implemented through contrastive language-image pre-training techniques. As described above, the picking task information 410 can include the type, quantity, and image information of the item to be picked. In some embodiments, the picking task information 410 (e.g., “pick item labeled 3 on the second layer of the S1 shelf in the second column of the picking area A” and the image information of the target item) can be fed into the task information encoder 412 to obtain the task information representation 414.

[0046] Referring to FIG. 4, similarly, feeding the current state 420 of the robotic arm into a state encoder 422 can result in a state representation 424. In some embodiments, the state encoder 422 can include multiple multi-layer perceptrons. These multi-layer perceptrons can respectively receive the pose of the picking robotic arm and the state of the gripper. In some embodiments, the pose of the picking robotic arm can be represented as [posX, posY, posZ, toX, toY, toZ], where pos(X, Y, Z) represents the position of the picking robotic arm, and to(X, Y, Z) represents the orientation of the picking robotic arm. Similarly, the state of the gripper can be represented as [0], [1], where 0 is the open state, and 1 is the closed state.

[0047] In some embodiments, the data related to the image in the current state, such as the image 430 picked up by the picking robot, can be picked up by an image pickup device of the picking robot. For example, the image pickup device can be mounted on the arm of the picking robot. In some embodiments, the image 430 picked up by the picking robot further includes a real-time image of the target item to be picked.

[0048] With continued reference to FIG. 4, feeding the image 430 picked up by the picking robot into an image encoder 432 can result in intermediate features related to the image. In some embodiments, the image picked up by the picking robot includes an image of the pose of the picking robotic arm, and further includes a real-time image of the item to be picked at the target location. In some embodiments, the image encoder 432 can first generate a one-off embedding for the input image 430 picked up by the picking robot, and then learn intermediate features representative of the content of the image by masking the portion of the image not covered by the mask. In some embodiments, in order to make the one-off embedding generated by the image encoder 432 have the same length, a perceptron resampler can be used to convert the embeddings of different scales into a uniform length, so that intermediate features of a uniform format can be obtained.

[0049] With continued reference to FIG. 4, feeding the task information representation 414 and the state representation 424 into an action decoder 416 can result in an action sequence 418 to be performed by the picking robotic arm. In some embodiments, the action decoder 416 can include multiple multi-layer perceptrons, one of which can receive action-related features, and the other multi-layer perceptrons can respectively output pose-related features EARM of the picking robotic arm and tool state-related features ETOOL. When the action sequence 418 is generated, the picking robot can perform a specific picking task according to the action sequence 418.

[0050] Continuing to refer to FIG. 4, after the task information representation 414, the state representation 424, and the intermediate features output by the image encoder 432 are input into the image decoder 426, the image prediction 428 about the scene of the output action sequence 418 to be performed can be obtained. In some embodiments, the image decoder 426 can receive the intermediate features output by the image encoder 432, the state representation 424, and the task information representation 414 to obtain a plurality of image blocks to generate the corresponding image prediction 428.

[0051] Returning to FIG. 3, after the action sequence to be performed by the picking robot arm is generated, the picking robot can pick the target items according to the action sequence. In some embodiments, during the picking process, images can also be captured by the vision sensing device or image capturing device of the picking robot, and the action sequence to be performed by the picking robot arm can be updated in real time.

[0052] Continuing to refer to FIG. 3, at 340, the items are loaded and transferred to the packing area. In some embodiments, the picking robot is configured with a loading table, and when the picking robot finishes picking the target items, the picking robot can place the target items on the loading table, and then the picking robot can return to the warehouse area, i.e., the packing area, so that the picked items can be packed and shipped by a worker or a packing robot. At 350, when the picking robot finishes the picking task for this round, the picking robot can wait to receive the next picking task, so that the task can be executed repeatedly, which can improve the operation efficiency of the warehouse logistics and also reduce the operation cost.

[0053] FIG. 5 shows a schematic diagram of an example 500 for training an action model according to some embodiments of the present disclosure. In order to make the arm of the picking robot controlled by the action model have the same effect as a human picking items, the action model can be trained using videos with human actions and picking task information.

[0054] Referring to FIG. 5, a training set of training human videos can be collected first, which describe how a human picks items, especially items of different sizes and types mixed together, using an arm. For example, the training human video 510 can be used as a raw material for training the action model 550. For the training human video 510, two sets of training frames can be extracted. The first set of training frames 520 is a series of consecutive frames representing the initial part of a picking action, and this set of training frames is to make the action model 550 understand the current picking action state. The second set of training frames 530 is a series of frames immediately following the first set of training frames 520, which is the development or change of the picking action after the first set of frames. The second set of training frames will be the target of the prediction of the action model 550, i.e., the action model 550 can predict the content in the second set of frames based on the first set of frames.

[0055] With reference back to FIG. 5, the first set of frames 520 and the training picking task information 540 are fed into the action model 550, and the predicted frames 560 for the second set of training frames can be obtained. In some embodiments, the training picking task information here can be “picking a hexagon on the third layer of shelves”, and the first set of training frames also describe the corresponding scene. At this time, the predicted frames 560 for the second set of training frames need to be compared with the real second set of training frames 530 (ground truth) to obtain the first loss 570. Then, the first loss 570 can be used to update the parameters of the action model 550, and this process can be achieved by means of back propagation and parameter optimization.

[0056] By training the action model 550 by means of human picking actions, the action model can better understand the picking action sequence, so that the trained action model has good practical performance.

[0057] FIG. 6 shows a flowchart of an example process 600 for fine-tuning an action model, according to some embodiments of the present disclosure. In order to make the application of the action model more in line with the actual picking situation, the trained action model 550 can be fine-tuned. As shown in FIG. 6, first, a third set of training frames 620 and a fourth set of training frames 630 after the third set of training frames are extracted from a video 610 of training a picking robot arm, respectively. These training frames can be regarded as image data at different time points of the picking robot during the execution of the picking task, and each set of training frames contains a plurality of consecutive image frames for capturing the action of the picking robot arm and the environmental state of the to-be-picked items.

[0058] With reference back to FIG. 6, the trained action model 650 determines the predicted frames 660 for the fourth set of training frames based on the training picking task information 640 and the third set of training frames 620. The action model 650 can learn a large amount of training data and can predict future image frames according to the current training picking task information 640 and the existing image frames. The training picking task information 640 includes the image, quantity, and other constraints of the target items of the picking task, which can help the action model 650 better understand the scene of the current picking task.

[0059] Referring back to FIG. 6, the action model 650 is then updated based on a second loss 670 between the predicted frames 660 of the fourth set of training frames and the actual fourth set of training frames 630. The second loss here can be a measure of the difference between the predicted result and the actual result, such as mean squared error, etc. By calculating this loss value, the degree of deviation of the prediction of the action model 650 from the actual situation can be determined. Then, by means of an optimization algorithm (such as gradient descent), the parameters of the action model 650 are adjusted according to the loss value, so that the action model 650 can more accurately approach the actual result when predicting next time. By repeating this process, the action model 650 can gradually improve its prediction ability, so as to better adapt to the requirements of the picking task.

[0060] By this method of training the action model based on the training of the person video and then training the action model through the video of the picking robot arm, the accuracy of the action model 650 can be further improved.

[0061] FIG. 7 shows a flowchart of an example process 700 for further fine-tuning the action model, according to some embodiments of the present disclosure. In order to further ensure the performance of the action model, further fine-tuning can be performed on the action model.

[0062] As shown in FIG. 7, the training current state 710 of the training picking robot arm, the third set of training frames 720 that have been extracted, and the training picking task information 640 can be input into the action model 650 that has been fine-tuned to obtain the prediction of the training execution action sequence. Then, the third loss 740 between the training execution action sequence 620 and the prediction 730 of the training execution action sequence is used to continue adjusting the action model 650 that has been fine-tuned.

[0063] In some embodiments, the training current state 710 of the training picking robot arm can include the training current pose of the picking robot arm and the training current state of the gripper at the end of the picking robot arm, which can be represented in the form of a vector, for example, the training current state of the training picking robot arm can be [posX, posY, posZ, toX, toY, toZ], and the state of the gripper can be [0] or [1]. In some embodiments, changes in these current states can also be included, such as changes in the pose of the picking robot arm and changes in whether the gripper is open or closed. In some embodiments, the third set of training frames 720 can depict a picture of the picking robot arm picking a pentacle on a shelf, with the gripper in an open state.

[0064] FIG. 8 illustrates a block diagram of an apparatus 800 for picking an item, according to some embodiments of the present disclosure. As shown in FIG. 8, the apparatus 800 includes a picking task information generation module 802 configured to, in response to a picking robot receiving a picking task, generate picking task information based on the picking task, where the picking task information comprises a number of items required to be picked, item images, and a designated storage location. The apparatus 800 further includes a movement path determination module 804 configured to determine a movement path of the picking robot based on a predefined warehouse map and the designated storage location. The apparatus 800 further includes an action sequence determination module 806 configured to, in response to the picking robot moving to the designated storage location by the movement path, determine, by an action model, an action sequence to be performed by an arm of the picking robot based on a current state of the arm of the picking robot and the picking task information. In addition, the apparatus 800 further includes an item picking module 808 configured to pick, by the picking robot, the items required to be picked based on the determined action sequence.

[0065] In some embodiments, the apparatus 800 further includes an acquisition module configured to acquire the current state of the arm of the picking robot, the current state comprising at least any one of: an image of the arm of the picking robot, a pose of the arm of the picking robot, a state of a gripper configured to the arm of the picking robot, and a change of the pose and the state of the gripper involved in the action to be performed.

[0066] In some embodiments, the action model comprises a task information encoder, a state encoder, and an action decoder, where the action sequence determination module 806 comprises: a task information representation determination module configured to determine, based on the task information encoder, a task information representation of the picking task information; a state representation determination module configured to determine, based on the state encoder, a state representation of the current state of the arm of the picking robot; and a first determination module configured to determine, by the action decoder, the action sequence to be performed by the arm of the picking robot based on the task information representation and the state representation.

[0067] In some embodiments, the action sequence determination module 806 further comprises: an image prediction determination module configured to determine, by the action model, an image prediction of a scene of the action sequence to be performed by the arm of the picking robot based on the picking task information and the current state of the arm of the picking robot.

[0068] In some embodiments, the action model further comprises an image decoder, where the image prediction determination module comprises a second determination module configured to determine, by the image decoder, the image prediction based on the task information representation and the state representation.

[0069] In some embodiments, the apparatus 800 further includes a training module configured to train the action model by training the person videos and the training picking task information, and a fine-tuning module configured to fine-tune the action model by training the picking robot arm videos and the training picking task information.

[0070] In some embodiments, the training module includes a first extraction module configured to extract a first set of training frames and a second set of training frames after the first set of training frames, respectively, from the training person videos, a third determination module configured to determine, by the action model, a prediction of the second set of training frames based on the training parsing task and the first set of training frames, and a first update module configured to update the action model based on a first loss between the prediction of the second set of training frames and the second set of training frames.

[0071] In some embodiments, the fine-tuning module includes a second extraction module configured to extract a third set of training frames and a fourth set of training frames after the third set of training frames, respectively, from the training picking robot arm videos, a fourth determination module configured to determine, by the trained action model, a prediction of the fourth set of training frames based on the training picking task information and the third set of training frames, and a second update module configured to update the action model based on a second loss between the prediction of the fourth set of training frames and the fourth set of training frames.

[0072] In some embodiments, the fine-tuning module includes a first acquisition module configured to acquire a training current state of the training picking robot arm and a training execution action sequence, a fifth determination module configured to determine, by the trained action model, a prediction of the training execution action sequence based on the training current state, the training picking task information, and the third set of training frames, and a third update module configured to update the action model based on a third loss between the prediction of the training execution action sequence and the training execution action sequence.

[0073] In some embodiments, the apparatus 800 further includes a loading module configured to load the item to a holding table of the picking robot in response to detecting that the picking of the required picking item by the picking robot is completed, and a returning module configured to return to a source point based on the movement path, the source point being a departure point of the picking robot.

[0074] In some embodiments, the movement path determination module 804 further includes a sixth determination module configured to traverse each node in the predefined warehouse map to determine the movement path.

[0075] In some embodiments, the apparatus 800 further includes a seventh determination module configured to determine the location of the required picking item based on image recognition.

[0076] It can be understood that at least one of the many advantages as can be achieved by the method or process described above can be achieved by utilizing the device 800 of the present disclosure. For example, the picking robot is not only able to directly locate and navigate to the location of the item, but also to accurately pick out the item that matches the picking task, thereby reducing the picking cost and improving the experience of managing warehouse logistics.

[0077] FIG. 9 shows a block diagram of a device 900 that can implement various embodiments of the present disclosure. The device 900 may, for example, be a processing unit of the picking robot 102 as shown in FIG. 1. As shown in FIG. 9, the device 900 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 901 that can perform various appropriate actions and processes in accordance with computer program instructions stored in a read-only memory (ROM) 902 or loaded into a random access memory (RAM) 903 from a storage unit 908. Various programs and data required for operation of the device 900 can also be stored in the RAM 903. The CPU / GPU 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904. Although not shown in FIG. 9, the device 900 can also include a coprocessor.

[0078] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906 such as a keyboard, a mouse, etc., an output unit 907 such as various types of displays, speakers, etc., a storage unit 908 such as a magnetic disk, an optical disk, etc., and a communication unit 909 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0079] The various methods or processes described above can be performed by the CPU / GPU 901. For example, in some embodiments, the methods can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the CPU / GPU 901, one or more steps or actions of the methods or processes described above can be performed.

[0080] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions embodied therewith.

[0081] A computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0082] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0083] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0084] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0085] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0086] The computer program product of the second aspect can include a computer readable storage medium. The computer readable storage medium can include instructions. The instructions can include one or both of: instructions for causing a computer to enable a user equipment device to receive a configuration message from a base station, the configuration message comprising an indication of a set of one or more parameters for a first type of hybrid automatic repeat request process, the first type of hybrid automatic repeat request process being associated with a first type of data; and instructions for causing a computer to enable a user equipment device to receive a configuration message from a base station, the configuration message comprising an indication of a set of one or more parameters for a first type of hybrid automatic repeat request process, the first type of hybrid automatic repeat request process being associated with a first type of data.

[0087] Embodiments of the present disclosure have been described above, with the understanding that these embodiments are exemplary only, and are not restrictive, and are not limited to the disclosed embodiments. Many modifications and changes to the described embodiments are possible, without departing from the scope and spirit of the described embodiments. The selection of terms to be used herein is intended to best explain the principles of the embodiments, practical application, or technical improvement over the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A method for picking items, comprising: in response to a picking robot receiving a picking task, generating picking task information based on the picking task, the picking task information comprising a number of items required to be picked, item images, and a designated storage location; based on a predefined warehouse map and the designated storage location, determining a movement path for the picking robot; in response to the picking robot moving through the movement path to the designated storage location, determining, by an action model, a sequence of actions to be performed by a picking robot arm based on a current state of the picking robot arm and the picking task information; and based on the determined sequence of actions, picking, by the picking robot, the required items.

2. The method of claim 1, further comprising: obtaining the current state of the picking robot arm, the current state comprising at least any one of: an image of the picking robot arm, a pose of the picking robot arm, a state of a gripper configured to the picking robot arm, and a change in the pose and the state of the gripper involved in the actions to be performed.

3. The method of claim 2, the action model comprising a task information encoder, a state encoder, and an action decoder, wherein in response to the picking robot moving through the movement path to the designated storage location, determining, by an action model, a sequence of actions to be performed by a picking robot arm based on a current state of the picking robot arm and the picking task information comprises: based on the task information encoder, determining a task information representation of the picking task information; based on the state encoder, determining a state representation of the current state of the picking robot arm; and based on the task information representation and the state representation, determining, by the action decoder, the sequence of actions to be performed by the picking robot arm.

4. The method of claim 3, further comprising: based on the picking task information and the current state of the picking robot arm, determining, by the action model, an image prediction of a scene of the sequence of actions to be performed by the picking robot arm.

5. The method of claim 4, the action model further comprising an image decoder, wherein based on the picking task information and the current state of the picking robot arm, determining, by the action model, an image prediction of a scene of the sequence of actions to be performed by the picking robot arm comprises: based on the task information representation and the state representation, determining, by the image decoder, the image prediction.

6. The method of claim 1, further comprising: training the action model by training person videos and training picking task information; and fine-tuning the action model by training picking robot arm videos and training picking task information.

7. The method of claim 6, wherein training the action model by training videos and training picking task information comprises: extracting a first set of training frames and a second set of training frames after the first set of training frames from the training person videos, respectively; ​ ​ ​ determining, by the action model, a prediction of the second set of training frames based on the training parsing task and the first set of training frames; and updating the action model based on a first loss between the prediction of the second set of training frames and the second set of training frames. 8.The method of claim 7, wherein fine-tuning the action model by training picking robot arm videos and training picking task information comprises: extracting a third set of training frames and a fourth set of training frames after the third set of training frames from the training picking robot arm videos, respectively; determining, by the trained action model, a prediction of the fourth set of training frames based on the training picking task information and the third set of training frames; and updating the action model based on a second loss between the prediction of the fourth set of training frames and the fourth set of training frames. 9.The method of claim 8, further comprising: obtaining a training current state of the training picking robot arm and a training execution action sequence; determining, by the trained action model, a prediction of the training execution action sequence based on the training current state, the training picking task information, and the third set of training frames; and updating the action model based on a third loss between the prediction of the training execution action sequence and the training execution action sequence. 10.The method of claim 1, further comprising: loading the item to a holding platform of the picking robot in response to detecting that the picking of the item by the picking robot is completed; and returning to a source point based on the movement path, the source point being a departure location of the picking robot. 11.The method of claim 1, wherein determining the movement path of the picking robot based on a predefined warehouse map and the designated storage location comprises: traversing each node in the predefined warehouse map to determine the movement path. 12.The method of claim 1, further comprising: determining a location of the item to be picked based on image recognition. 13.An apparatus for picking an item, comprising: a picking task information generation module configured to generate picking task information based on a picking task in response to a picking robot receiving the picking task, the picking task information comprising a quantity of items to be picked, an image of the item, and a designated storage location; a movement path determination module configured to determine a movement path of the picking robot based on a predefined warehouse map and the designated storage location; an action sequence determination module configured to determine, by an action model, an action sequence to be executed by a picking robot arm based on a current state of the picking robot arm and the picking task information in response to the picking robot moving to the designated storage location via the movement path; and an item picking module configured to pick, by the picking robot, the item to be picked based on the determined action sequence. 14.An electronic device, comprising: a processor; and ​ ​ ​ a memory coupled with the processor, the memory having stored therein instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1-12.

15. A computer program product, tangibly stored on a non-transitory computer- readable medium and comprising machine-executable instructions that, when executed, cause a machine to implement the method of any one of claims 1-12.