Article obtaining method and device, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2024-09-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing robotic devices struggle to perform flexible tasks in complex environments according to user needs, especially lacking flexibility and accuracy when identifying and acquiring target objects.
The robot receives user tasks, uses image acquisition units and environmental information, and combines language and motion models to identify target objects and adjust them to appropriate poses, thereby enabling the acquisition and placement of target objects.
It improves the flexibility and accuracy of robotic equipment in complex environments, enabling it to accurately acquire and place items according to user needs, thereby enhancing the operational efficiency and safety of robotic equipment.
Smart Images

Figure CN122029009A_ABST
Abstract
Description
Method, apparatus, device and medium for acquiring an item TECHNICAL FIELD
[0001] Exemplary implementations of the present disclosure generally relate to the field of robotics, and in particular, to a method, apparatus, device and computer-readable storage medium for acquiring an item using a robot. BACKGROUND
[0002] Robot technology has been rapidly developed and has been widely used in multiple technical fields. Currently, a variety of special-purpose robot devices have been developed, for example, in an industrial environment, robots can be used to perform various tasks such as processing, grabbing, sorting, packaging, etc. For another example, in a home environment, a sweeping robot, a glass wiping robot, etc. have been developed. However, robots can usually only perform pre-set fixed tasks and cannot perform different user tasks according to user needs.
[0003] SUMMARY
[0004] In a first aspect of the present disclosure, a method for acquiring an item is provided. In the method, in response to receiving a user task, a robot device moves to a target location specified by the user task, the user task instructing the robot device to acquire a target item from the target location. The robot device converts to a pose for acquiring the target item. The robot device detects the target item based on attribute information of the target item and environmental information of the location. In response to detecting the target item, the robot device places the target item into a target storage space.
[0005] In a second aspect of the present disclosure, an apparatus for acquiring an item is provided. The apparatus comprises: a moving module configured to, in response to receiving a user task, move a robot device to a target location specified by the user task, the user task instructing the robot device to acquire a target item from the target location; a converting module configured to convert the robot device to a pose for acquiring the target item; a detecting module configured to detect the target item based on attribute information of the target item and environmental information of the location; and a placing module configured to, in response to detecting the target item, place the target item into a target storage space.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.
[0008] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to the first aspect of the present disclosure.
[0009] It is to be understood that the details set forth herein are not intended to limit the key or critical features of the implementations of the present disclosure or to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description, which is given by way of example only. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, aspects, and advantages of various implementations of the present disclosure will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numbers represent like elements throughout. In the drawings:
[0011] FIG. 1 shows a block diagram of an application environment according to one example implementation of the present disclosure;
[0012] FIG. 2 shows a block diagram of acquiring an item by a robotic device according to some implementations of the present disclosure;
[0013] FIG. 3 shows a block diagram of an image collection process according to some implementations of the present disclosure;
[0014] FIG. 4 shows a block diagram of a process of invoking a language model according to some implementations of the present disclosure;
[0015] FIG. 5 shows a block diagram of a process of a user selecting a target item according to some implementations of the present disclosure;
[0016] FIG. 6 shows a block diagram of a process of invoking an action model according to some implementations of the present disclosure;
[0017] FIG. 7A and FIG. 7B respectively show a block diagram of a process of acquiring an item according to some implementations of the present disclosure;
[0018] FIG. 8 shows a flowchart of a method of acquiring an item by a robotic device according to some implementations of the present disclosure;
[0019] FIG. 9 shows a block diagram of an apparatus of acquiring an item by a robotic device according to some implementations of the present disclosure; and
[0020] FIG. 10 shows a block diagram of a device capable of implementing the implementations of the present disclosure. DETAILED DESCRIPTION
[0021] Implementations of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several implementations of the present disclosure are described, it should be understood that the present disclosure can be embodied in various forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and implementations described are merely for illustrative purposes and are not intended to limit the scope of the present disclosure.
[0022] In the description of implementations of the present disclosure, the term "includes" and its derivatives mean "including but not limited to". The term "based on" means "based at least in part on". The term "one implementation" or "the implementation" means "at least one implementation". The term "some implementations" means "at least some implementations". Other explicit or implicit definitions can also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions known at present and / or to be developed in the future.
[0023] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0024] It can be understood that before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scene of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0025] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that executes the operation of the technical solutions of the present disclosure according to the prompt information.
[0026] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending prompt information to the user, for example, can be the way of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0027] It can be understood that the above notification and user authorization obtaining process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0028] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is met. It will be understood that the timing of the execution of the subsequent action performed in response to the event or condition is not necessarily strongly associated with the time at which the event occurs or the condition is established. For example, in some cases, the subsequent action can be performed immediately when the event occurs or the condition is established; in other cases, the subsequent action can be performed after a period of time after the event occurs or the condition is established.
[0029] Example environment
[0030] In recent years, robot technology and machine learning technology have been widely applied to multiple application scenarios. However, robots can usually only perform pre-set fixed tasks and cannot perform different user tasks according to user needs. Especially in complex application environments, robot devices are difficult to determine user needs and then perform corresponding tasks.
[0031] Simple robot devices that perform specific tasks have been developed at present, however, such simple robot devices cannot understand complex user instructions and cannot perform desired tasks according to user instructions in a complex physical space 160. At this time, it is desirable to control the operation of the robot in an effective manner and then perform the desired task.
[0032] According to one example implementation of the present disclosure, a method for obtaining an item by a robot device is proposed. Referring to FIG. 1, an application environment according to one example implementation of the present disclosure is described, and FIG. 1 shows a block diagram 100 of an application environment according to one example implementation of the present disclosure. As shown in FIG. 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform multiple tasks. The physical space 160 can include but is not limited to one or more rooms. For example, in a warehouse environment, the physical space 160 can include one or more warehouses; in a home environment, the physical space 160 can include but is not limited to a living room, a bedroom, a study, a kitchen, a bathroom, etc., or include a combination of one or more of the above.
[0033] As shown in FIG. 1, the robotic device 110 can include multiple parts. For example, a control unit 111 can serve as a control center of the robotic device 110, and an application program can be loaded into the control unit 111 so as to control various parts in the robotic device. A user 120 can use an interaction unit 112 to interact with the robotic device 110, for example, to input control instructions to the robotic device 110 so as to perform a desired task with the robotic device 110. The robotic device 110 can include an arm 113 for performing actions such as grabbing, releasing, etc. For example, the arm 113 can grab a certain item and move the item to a desired location, etc.
[0034] Alternatively and / or additionally, the robotic device 110 can also include a collection unit 114. Here, the collection unit 114 can include various types, for example, an image collection unit, a sound collection unit, etc. Alternatively and / or additionally, the robotic device 110 can further include a sensing unit for detecting surrounding objects, for example, a distance between the robotic device and surrounding objects can be detected based on laser, etc. The robotic device 110 can also include a driving unit 115, for example, the robotic device 110 can be deployed on a movable base, and the driving unit 115 can drive wheels of the base to move along a desired path.
[0035] The physical environment 160 can include one or more collection units 130, …, and 132, for example, one or more image collection devices can be deployed in a room to collect images of the room from various angles. The physical environment 160 can include a control device 140, which can control the one or more collection units 130, …, and 132, etc. via a network (not shown). Alternatively and / or additionally, under a warehouse, the control device 140 can control various devices in the physical space 160, for example, a transportation device for transporting items, a lifting device at a shelf, etc.
[0036] Alternatively and / or additionally, a machine learning model (e.g., the model 150) can be provided to manage the physical space 160. It should be understood that although FIG. 1 shows that the model 150 is located inside the physical space 160, alternatively and / or additionally, the model 150 can be located at a remote device outside the physical space 160, and the control device 140, the robotic device 110, or other devices can access the remote model 150 via a network.
[0037] The model 150 can include one or more models. If the model 150 includes multiple models, the multiple models can include multiple types of models. The model 150 can include, for example, at least a language model (LM) and an action model. The language model can have the ability to answer questions by learning from a large amount of corpus. The action model can control the robot device 110 to perform various actions. The model 150 can further include, for example, an image recognition model, a text recognition model, and the like.
[0038] As shown in FIG. 1, the user 120 can instruct the robot device 110 to operate various items in the physical space 110. For example, the user 120 can instruct the robot device 110 to find a certain item in the physical space 160; for another example, the user 120 can instruct the robot device 110 to place the found item to a designated location, and the like.
[0039] Obtaining a summary of an item
[0040] To at least partially address the deficiencies in the prior art, according to one example implementation of the present disclosure, a method for obtaining an item by a robot device is proposed. A summary according to one example implementation of the present disclosure is described with reference to FIG. 2, which shows a block diagram 200 for obtaining an item by a robot device according to some implementations of the present disclosure.
[0041] As shown in FIG. 2, the robot device 110 in the physical space 160 (also referred to as a first physical space) can receive a user task 210 from the user 120. At this time, the user task 210 can instruct or control the robot device 110 to obtain a target item. For example, the user 120 can speak "get a bottled water from the shelf" in natural language, the target location 230 is "the shelf" in the example of FIG. 2, and the target item 220 is "a bottled water". The robot device 110 can obtain an image of the physical space where the robot device is located. For example, the image can be obtained via at least any one of the acquisition units 114, 130, …, and 132.
[0042] In response to receiving the user task, the robot device moves to the target location specified by the user task, which instructs the robot device to obtain the target item from the target location. The robot device converts to a pose for obtaining the target item, for example, the robot device can rotate to face the shelf. The robot device detects the target item based on attribute information of the target item (e.g., name, image, features, etc. of the target item) and environmental information of the location (e.g., images obtained by the acquisition device 114, 130, 132, etc.). In response to detecting the target item, the robot device places the target item into the target storage space 240.
[0043] According to some implementations of the present disclosure, the above-described method can be executed at any computing device with computing capability. For example, the above-described method can be executed with an application deployed at the robotic device 110. Alternatively and / or additionally, the application can be deployed at the control device 140 for executing the above-described method. Specifically, the powerful processing capability of the model 150 can be invoked for going to the target location. In turn, the robotic device 110 can grasp the item 220 and place it to the storage space 240 (e.g., a storage cart at the robotic device, etc.). Further, the robotic device 110 can return to the location of the user 120 and provide the user 120 with the retrieved item 220. Alternatively and / or additionally, the storage space 240 can further include a storage space in a transportation tool such as a conveyor belt, etc. for transporting the item. At this time, the item 220 can arrive at the designated location via the transportation tool, etc.
[0044] With the example implementations of the present disclosure, the robotic device can perform the user task in a complex physical space. In this way, the robotic device can go to the target location and retrieve the target item according to the instruction. Specifically, the robotic device can find the matching target item according to the attribute information of the target item. In this way, the flexibility and accuracy of the robotic device in performing the task in a complex environment can be improved, thereby completing the intended user task.
[0045] Detailed procedure of acquiring the item
[0046] Having described the outline of acquiring the item, in the following, more details according to some implementations of the present disclosure are described with reference to the accompanying drawings. For ease of description, in the following, more details are described only by way of example of controlling the robotic device to grasp the item in a warehouse environment.
[0047] According to some implementations of the present disclosure, the target location can be represented with a variety of functional ways. Specifically, in response to determining that the target location is represented with a location code, the target location coordinates corresponding to the location code are determined. For example, the location code is, for example, a code of a shelf (e.g., a number of the shelf, a bar code, a QR code, etc.) in a physical space (i.e., a warehouse). At this time, the robotic device can look up a warehouse map for locating the specific coordinates of the shelf (i.e., the target location coordinates). Further, a navigation path between the current location coordinates of the robotic device and the target location coordinates can be determined, thereby causing the robotic device to move along the navigation path to the target location.
[0048] According to some implementations of the present disclosure, an image of the shelf can be taken so as to represent the target location. At this time, in response to determining that the target location is represented with the location image, the robotic device moves in the physical space including the target location to capture an image of the physical space. In response to determining that a location matching the location image is identified in the image, the robotic device moves to the location.
[0049] According to some implementations of the present disclosure, an image of the physical space can be captured by a capturing device. The image can be from at least any of: a capturing device at the robotic device 110, and a capturing device in the physical space, and so on. More details of the image capturing are described with reference to FIG. 3, which shows a block diagram 300 of an image capturing process according to some implementations of the present disclosure. As shown in FIG. 3, an image (e.g., one or more images 310) can be obtained from the capturing unit 114 at the robotic device 110. Since the robotic device 110 can move freely in the physical space 160, the capturing unit 114 can capture images of various locations in the physical space, thereby facilitating finding the target item.
[0050] Alternatively and / or additionally, a first image of the physical space 160 can be obtained from the capturing units 130, …, and 132. Here, the capturing units 130, …, and 132 can be pre-deployed at designated locations within the physical space 160, e.g., at the corners of the ceiling, and so on. In this way, an image of the physical space 160 taken from a top-down perspective can be obtained, thereby facilitating understanding the layout of the physical space 160 as a whole, thereby facilitating locating the target item.
[0051] According to some implementations of the present disclosure, whether an image matches a location image of a target location can be determined based on various manners. For example, a shelf can be identified from an image based on image recognition techniques. Alternatively and / or additionally, a prompt can be constructed and input to a model so as to invoke the processing power of the model to identify a shelf from an image. The prompt can be expressed as, for example, “please identify the shelf represented by the second image from the first image”, and the captured image (corresponding to the first image), the image of the shelf (corresponding to the second image), and the prompt are submitted to the model.
[0052] The model can process the image, and in the case that the image includes a shelf, the model can output a location where the shelf is located (e.g., a region coordinate of the shelf in the image, and / or directly output an image of the region where the shelf is located, etc.). At this time, the robotic device determines that a location matching the location image is recognized from the image, and then the robotic device moves to the location. If the image does not include a shelf, the robotic device can continue to move in the warehouse in order to find the target location. If the robotic device has traversed various locations of the warehouse and has not found the desired shelf, the robotic device can feedback to the user that the target location cannot be found. With some implementations of the present disclosure, whether the image includes a shelf can be detected based on a variety of ways, thereby improving the efficiency of the robotic device positioning.
[0053] The robotic device can move to the target location and transition to a pose for picking up the target item, e.g., the robotic device can rotate to face the shelf. The acquisition unit 114 can continuously acquire images in front of the robotic device in order to determine whether the pose of the robotic device is suitable for grabbing the item at the target location. The robotic device can continuously adjust the pose based on the acquired images. Alternatively and / or additionally, in a warehouse environment, the shelf height can be high, at which time the robotic device can reach a suitable pose for grabbing the acquisition by means of a device such as an elevator, etc. For example, the robotic device can enter the elevator and operate the elevator to reach a suitable height, etc. Alternatively and / or additionally, the robotic device can also send a message to the central control system of the warehouse in order for the central control system to operate the elevator to reach a suitable height, etc.
[0054] According to some implementations of the present disclosure, the environment information of the target location can be acquired by means of a plurality of acquisition units, and then the target item is detected. Specifically, the environment information of the target location includes an image of the target location, and detecting the target item includes identifying the target item from the image based on attribute information. In the example above, bottled water only represents the name of the target item, and assuming that a plurality of brands of bottled water are included on the shelf, the robotic device cannot determine which item is to be grabbed. Here, the attribute information can define more details of the target item, which can be used as a supplement to the target item so that the robotic device can more accurately find the target item.
[0055] More details about detecting the target item are described with reference to FIG. 4, which shows a block diagram 400 of a process of invoking a language model according to some implementations of the present disclosure. As shown in FIG. 4, it can be determined from the user task 210 that the item 220 to be fetched is “bottled water”, at which point the attribute information 430 of the item 220 can be fetched. According to some implementations of the present disclosure, the attribute information includes at least any of the following: a code of the target item, an image, and a feature. For example, the code can be a number of the item in a warehouse inventory, and based on the number, more details about the item can be looked up: a name, a volume, a manufacturer, an image, a product code, etc. In this way, the target item can be accurately found based on the details described above.
[0056] With reference to FIG. 4, the prompt 410 can be generated based on the image 310, the item 220, and the attribute information 430. The prompt 410 can be expressed as, for example, “Please detect ‘bottled water’ from the following image, volume***, manufacturer***, etc.” For another example, the prompt 410 can be expressed as “Please detect ‘bottled water’ from the first image, and the appearance is shown in the second image”, etc. The prompt 410, the image 310, and the corresponding appearance image can be input to the language model 420, so that the language model 420 finds the bottled water from the image 310.
[0057] According to some implementations of the present disclosure, the language model 420 is a model that is trained and fine-tuned, and has rich knowledge of performing tasks in multiple domains. The powerful processing capability and rich knowledge of the model can be utilized to determine the specific location of the target item at the target location. In this way, the efficiency and accuracy of locating the target item can be improved, thereby improving the overall efficiency of performing the user task. Alternatively and / or additionally, the target item can be determined via a dedicated item recognition model.
[0058] Alternatively and / or additionally, a feature (e.g., embedding) of the target item can be fetched. Here, the feature can describe the target item in an overall manner, for example, a dedicated field can be added in a database of the warehouse to store the feature. At this point, the feature can be utilized to more accurately describe the item desired to be fetched, thereby improving the accuracy of fetching the item. Alternatively and / or additionally, features of different granularities can be utilized to describe multiple items that are the same and / or similar. For example, one feature can represent bottled water of a certain brand, another feature can represent bottled water with a volume of 500 ml, and yet another feature can represent bottled water (without distinguishing the brand and the volume), etc. In this way, the one or more target items desired to be fetched can be specified in a more flexible manner.
[0059] According to some implementations of the present disclosure, the feature can be determined using an image encoder of a model. In this way, the image portion in which each item in the image to be processed is located and the image of the target item can be encoded into the same feature space using the same encoder, in this way, the accuracy of the later comparison and identification of the target image can be improved.
[0060] According to some implementations of the present disclosure, an image can be provided to a user, and a user input for the image can be received. The user input can indicate a location of the target item, at which point the attribute information of the target object can be determined based on the user input. More details are described with reference to FIG. 5, which shows a block diagram 500 of a process of a user selecting a target item according to some implementations of the present disclosure. As shown, an image 510 can be provided, which can include a plurality of items 220, 512, etc. located at a shelf. The robotic device can provide the image 510 to a user device 520, and can receive an interaction 530 between the user 120 and the image 510.
[0061] Specifically, the interaction 530 can include a variety of types, for example, the user 120 can use a box selection operation to select an item 220, for example, the user 120 can use a box selection tool to determine the bounding box of the item 220, thereby determining a specific target item. For another example, the user can use a brush tool to determine the mask of the item 220 in order to determine the target item. For another example, the user can use natural language to specify the item 220, at which point the interaction 530 can be expressed as: “grab the green bottle behind the can”, etc. In this way, by providing the user with a real image of the shelf, the user is supported to specify the target object in a more accurate manner.
[0062] According to some implementations of the present disclosure, in the process of obtaining the target item, a segmentation model can be used to determine a set of regions in the image, a region in the set of regions including a candidate item; and in response to receiving a selection of the region, obtaining the candidate item corresponding to the region. Specifically, the image can be divided into a plurality of regions using a pre-trained dedicated segmentation model, or using a language model. For example, the image 510 can be input to the segmentation model in order to determine the bounding box of each item. Alternatively and / or additionally, the image 510 and the corresponding prompt word can be input to the language model. The prompt word may, for example, be expressed as: “please identify each object in the following image”, “please determine the bounding box / mask of each object in the following image”, etc.
[0063] Referring to FIG. 5, a bounding box (as shown by the dashed line box) of each item 220 and 512 can be determined from the image, and the bounding box can be displayed to the user, thereby facilitating the user to select. At this time, the user can click the area where the bounding box is located, thereby specifying the target item. With some implementations of the present disclosure, the user can be prompted in a more explicit manner about the location of each item, thereby improving the accuracy of the user specifying the target item.
[0064] According to some implementations of the present disclosure, in response to not identifying the target item in the image, the robotic device moves the position of at least one item at the target position. Another image of the target position can be acquired, and the target item is identified from the other image. Specifically, the image can be input to the language model, and the next instruction is determined by the language model. For example, the instruction can drive the robotic device to move each item in the shelf in order to find the target item.
[0065] Specifically, the corresponding prompt word can be acquired and the model can be asked how to find the target item. The prompt word may, for example, be expressed as: “please determine whether the following image includes bottled water”, and the prompt word and the corresponding image can be sent to the model. At this time, the model can return: move the items on the shelf to find bottled water. In turn, the robotic device can be instructed to move the items and find bottled water. With some implementations of the present disclosure, the powerful processing capability of the model can be called upon to solve unknown problems in complex environments, thereby determining the action that needs to be performed by the robotic device. In this way, the ability of the robotic device to handle complex tasks can be improved, thereby performing the user task in a more accurate manner.
[0066] According to some implementations of the present disclosure, the action model can be used to determine the specific action performed by the robotic device. More details are described with reference to FIG. 6, which shows a block diagram 600 of a process of calling an action model according to some implementations of the present disclosure. As shown in FIG. 6, an action model 630 can be provided, which can determine the specific action to be performed by the robotic device based on the current state of the robotic device and the instruction, and the action model can be a pre-trained and fine-tuned model.
[0067] It should be appreciated that the current state can include data of multiple aspects, such as an image of the robotic device, an image of the environment of the robotic device, pose data of the robotic arm (e.g., positions of various joints of the robotic arm (POS1,...)), and a state of a tool (e.g., a gripper, a cutter, etc.) fixed at the end of the robotic arm. For example, 0 can be used to represent a closed state of the gripper, and 1 can be used to represent an open state of the gripper. The instruction and the current state can be input to the action model 630, and then an action to be performed by the robotic device is determined based on the instruction and the current state by using the action model. Here, the action can represent a difference between a current pose of the robotic device and a next pose, and a difference between a current state of the tool and a next state, etc.
[0068] The instruction 610 (e.g., "grab the item") can be input to the action model 630, where the instruction 610 can be expressed in natural language and can be determined from the response of the language model. Further, a current state of the robotic device can be obtained, and the action model 630 can determine a corresponding action 640 based on the input data. For example, orientations, positions, velocities, accelerations, etc. of various joints in the arm, and / or wheels and / or other movable devices of the robotic device at a next point in time can be determined. Further, the determined action 640 can be used to control a state of the robotic device at the next point in time.
[0069] With some implementations of the present disclosure, a correlation can be established between the language model and the action model, and a user task initially input by a user in natural language and / or a response returned by the language model can be converted into specific actions executable by the robotic device. In this way, actions of the robotic device can be accurately controlled, and the user task can be performed with higher efficiency.
[0070] According to some implementations of the present disclosure, the robotic device can grab a bottled water. More details are described with reference to FIGS. 7A and 7B, which shows a block diagram 700A of a process of moving an item according to some implementations of the present disclosure. As shown in FIG. 7, in the image 710, the item 220 is not blocked by other items, and the item 220 can be directly grabbed.
[0071] A new instruction and a current state can be input to the action model 630. At this time, the instruction can include "grab the bottled water", and the current state can represent a state of the robotic device in front of the shelf. Further, the action model 630 can generate a new action to control the robotic device to grab the bottled water from the shelf. With some implementations of the present disclosure, a new instruction and a state can be continuously input to the action model, and then subsequent actions are determined.
[0072] According to some implementations of the present disclosure, a constraint 720 can be determined, i.e., a constraint that should be followed during performing the action. For example, it can be determined that the original posture of the bottled water should be maintained (e.g., keep the vertical direction, not be tilted) during moving the bottled water. The constraint 720 can be determined by the language model. For example, a prompt can be generated: “please determine the constraint that should be followed during moving the bottled water based on the following image”, or “please determine the precautions during moving the bottled water”, etc. At this time, the language model can generate the corresponding constraint. The constraint can be input to the action model 630, at this time, the series of actions output by the action model 630 will grasp the bottled water while ensuring the constraint. According to some implementations of the present disclosure, the safety during the operation of the robotic device can be ensured, so as to avoid causing accidental damage to an object, etc.
[0073] According to some implementations of the present disclosure, if the bottled water is blocked by other objects, the object can be removed first, and then the bottled water is retrieved. Specifically, in response to determining that the collected image indicates that the target object is blocked by another object, the other object can be moved to obtain the target object. More details are described with reference to FIG. 7B, which shows a block diagram 700B of a process of moving an object according to some implementations of the present disclosure. As shown in FIG. 7B, the object 220 in the image 730 is the bottled water to be retrieved, and the object 750 is located in front of the object 220 and blocks the object 750. At this time, the robotic device can be instructed to move the object 750 from the position 760 to a position that does not hinder the retrieval of the object 220 (e.g., the position 760’ in the image 740).
[0074] According to some implementations of the present disclosure, a placement position for placing the object 750 can be determined, and the robotic device can be instructed to move the object 750 to the placement position. At this time, the motion model will generate actions to control the robotic device to move the object 750 from the position 760 to the position 760’. In this way, the robotic device can be supported to handle complex problems in a complex environment, so as to perform the user task in a more accurate manner.
[0075] According to some implementations of the present disclosure, during the process of moving the object 750, a constraint during moving the object 750 can be determined based on the posture of the object 750, and the robotic device can be instructed to move the object 750 under the constraint. Similar to the process described above, the object 750 can be moved while complying with the constraint 720. In this way, it can be ensured that each action of the robotic device in a complex environment complies with the safety specification. Alternatively and / or additionally, after placing the object 220 to the storage space, the robotic device can restore the position of the object 750. Specifically, the arm of the robotic device can move the object 750 from the position 760’ back to the position 760.
[0076] According to some implementations of the present disclosure, during the process of grabbing the target item, as the robot arm moves, images of the target item can be captured in real time, and other sensors at the robot device can determine the distance between the arm and the target item in real time, and the pose of the arm is adjusted accordingly, so as to grab the target item.
[0077] According to some implementations of the present disclosure, in response to detecting multiple target items at the target location, a rule for obtaining the target item can be received; and the target item is selected from the multiple target items based on the rule. For example, the robot device can feedback to the user that multiple bottles of water are detected in the shelf, and ask the user how many bottles of water are needed; alternatively and / or additionally, the user can be asked how much capacity of the bottle of water is needed, or which brand of the bottle of water is needed, etc. Then, the robot device can grab one or more bottles of water according to the user's answer.
[0078] According to some implementations of the present disclosure, the target item can include multiple target items, and during the process of placing the target item into the target storage space by the robot device: a sequence for obtaining the multiple target items can be determined based on a rule; and the robot device places the multiple target items into the storage space according to the sequence. Assuming that a box of bottles of water is found at the shelf, the robot device can ask in what order to grab the bottles of water. For example, the rule can indicate grabbing the bottles of water in the order of rows, in the order of columns, and / or in a random order at this time. In this way, the items in the shelf can be ensured to be neat during the grabbing process, thereby avoiding the problem that multiple items placed in disorder increase the difficulty of processing by the robot device.
[0079] According to some implementations of the present disclosure, in response to determining that the target item has been obtained from the target location, the state data of at least one item at the target location is updated. Specifically, the inventory data of the warehouse can be updated, for example, when the target item at a certain shelf is removed, the number of items at the shelf can be updated. Alternatively and / or additionally, the robot device can capture the image of the shelf again, and store the image into the warehouse management database. Alternatively and / or additionally, the specific positions of each item at the shelf can be recorded in detail. For example, if other items are moved during the search for the target item by the robot device, the updated positions of the other items can be stored into the database for later processing by the robot device, etc.
[0080] According to some implementations of the present disclosure, in response to determining that the target item is not detected at the target location, a message can be provided to indicate that the target item is not detected. At this time, the arm of the robot device can be restored to the default position, and the user is reported that the target item is not found.
[0081] According to some implementations of the present disclosure, the user task further includes a destination location. At this time, the robotic device can move to a neighboring location corresponding to the destination location; and the robotic device moves the target item in the storage space to the destination location. For example, the user task can instruct the robotic device to retrieve an item, and place the item on a specified table. At this time, the robotic device can move to the vicinity of the table, and place a bottled water in the cart to the table. In this way, the robotic device can move within the warehouse range, and retrieve one or more items as desired.
[0082] Specifically, the image can be captured in real time, and the user location can be located in the image. Further, the corresponding instruction can be determined based on the current location of the robotic device (e.g., location A) and the destination location (e.g., location B), and in turn, the corresponding instruction. At this time, the instruction can be represented as: move from location A to location B. At this time, the action model 630 will generate a corresponding action, which can control the robotic device 110 to move from location A to location B according to the determined trajectory. In this way, the robotic device 110 completes the user task.
[0083] Alternatively and / or additionally, the target storage space can include a storage space of a transport tool for transporting the item. For example, the robotic device can place the bottled water at a conveyor belt, so that the bottled water reaches the destination location along with the conveyor belt. Alternatively and / or additionally, an image of the target item can be captured, and the image can be associated with the user task. In this way, it can be facilitated to distinguish multiple target items received at the destination location, and dispatch each target item to a corresponding user according to the above association.
[0084] It should be understood that although the above describes one example implementation of the present disclosure in a Chinese language environment. Alternatively and / or additionally, the technical solutions of one example implementation of the present disclosure can be performed in a variety of language environments. For example, the robot can be controlled in a Chinese, English, Japanese, French, etc. environment. Specifically, the robot can be controlled in different language application environments based on the multi-language capabilities provided by the machine learning technology. Further, although the above describes the process of performing the user task using the robotic device with the example of retrieving the bottled water, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as finding other items in the room, placing a certain item to a specified location, etc.
[0085] According to some implementations of the present disclosure, a user can interact with the robotic device via language, motion, gesture, etc. For example, the user can speak a user task that he / she desires to perform, predefine a certain motion to specify the user task, etc. Specifically, the user can specify a motion of holding a water bottle and drinking water, which can be used as a trigger for the robotic device to retrieve a bottled water as a user task. When the motion is recognized from the captured image sequence, the robotic device can automatically ask the user if he / she needs a bottled water, and in case of a positive reply, the robotic device can retrieve a bottled water.
[0086] Alternatively and / or additionally, the user can interact with the robotic device via the interaction unit 112, e.g., the user inputs a task in text and / or image representation, and controls the robotic device to perform the task. Alternatively and / or additionally, the user can specify a condition for the task to be performed, e.g., the task is performed immediately, the task is performed after a predetermined time, or the task is performed upon determining that a predetermined condition is met (e.g., 10 am), etc.
[0087] According to some implementations of the present disclosure, the robotic device can provide a variety of messages to the user, e.g., assuming the robotic device finds bottled waters of multiple brands, it can ask the user which brand he / she needs. As another example, assuming the robotic device does not find bottled waters, and only finds a barrel of water, the robotic device can ask the user if he / she needs the barrel of water, etc. Alternatively and / or additionally, the robotic device can ask the user where the desired item can be found, and go to the location specified by the user to find the desired item. Alternatively and / or additionally, if the desired item cannot be found, the robotic device can ask the user if he / she needs to restock, etc.
[0088] According to some implementations of the present disclosure, a variety of positioning algorithms can be utilized to determine the location of the robotic device, and of various items in the physical environment. For example, a global positioning system (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise location of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, the location of which can be determined by means of signals between the communication unit and a base station, and utilizing a communication network. Alternatively and / or additionally, Wi-Fi access points can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspots to determine the location via the Wi-Fi signal strength and the location of known Wi-Fi access points. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case the location of nearby devices can be determined using Bluetooth signals and known Bluetooth device locations. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and compute the movement and orientation of the device in space, and in turn determine the location of the robotic device.
[0089] Alternatively and / or additionally, a visual positioning system can be used to determine the position of the robotic device and / or the individual items. A map of the physical space can be pre-acquired, and the position of the individual items can be marked in the map. The robotic device can detect the distance from the surrounding items using the echo detection unit, and determine the specific position of the individual items in combination with the acquired images and the map of the physical space. Specifically, computer aided design (CAD) and geographic information system (GIS) can be used, and a positioning algorithm can be used to determine the position. Alternatively and / or additionally, a tracking unit can be deployed at important items in the physical space, for example, a tracking unit can be added at the remote controller of a household appliance (e.g., a TV remote controller, an air conditioner remote controller), so that the robotic device can timely acquire the accurate position of the important item, and the like.
[0090] According to some implementations of the present disclosure, the original position of the robotic device itself and the destination position to which the robotic device is expected to go can be determined based on the methods described above. The robotic device can determine a path from the original position to the destination position. For example, the surrounding environment images can be constantly acquired, and the path can be constantly updated while ensuring to avoid obstacles, and the robotic device can be caused to move along the path to the destination position.
[0091] According to some implementations of the present disclosure, after reaching the destination position, the robotic device can perform a specified task. For example, a specified item can be acquired and moved to a corresponding position. The constraint condition, i.e., the constraint condition that should be followed during the execution of the task, can be determined using the language model and / or the knowledge base. For example, the image and the corresponding prompt word can be acquired, and the image and the prompt word can be input to the language model, and then the constraint condition can be received from the language model. For example, the prompt word can be determined as: “please determine the constraint condition that should be followed during the movement of XXX item based on the following image”, or “please determine the precautions during the movement of XXX item”, and the like.
[0092] At this time, it can be determined that the original posture of the item (e.g., bottled water, beverage, etc.) should be maintained (e.g., the vertical direction should be maintained and the item should not be tilted) during the movement of the item. Further, the constraint condition can be input to the action model, and a series of actions output by the action model will perform the corresponding task while ensuring the constraint condition. Using some implementations of the present disclosure, the safety during the operation of the robotic device can be ensured, so that accidental damage to an item can be avoided, and the like.
[0093] With the example implementations of the present disclosure, a robotic device can perform a user task in a complex physical space. In this way, the robotic device can go to a target location and retrieve a target item according to instructions. Specifically, the robotic device can seek a matching target item according to attribute information of the target item. The flexibility and accuracy of the robotic device in performing a task in a complex environment can be improved, thereby completing an intended user task.
[0094] Example process
[0095] FIG. 8 illustrates a flowchart of a method 800 of retrieving an item by a robotic device, according to some implementations of the present disclosure. At block 810, in response to receiving a user task, the robotic device moves to a target location specified by the user task, the user task instructing the robotic device to retrieve a target item from the target location. At block 820, the robotic device transitions to a pose for retrieving the target item. At block 830, the robotic device detects the target item based on attribute information of the target item and environmental information of the location. At block 840, in response to detecting the target item, the robotic device places the target item into a target storage space.
[0096] According to some implementations of the present disclosure, the environmental information of the target location includes an image of the target location, and detecting the target item includes identifying the target item from the image based on the attribute information.
[0097] According to some implementations of the present disclosure, the method 800 further includes, in response to not identifying the target item in the image, moving a position of at least one item at the target location, obtaining another image of the target location, and identifying the target item from the other image.
[0098] According to some implementations of the present disclosure, the attribute information includes at least any of a code, an image, and a feature of the target item.
[0099] According to some implementations of the present disclosure, obtaining the attribute information includes receiving a user input for the image, the user input indicating a position of the target item.
[0100] According to some implementations of the present disclosure, detecting the target item includes determining a set of regions in the image, a region in the set of regions including a candidate item, and in response to receiving a selection for the region, obtaining the candidate item corresponding to the region.
[0101] According to some implementations of the present disclosure, the method 800 further includes, in response to determining that the target item is not detected at the target location, providing a message to indicate that the target item is not detected.
[0102] According to some implementations of the present disclosure, the method 800 further includes: in response to detecting the plurality of target items at the target location, receiving a rule for picking the target items; and selecting the target items from the plurality of target items based on the rule.
[0103] According to some implementations of the present disclosure, the target items include a plurality of target items, and the placing the target items into the target storage space by the robotic device includes: determining an order for picking the plurality of target items based on the rule; and placing the plurality of target items into the storage space by the robotic device in the order.
[0104] According to some implementations of the present disclosure, the moving to the target location by the robotic device includes: in response to determining that the target location is represented by a location code, determining target location coordinates corresponding to the location code; determining a navigation path between current location coordinates of the robotic device and the target location coordinates; and moving to the target location by the robotic device along the navigation path.
[0105] According to some implementations of the present disclosure, the moving to the target location by the robotic device includes: in response to determining that the target location is represented by a location image, moving by the robotic device in a physical space including the target location to capture an image of the physical space; and in response to determining that a location matching the location image is identified in the image, moving to the location by the robotic device.
[0106] According to some implementations of the present disclosure, the user task further includes a destination location, and the method 800 further includes: moving to a neighboring location corresponding to the destination location by the robotic device; and moving the target items in the storage space to the destination location by the robotic device.
[0107] According to some implementations of the present disclosure, the method 800 further includes: in response to determining that the target items have been picked from the target location, updating state data of at least one item at the target location.
[0108] Example apparatuses and devices
[0109] FIG. 9 illustrates a block diagram of an item picking apparatus 900 by a robotic device according to some implementations of the present disclosure. The apparatus includes: a moving module configured for moving by the robotic device to a target location specified by a user task in response to receiving the user task, the user task instructing the robotic device to pick target items from the target location; a transforming module configured for transforming by the robotic device to a pose for picking the target items; a detecting module configured for detecting by the robotic device the target items based on attribute information of the target items and environment information of the location; and a placing module configured for placing by the robotic device the target items into a target storage space in response to detecting the target items.
[0110] According to some implementations of the present disclosure, the environment information of the target location includes an image of the target location, and the detection module is further configured to: identify the target item from the image based on the attribute information.
[0111] According to some implementations of the present disclosure, the apparatus further includes an operation module configured to: in response to that the target item is not identified in the image, cause the robotic device to move a position of at least one item at the target location; acquire another image of the target location; and identify the target item from the other image.
[0112] According to some implementations of the present disclosure, the attribute information includes at least any of the following: a code of the target item, an image, and a feature.
[0113] According to some implementations of the present disclosure, the detection module is further configured to: receive a user input for the image, the user input indicating a position of the target item.
[0114] According to some implementations of the present disclosure, the detection module is further configured to: determine a set of regions in the image, a region in the set of regions including a candidate item; and in response to receiving a selection for the region, acquire the candidate item corresponding to the region.
[0115] According to some implementations of the present disclosure, the operation module is further configured to: in response to determining that the target item is not detected at the target location, provide a message to indicate that the target item is not detected.
[0116] According to some implementations of the present disclosure, the operation module is further configured to: in response to detecting a plurality of target items at the target location, receive a rule for acquiring the target item; and select the target item from the plurality of target items based on the rule.
[0117] According to some implementations of the present disclosure, the target item includes a plurality of target items, and the placement module is further configured to: determine an order for acquiring the plurality of target items based on the rule; and cause the robotic device to place the plurality of target items into the storage space in the order.
[0118] According to some implementations of the present disclosure, the movement module is further configured to: in response to determining that the target location is represented by a position code, determine a target location coordinate corresponding to the position code; determine a navigation path between a current location coordinate of the robotic device and the target location coordinate; and cause the robotic device to move to the target location along the navigation path.
[0119] According to some implementations of the disclosure, the mobile module is further configured to: in response to determining that the target location is represented with a location image, cause the robotic device to move in a physical space that includes the target location to capture an image of the physical space; and in response to determining that a location matching the location image is identified in the image, cause the robotic device to move to the location.
[0120] According to some implementations of the disclosure, the user task further includes a destination location, and the operation module is further configured to: cause the robotic device to move to a neighboring location corresponding to the destination location; and cause the robotic device to move a target item in the storage space to the destination location.
[0121] According to some implementations of the disclosure, the operation module is further configured to: in response to determining that the target item has been retrieved from the target location, update the state data of at least one item at the target location.
[0122] FIG. 10 illustrates a block diagram of a device 1000 capable of implementing a number of implementations of the present disclosure. It should be understood that the computing device 1000 illustrated by FIG. 10 is merely an example and should not be construed to limit the functionality and scope of the implementations described herein. The computing device 1000 illustrated by FIG. 10 can be used to implement the methods described above.
[0123] As shown in FIG. 10, the computing device 1000 is in the form of a general- purpose computing device. Components of the computing device 1000 can include, but are not limited to, one or more processors or processing units 1010, a memory 1020, a storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. The processing unit 1010 can be a real or virtual processor and is capable of executing various processing according to programs stored in the memory 1020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 1000.
[0124] The computing device 1000 typically includes a plurality of computer storage media. Such media can be removable computer storage media 1030 and / or non-removable computer storage media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Storage media 1030 is typically implemented as one or more computer available storage media 1030 such as volatile computer storage media (e.g., random access memory (RAM)), non-volatile computer storage media (e.g., read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory, or other solid state memory technology), or some combination thereof. The storage media 1030 can be removable and / or non-removable. The storage media 1030 can be machine readable media that can be accessed by the computing device 1000. The storage media 1030 can include a computer program product 1025 that includes one or more programs 1026 configured to carry out the various implementations of the methods or actions outlined herein.
[0125] The computing device 1000 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not specifically shown in FIG. 10, these can include, but are not limited to, magnetic floppy diskettes, other magnetic disks, optical disks such as CD-ROMs, DVDs, or Blu-ray disks, or other removable, non-removable media implemented in a memory technology, a Flash memory, or the like. Such media can be connected by one or more data media interfaces 1042 to the bus 1010. The storage media 1030 can include a computer program product 1025 that includes one or more programs 1026 configured to carry out the various implementations of the methods or actions outlined herein.
[0126] The communication unit 1040 enables communications with other computing devices over a communication media. Additionally, the functionality of the components of the computing device 1000 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. As such, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.
[0127] The input device 1050 can be one or more input devices such as a mouse, a keyboard, a trackball, etc. The output device 1060 can be one or more output devices such as a display, a speaker, a printer, etc. The computing device 1000 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 1040, and with one or more devices that enable a user to interact with the computing device 1000, or any devices (e.g., a network card, a modem, etc.) that enable the computing device 1000 to communicate with one or more other computing devices. Such communication can be carried out through an input / output (I / O) interface (not shown).
[0128] According to example implementations of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.
[0129] Various aspects of the disclosure are now described with reference to the drawings. In general, the drawings described below are diagrammatic and schematic representations of actual or conceptual structures and processes, and are not limiting of the scope of the present disclosure. In the drawings, the same reference numerals are used to represent similar or like items.
[0130] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer readable storage medium having no data signals on it. The instructions can be executed by one or more processors to produce a computer-implemented process such that the instructions, which execute via one or more processors of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0131] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0132] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk, or memory stick, can also be used to implement the present disclosure. The computer program product of the present disclosure can also be provided as a service to download and use the computer program over a network, such as the Internet.
[0133] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and substitutions are possible without departing from the scope and spirit of the described implementations. Many modifications and variations of the present disclosure are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims, the present disclosure can be practiced otherwise than as specifically described. While the present disclosure has been described with reference to the implementation figures, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt to a particular situation and the teachings of the present disclosure to a specific implementation, without departing from the central novel teachings of the application. The implementation(s) illustrated and described herein are meant only to serve as examples. Departures in form and detail are within the scope of the disclosure. Therefore, one skilled in the art can restructure the implementation(s) as needed, while still adhering to the principles of the present disclosure.
Claims
1. A method for retrieving an item, comprising: in response to receiving a user task, a robotic device moving to a target location specified by the user task, the user task instructing the robotic device to retrieve a target item from the target location; the robotic device transitioning to a pose for retrieving the target item; the robotic device detecting the target item based on attribute information of the target item and environment information of the location; and in response to detecting the target item, the robotic device placing the target item into a target storage space. identifying the target item from the image based on the attribute information.
2. The method of claim 1, wherein the environment information of the target location comprises an image of the target location, and detecting the target item comprises:
3. The method of claim 2, further comprising: in response to not identifying the target item in the image, the robotic device moving a location of at least one item at the target location; retrieving another image of the target location; and identifying the target item from the other image.
4. The method of claim 2, wherein the attribute information comprises at least any of: a code, an image, and a feature of the target item. receiving user input for the image, the user input indicating a location of the target item.
5. The method of claim 2, wherein obtaining the attribute information comprises:
6. The method of claim 2, wherein detecting the target item comprises: determining a set of regions in the image, a region in the set of regions comprising a candidate item; and in response to receiving a selection for the region, retrieving the candidate item corresponding to the region. in response to determining that the target item is not detected at the target location, providing a message to indicate that the target item is not detected.
8. The method of claim 1, further comprising: in response to detecting a plurality of target items at the target location, receiving a rule for retrieving the target items; and 7. The method of claim 2, further comprising: selecting the target item from the plurality of target items based on the rule.
9. The method of claim 8, wherein the target item comprises a plurality of target items, and the robotic device placing the target item into the target storage space comprises: determining an order for retrieving the plurality of target items based on the rule; and the robotic device placing the plurality of target items into the storage space in the order.
10. The method of claim 1, wherein the robotic device moving to the target location comprises: in response to determining that the target location is represented with a location code, determining target location coordinates corresponding to the location code; determining a navigation path between current location coordinates of the robotic device and the target location coordinates; and the robotic device moving to the target location along the navigation path.
11. The method of claim 1, wherein the robotic device moving to the target location comprises: in response to determining that the target location is represented with a location image, the robotic device moving in a physical space comprising the target location to capture an image of the physical space; and In response to determining that a location matching the location image is recognized in the image, the robotic device moves to the location. 12.The method of claim 1, wherein the user task further comprises a destination location, and the method further comprises: the robotic device moving to a neighboring location corresponding to the destination location; and the robotic device moving the target item in the storage space to the destination location.
13. The method of claim 1, further comprising: In response to determining that the target item has been retrieved from the target location, updating state data of at least one item at the target location. 14.An apparatus for retrieving an item, comprising: a moving module configured for, in response to receiving a user task, a robotic device moving to a target location specified by the user task, the user task instructing the robotic device to retrieve a target item from the target location; a converting module configured for a robotic device converting to a pose for retrieving a target item; a detecting module configured for a robotic device detecting the target item based on attribute information of the target item and environment information of the location; and a placing module configured for, in response to detecting the target item, the robotic device placing the target item into a target storage space. 15.An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1-13. 16.A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to any one of claims 1-13. 17.A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-13.