Method and apparatus for executing user task, and device and medium
By acquiring physical space images and utilizing language and motion models, robotic devices can autonomously determine objects and operational steps in complex environments, solving the problem that existing robotic devices cannot understand user instructions and achieving more efficient task execution.
Patent Information
- Application Number
- PCT/CN2024/106260
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-22
AI Technical Summary
Existing robotic devices struggle to perform multiple tasks in complex environments according to user needs, especially in understanding complex user instructions and executing corresponding tasks.
By acquiring images of the physical space where the robot is located, language and motion models are used to identify objects associated with the user's task, and operation steps are determined based on the object images. The robot then executes these steps autonomously to complete the task.
It improves the flexibility and accuracy of robotic equipment in complex environments, enabling it to autonomously complete various tasks according to user needs.
Smart Images

Figure CN2024106260_22012026_PF_FP_ABST
Abstract
Description
Methods, apparatus, devices, and media for performing user tasks Technical Field
[0001] The exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks. Background Technology
[0002] Robotics technology has developed rapidly and is widely used in many technological fields. Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging. In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed. However, robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.
[0003] Summary of the Invention
[0004] In a first aspect of this disclosure, a method for performing a user task is provided. In this method, a first image of the physical space in which a robot device is located is acquired; based on the first image, a first object associated with the user task is determined; based on a second image of the first object, a set of steps for manipulating the first object to perform the user task is determined; and the robot device performs the set of steps to perform the user task.
[0005] In a second aspect of this disclosure, an apparatus for performing a user task is provided. The apparatus includes: an acquisition module configured to acquire a first image of a physical space in which a robotic device is located; an object determination module configured to determine a first object associated with the user task based on the first image; a step determination module configured to determine a set of steps for manipulating the first object to perform the user task based on a second image of the first object; and an execution module configured to cause the robotic device to perform the set of steps to perform the user task.
[0006] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0012] Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks;
[0013] Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure;
[0014] Figure 4 shows a flowchart of the process of invoking a language model according to some implementations of this disclosure;
[0015] Figure 5 shows a block diagram of a process for determining a first object from a plurality of first objects according to some implementations of this disclosure;
[0016] Figure 6 shows a block diagram of the process of invoking the action model according to some implementations of this disclosure;
[0017] Figure 7 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure;
[0018] Figure 8 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure; and
[0019] Figure 9 shows a block diagram of a device capable of implementing various implementations of the present disclosure. Detailed Implementation
[0020] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0021] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.
[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0023] It is understood that before using the technical solutions disclosed in each implementation of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0024] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0025] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0027] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
[0028] Example Environment
[0029] In recent years, robotics and machine learning technologies have been widely applied in various scenarios. However, robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs. In particular, in complex application environments, robotic devices struggle to determine user requirements and thus perform corresponding tasks.
[0030] Simple robotic devices have been developed to perform specific tasks. However, these devices cannot understand complex user instructions, nor can they execute the desired tasks according to user commands in complex physical spaces. Therefore, it is desirable to control the robot's operation in an effective way to perform the desired tasks.
[0031] According to an exemplary implementation of this disclosure, a method for performing user tasks is proposed. Referring to Figure 1, which describes an application environment according to an exemplary implementation of this disclosure, Figure 1 shows a block diagram 100 of the application environment according to an exemplary implementation of this disclosure. As shown in Figure 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, bedroom, study, kitchen, toilet, etc., or a combination of one or more of the above.
[0032] As shown in Figure 1, the robot device 110 may include multiple parts. For example, the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device. The user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks. The robot device 110 may include an arm 113 for performing actions such as grasping and releasing. For example, the arm 113 can grasp an object and move it to a desired position, and so on.
[0033] Alternatively and / or additionally, the robot device 110 may also include a data acquisition unit 114. Here, the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc. Alternatively and / or additionally, the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc. The robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.
[0034] Physical environment 160 may include one or more acquisition units 130, ..., and 132. For example, one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles. Physical environment 160 may include control device 140, which can control one or more acquisition units 130, ..., and 132, etc., via a network (not shown). Alternatively and / or additionally, in a smart home environment, control device 140 can control various electrical devices in physical space 160.
[0035] Alternatively and / or additionally, a machine learning model (e.g., model 150) may be provided to manage physical space 160. It should be understood that although Figure 1 shows model 150 located inside physical space 160, alternatively and / or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.
[0036] Model 150 may include one or more models. If model 150 includes multiple models, these multiple models may include multiple types of models. Model 150 may, for example, include at least a language model (LM) and an action model. The language model, by learning from a large corpus, is capable of question answering. The action model can control the robotic device 110 to perform various actions. Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.
[0037] As shown in Figure 1, user 120 can instruct robot device 110 to manipulate various objects in physical space 110. Here, objects can be various items in the home environment. For example, user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.
[0038] Summary of the task to be performed
[0039] To at least partially address the shortcomings of the prior art, a method for performing user tasks is proposed according to an exemplary implementation of this disclosure. Referring to Figure 2, which describes an overview of an exemplary implementation of this disclosure, Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.
[0040] As shown in Figure 2, the robot device 110 in physical space 160 (also referred to as the first physical space) can receive a user task 210 from user 120. The user task 210 can instruct the robot device 110 to perform a series of operations. The robot device 110 can acquire a first image of its own physical space 160. For example, the first image can be acquired via at least any one of acquisition units 114, 130, ..., and 132.
[0041] Based on the first image, a first object associated with user task 210 can be determined. For example, user 120 might say "Give me a cup of coffee" in natural language, and user task 210 would then be "Give me a cup of coffee." For instance, the first object associated with user task 210 can be determined to be a coffee machine (e.g., object 220) based on the first image. A second image of the first object can then be obtained.
[0042] The second image can be acquired in real time. According to some implementations of this disclosure, a second image of the first object can be acquired simultaneously with the first image of the physical space 160 where the robot device 110 is located. For example, a second image of the first object can be cropped from the first image. Alternatively or additionally, according to some implementations of this disclosure, a second image of the first object can also be acquired in response to the robot device 110 moving to the vicinity of the first object. For example, in response to the robot device 110 moving to the vicinity of the first object, a second image of the first object can be acquired using the acquisition unit 114 at the robot device 110.
[0043] Alternatively or additionally, the second image may also be acquired from other devices (e.g., the cloud) or pre-stored locally. In this case, the second image may have been pre-captured. A second image of the first object may be acquired from other devices or locally in response to identifying the first object.
[0044] Based on a second image of the first object, a set of steps for operating the first object to perform user task 210 can be determined. Taking user task 210 as "Give me a cup of coffee" and the first object as a coffee machine as an example, a set of steps for operating the coffee machine to provide coffee to user 120 can be determined based on an image of the coffee machine. Furthermore, the robot device 110 can be instructed to perform a set of steps to execute user task 210. Taking user task 210 as "Give me a cup of coffee" and the first object as a coffee machine as an example, the robot device 110 can be instructed to operate the coffee machine to make coffee and provide the made coffee to user 120.
[0045] According to some implementations of this disclosure, the methods described above can be executed at any computing device with computing capabilities. For example, the methods described above can be executed using an application deployed at robot device 110. Alternatively and / or additionally, an application can be deployed at control device 140 to execute the methods described above. Specifically, the powerful processing capabilities of model 150 can be invoked to determine a first object and a set of steps for manipulating the first object to perform a user task. Subsequently, robot device 110 can travel along path 232 to the coffee machine (i.e., object 220), operate the coffee machine to make coffee, and then travel along path 234 to the location of user 120, and serve the prepared coffee to user 120.
[0046] Using the exemplary implementations of this disclosure, robotic devices can perform user tasks in complex physical spaces. In this way, an object for performing the user task can be identified, and a set of steps for operating that object can be determined automatically. The robotic device can then be instructed to perform this set of steps to execute the user task. This approach improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby completing the intended user task.
[0047] Detailed process of executing the task
[0048] Having described an outline of some implementations according to this disclosure, further details regarding the execution of user tasks will be described below. For ease of description, the following example uses controlling the robotic device 110 to operate a coffee machine to serve coffee to a user as an example to illustrate further details of the execution of user tasks.
[0049] The first image can come from at least one of the following: an acquisition device at the robot device, an acquisition device in the first physical space, or an acquisition device in the second physical space. See Figure 3 for further details of image acquisition, which shows a block diagram 300 of an image acquisition process according to some implementations of this disclosure. As shown in Figure 3, a first image (e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely within the physical space 160, the acquisition unit 114 can acquire images from various locations within the physical space, thereby facilitating the search for the target object.
[0050] Alternatively and / or additionally, a first image of the physical space 160 can be acquired from acquisition units 130, ..., and 132. Here, acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as a corner of the ceiling, etc. In this way, an image of the physical space 160 taken from a top-down angle can be obtained, facilitating an overall understanding of the layout of the physical space 160 and thus facilitating the location of target objects.
[0051] A first object associated with a user task can be determined based on a first image in various ways. According to some implementations of this disclosure, the text corresponding to user task 210 can be determined, and keywords can be extracted from user task 210. Keywords can be, for example, the text in user task 210 used to describe the object. Keywords can be extracted using any appropriate method. For example, keywords can be extracted based on any predetermined rules or algorithms, keywords can be extracted using a model, prompts can be provided to the user to manually determine the keywords, and so on. Taking the extraction of keywords using model 150 as an example, the prompt word could be, for example, "Please determine the keywords in the task 'Give me a cup of coffee'". The prompt word can be provided to model 150 for keyword extraction. Taking user task 210 as "Give me a cup of coffee" as an example, the keyword could be, for example, "coffee". The model output of model 150 could be, for example, "The prompt word is 'coffee'". If the keyword cannot be extracted, model 150 can output a response such as "Extraction failed".
[0052] A first object corresponding to a keyword can be searched within the first image. Any suitable method can be used to search for the first object corresponding to the keyword within the first image. For example, the first object can be searched using a model (e.g., model 150), manually by user 120, or determined using a knowledge base and / or image recognition, etc. For instance, robot device 110 can provide the first image and keywords to user 120 and ask user 120: "What objects in the image correspond to 'coffee,' and where are they?" As another example, a knowledge base can store mapping relationships between keywords and objects (e.g., keyword "coffee" --- object "coffee machine"). A search can be performed in the knowledge base based on the keyword to determine the object corresponding to the keyword, and then the object can be identified from the first image.
[0053] If model 150 is used to determine the first object, model 150 can have the ability to process images. According to some implementations of this disclosure, a cue word can be constructed based on the first image and keywords. This cue word can instruct model 150 to search for an object in the first image that corresponds to the keywords. The cue word can be provided to model 150 to determine the first object. The response of model 150 to the cue word, i.e., the model output of model 150 to the cue word, can indicate the first object.
[0054] For example, a prompt can be generated for searching the first object: "Please search for the object corresponding to the keyword 'coffee' in the image below," to utilize model 150 for object searching. Model 150 can output the location of the first object (e.g., the region coordinates of the first object in the image, and / or directly output the image of the region where the first object is located, etc.). If the image does not contain the first object, model 150 can output a response such as "not found." Thus, the first object in the image can be searched based on multiple methods, thereby improving the performance of the robotic device in searching the location of the first image.
[0055] Referring to Figure 4 for further details regarding the determination of the first object, Figure 4 illustrates a block diagram 400 of the process of invoking a language model according to some implementations of this disclosure. As shown in Figure 4, the keyword 410 "coffee" can be extracted from user task 210, and a corresponding prompt 420 can be generated based on image 310 and keyword 410. The prompt 420 can be represented, for example, as: "Please identify the object corresponding to the keyword 'coffee' from the following images." The prompt 420 and image 310 can be input to language model 430 so that language model 430 can find the object corresponding to 'coffee' from image 310. Language model 430 can be, for example, a model included in model 150.
[0056] According to some implementations of this disclosure, the language model 430 is a trained and fine-tuned model with rich knowledge of performing tasks across multiple domains. The language model 430 can determine that image 310 includes a coffee machine and identify the coffee machine as a first object. Figure 4 is merely illustrative; the language model 430 can process one or more images from different acquisition devices and find the object corresponding to the keyword "coffee."
[0057] According to some implementations of this disclosure, fine-tuning operations can be performed on the model. For example, the robot device 110 can be instructed to pre-collect images of various parts in the first physical space. For instance, the robot device 110 can enter the kitchen, open the refrigerator (or cabinets, drawers), etc., and collect images related to each object. Then, the collected images can be used to fine-tune the model. In this way, the model can determine the specific location of each object, thereby improving the accuracy of identifying the first object.
[0058] Alternatively or additionally, according to some implementations of this disclosure, the first prompt word can also be obtained directly based on the first image and user task 210. This first prompt word can, for example, be used to determine the object from the first image used to perform user task 210. Continuing with the example of user task 210 being "Give me a cup of coffee," the first prompt word could be represented as: "Identify the equipment, appliances, etc. involved in providing coffee from the image." The first prompt word can be provided to a machine learning model (e.g., model 150). The first response of model 150 to the prompt word, i.e., the model output of model 150 to the prompt word, can indicate the first object (e.g., a coffee machine). Thus, the first response of the machine learning model to the first prompt word can be received to determine the first object.
[0059] Further details regarding the identification of the first object are described in Figure 4. As shown in Figure 4, a corresponding prompt 420 can also be generated directly based on image 310 and user task 210. The prompt 420 could be, for example, expressed as: "Identify the equipment, appliances, etc. involved in providing coffee from the image," or "Please identify the object corresponding to the user task 'Give me a cup of coffee' from the following images," etc. The prompt 420 and image 310 can be input into the language model 430 so that the language model 430 can find the object needed to provide coffee from image 310.
[0060] It's important to note that user tasks can include one or more keywords. If only one keyword can be extracted from the user task, the first object corresponding to that keyword is directly searched in the physical space. If multiple keywords can be extracted from the user task, multiple objects corresponding to those keywords can be searched in the physical space. For example, if the user task is "Give me a cup of coffee and a cup of ice water," then the keywords include "coffee" and "ice water." As another example, if the user task is "Give me a cup of iced coffee," then the keywords include "coffee" and "ice." In this case, the object corresponding to "coffee" can be identified as a "coffee machine," and the object corresponding to "ice" can be identified as an "ice maker" or a "refrigerator."
[0061] Alternatively or additionally, a target keyword (which may include at least one keyword) can be determined from these multiple keywords, and the target object corresponding to the target keyword can be determined only from the physical space. It is understood that any appropriate method can be used to determine the target keyword from these multiple keywords. For example, a prompt message can be provided to the user to manually determine the target keyword from these multiple keywords. Another example is that the priority of these multiple keywords can be determined, and the target keyword can be determined based on this priority. Yet another example is that a model can be used to determine the target keyword from these multiple keywords.
[0062] It's also important to note that for each keyword, there can be one or more corresponding first objects in the physical space. If multiple first objects exist in the physical space for each keyword, one can be selected from these first objects to perform the user task. Any suitable method can be used to select the first object from these multiple first objects. For example, a model can be used to select the first object from these multiple first objects. For instance, a prompt phrase like "Please select an object from the following multiple objects to perform the user task" can be generated and provided to the model to determine the first object for performing the user task.
[0063] For example, a first message can be provided to the user to indicate the existence of multiple first objects in physical space. The prompt message may include, for example, the text "Please select one object from the following multiple objects to perform the user's task" along with the names, images, and / or identifiers of each object. A first response from the user to the first message can be received. This first response can indicate the first object selected by the user. Thus, the first object can be determined based on the first response.
[0064] Referring to Figure 5 for further details, Figure 5 illustrates a block diagram 500 of a process for determining a first object from a plurality of first objects according to some implementations of this disclosure. As shown in Figure 5, object 220 (coffee machine) and object 510 (canned coffee) are identified from image 310. A first message, "Please determine the object for serving coffee from the coffee machine and canned coffee shown in the figure," along with image 310, can be provided to the user to instruct the user to determine the object for performing a user task from objects 220 and 510. In this way, the accuracy of locating the first object can be improved, thereby improving the accuracy of performing the user task.
[0065] According to some implementations of this disclosure, after determining the first image, a second image of the first object can be acquired. For example, the second image can be acquired from at least any one of acquisition devices 114, 130, ..., and 132. Based on the second image of the first object, a set of steps for manipulating the first object to perform a user task can be determined. This set of steps can be determined in any suitable manner. For example, a set of steps for manipulating the first object can be retrieved from a knowledge base, a prompt message can be provided to the user to instruct the user to manually determine a set of steps, a model can be used to determine a set of steps, and so on.
[0066] To determine a set of steps using a model, a second cue word can be obtained based on a second image and the user task, used to determine the set of steps from the second image. Taking a coffee machine as an example, the second cue word could be expressed as: "Determine a set of steps to operate the coffee machine shown in the image." The second cue word can be provided to a machine learning model (e.g., model 150), and a second response from the machine learning model can be received in response to the second cue word. The second response can indicate a set of steps for operating the first object. Thus, a set of steps can be determined using a model.
[0067] According to some implementations of this disclosure, in response to determining that the first step in a set of steps is to provide a second object associated with a user task to a first object, the robot device 110 may be instructed to retrieve the second object and to provide the second object to the first object. For example, if the first step is "place the coffee cup under the coffee machine," then the second object can be determined to be "coffee cup." The robot device 110 may be instructed to retrieve the coffee cup and place it under the first object, "coffee machine."
[0068] The robot device 110 can be instructed to move to the location of the second object, thereby acquiring the second object. Specifically, the robot device 110 can determine, according to a pre-acquired map, how to move from its current location to the second location of the second object in physical space (i.e., the location of the coffee cup), and then from the second location to the first location of the first object in physical space (i.e., the location of the coffee machine). For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, determine the path to the second location and the path to the first location.
[0069] In some cases, robot device 110 can directly obtain the second object. For example, robot device 110 can directly go to the location of the coffee cup and obtain it. In other cases, robot device 110 cannot directly obtain the second object, and in such cases, robot device 110 needs to determine the specific method for obtaining the second object. For example, if the coffee cup is placed in a cabinet, robot device 110 needs to determine the specific method for opening the cabinet to obtain the coffee cup. The specific method for obtaining the second object can be determined in any way. For example, a prompt message can be provided to the user to receive user input on the specific method for obtaining the second object. Alternatively, a model can be used to determine the specific method for obtaining the second object. Similarly, any appropriate method can be used to instruct robot device 110 to provide the second object to the first object.
[0070] According to some implementations of this disclosure, corresponding prompts can be constructed to inquire of the model how to obtain the second object and how to provide the second object to the first object. For example, the prompts could be represented as: "Determine how to obtain the coffee cup and how to provide the coffee cup to the coffee machine from the following images." The prompts and corresponding images (e.g., the first image) can be sent to the model. The model can then return: open the cabinet, take out the coffee cup, place the coffee cup under the coffee machine, etc. Utilizing some implementations of this disclosure, the powerful processing capabilities of the model can be invoked to solve unknown problems in complex environments, thereby determining the actions that the robotic device needs to perform. In this way, the ability of the robotic device to handle complex tasks can be improved, thus enabling it to execute user tasks in a more accurate manner.
[0071] According to some implementations of this disclosure, a motion model can be used to determine the specific actions to be performed by the robotic device. See Figure 6 for further details, which shows a block diagram 600 illustrating the process of invoking a motion model according to some implementations of this disclosure. As shown in Figure 6, a motion model 630 can be provided, which can determine the specific actions to be performed by the robotic device based on the current state and instructions of the robotic device. This motion model can be a pre-trained and fine-tuned model. The motion model 630 can also be one of the models included in model 150.
[0072] It should be understood that the current state may include data from multiple aspects, such as an image of the robot device, an image of the robot device's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm. For example, 0 can be used to represent the gripper's closed state, and 1 can be used to represent the gripper's open state. Instructions and the current state can be input into the motion model 630, which then uses the motion model to determine the action to be performed by the robot device based on the instructions and the current state. Here, the action can represent the difference between the robot device's current pose and the next pose, and the difference between the tool's current state and the next state, etc.
[0073] If a coffee cup is placed in a cabinet, an instruction 610 (e.g., "open the cabinet") can be input to the motion model 630. Here, instruction 610 can be expressed in natural language, and the instruction 610 can be determined from the response of the language model. Furthermore, the current state of the robot device can be obtained, and the motion model 630 can determine the corresponding action 640 based on the input data. For example, the orientation, position, speed, acceleration, etc., of the various joints in the arm, and / or the wheels and / or other movable devices of the robot device at the next time point can be determined. Furthermore, the determined action 640 can be used to control the state of the robot device at the next time point.
[0074] Using some implementations of this disclosure, a relationship can be established between the language model and the action model, and the user's initial input, expressed in natural language, can be converted into specific actions that can be performed by the robotic device. In this way, the actions of the robotic device can be precisely controlled, thereby performing the user task more efficiently. Similarly, the action model 630 can be used to determine actions corresponding to the following steps: taking out a coffee cup, placing the coffee cup under the coffee machine, and the robotic device 110 placing the coffee cup under the coffee machine. Furthermore, subsequent steps for the robotic device to operate the coffee machine can be determined.
[0075] According to some implementations of this disclosure, in response to determining that the second step in a set of steps is a candidate control element for operating the first object, a candidate control element can be located in the second image. It is understood that any suitable method can also be used to locate the candidate control element. For example, if a knowledge base stores descriptive information about different objects (e.g., object specifications), the first object can be retrieved from the knowledge base to obtain its descriptive information. The candidate control element can then be determined based on the descriptive information of the first object.
[0076] For example, candidate control elements can be located in the second image using a model, or manually, and so on. Taking a coffee machine as the first object and "starting the coffee machine" as an example, a candidate control element could be, for example, the start button for starting the coffee machine. After determining the candidate control element, the robot device 110 can be instructed to operate the candidate control element. For example, the robot device 110 can be instructed to press the start button of the coffee machine.
[0077] According to some implementations of this disclosure, the found candidate control elements can also be detected. For example, taking a coffee machine as the first object, if the coffee machine is to be operated to make coffee, it can be detected whether the candidate control element is a control element used for making coffee. If the candidate control element is a control element used for making coffee (e.g., a start button), the robot device 110 can be instructed to operate the candidate control element. If the candidate control element is not a control element used for making coffee (e.g., a power off button), it can be determined that the candidate control element was not successfully located.
[0078] According to some implementations of this disclosure, one or more candidate control elements can be determined in the second image. If multiple candidate control elements can be determined in the second image, the operation sequence associated with the multiple candidate control elements can be obtained, thereby instructing the robot device 110 to operate the multiple candidate control elements in the operation sequence. The operation sequence associated with the multiple candidate control elements can be obtained in any suitable manner. For example, the operation sequence associated with the multiple candidate control elements of the first object can be retrieved from a knowledge base, the operation sequence can be determined using a model, the operation sequence can be determined manually, and so on.
[0079] For example, still taking a coffee machine as the first object, if multiple candidate control elements are determined, these multiple candidate control elements may include, for example, a coffee selection element (e.g., for selecting Americano, cappuccino, latte, etc.), a cup size selection element (e.g., for selecting large, medium, small, etc.), a sweetness selection element (e.g., 30% sweetness, 50% sweetness, 70% sweetness, normal, etc.), and so on. If the operation sequence indicates selecting coffee, type, and sweetness in sequence, the robot device 110 can be controlled to operate the coffee selection element, type selection element, sweetness selection element, and start button in sequence to operate multiple candidate control elements to make coffee.
[0080] According to some implementations of this disclosure, a second message associated with multiple candidate control elements can also be provided to the user. The second message can be, for example, expressed as: "What kind of coffee do you want? Large or small? What sweetness?". A second response from the user to the second message can be obtained. The second response can be, for example, "Latte, large, normal sweetness." Thus, the first object can be operated based on the second response.
[0081] Regarding the specific method of operating the first object based on the second answer, according to some implementations of this disclosure, the robot device 110 can be instructed to directly operate multiple candidate control elements based on the second answer. For example, taking the second answer as "I want a latte, large, and normal sweetness", the robot device 110 can be instructed to operate the coffee selection element to select "latte", operate the cup type selection element to select "large", operate the sweetness selection element to select "normal", and finally operate the start button to make coffee that meets the user's requirements.
[0082] Alternatively or additionally, according to some implementations of this disclosure, operation instructions for operating the first object can also be generated based on the second response, so as to operate the first object. For example, voice instructions for operating a coffee machine can be generated based on the second response. In the case of a smart home, voice instructions can be provided to the coffee machine to instruct it to make coffee. For example, voice instructions can be played directly to the coffee machine, which can receive the voice instructions via its own voice acquisition device (e.g., a microphone). Alternatively, the coffee machine can be connected to the internet to receive voice instructions via a network. The coffee machine can then automatically operate multiple control elements based on the received operation instructions to make coffee that meets the user's needs.
[0083] According to some implementations of this disclosure, in response to the completion of an instruction / step, a new image can be acquired and a new prompt word can be constructed to query the language model for the next instruction. The prompt word can be, for example, represented as: "Please determine the next instruction based on the following image," "What to do next," etc. For instance, assuming the coffee machine has already made coffee, the language model can return the instruction "Pick up the coffee cup and provide it to the user." At this point, based on this instruction and the current state of the robot device, a corresponding action can be generated to instruct the robot device to provide the coffee to the user.
[0084] It should be understood that although the foregoing description uses a Chinese language environment as an example to illustrate an exemplary implementation of this disclosure, alternatively and / or additionally, the technical solution of the exemplary implementation of this disclosure can be executed in multiple language environments. For example, the robot can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, the multilingual capabilities provided by machine learning technology can be used to control the robot in application environments of different languages. Furthermore, although the foregoing description uses retrieving bottled water as an example to illustrate the process of using a robotic device to perform user tasks, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as searching for other items in a room, placing an item in a designated location, etc.
[0085] According to some implementations of this disclosure, users can interact with the robotic device through language, actions, gestures, etc. For example, a user can state the user task they wish to perform, predefine a certain action to specify the user task, and so on. Specifically, a user can specify the action of picking up a coffee cup, which can be used as the user task to trigger the robotic device to serve coffee to the user. When the action is recognized from the acquired image sequence, the robotic device can automatically ask the user if they want coffee. If an affirmative answer is received, the robotic device can make coffee and serve it to the user.
[0086] Alternatively and / or additionally, a user can interact with the robot device via the interaction unit 112, for example, the user inputs a task represented by text and / or images, and controls the robot device to perform the task. Alternatively and / or additionally, the user can specify the execution conditions of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when predetermined conditions are determined to be met (e.g., after the user wakes up), etc.
[0087] According to some implementations of this disclosure, the robotic device can provide users with various messages. For example, assuming the robotic device does not find an object associated with coffee (e.g., a coffee machine, canned coffee, etc.) and only finds bottled water, the robotic device can ask the user if they need bottled water, and so on. Alternatively and / or additionally, the robotic device can ask the user where the desired object can be found and go to the location specified by the user to find the desired object. Alternatively and / or additionally, if the desired object cannot be found, the robotic device can ask the user if they want to purchase it, and so on.
[0088] According to some implementations of this disclosure, various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment. For example, a Global Positioning System (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network. Alternatively and / or additionally, a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.
[0089] Alternate and / or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the locations of each object can be marked on this map. The robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, along with positioning algorithms to determine the location. Alternate and / or additional locations can also be used to deploy tracking units at important objects in the physical space; for example, tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.
[0090] According to some implementations of this disclosure, the robot's initial position and desired destination can be determined based on the methods described above. The robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
[0091] According to some implementations of this disclosure, after reaching the destination location, the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location. Constraints, i.e., the constraints that should be followed during task execution, can be determined using a language model and / or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model. For example, prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX," or "Please determine the precautions during the movement of object XXX," etc.
[0092] At this point, it can be determined that during the movement of an object (e.g., bottled water, plate, bowl, etc.), the object's original posture should be maintained (e.g., remaining vertical and not tilted). Furthermore, constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met. Using some implementation methods of this disclosure, safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
[0093] Using the exemplary implementations of this disclosure, robotic devices can perform user tasks in complex physical spaces. In this way, an object for performing the user task can be identified, and a set of steps for operating that object can be determined automatically. The robotic device can then be instructed to perform this set of steps to execute the user task. This approach improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby completing the intended user task.
[0094] Example process
[0095] Figure 7 illustrates a flowchart of a method 700 for performing a user task according to some implementations of this disclosure. At block 710, a first image of the physical space where the robot device is located is acquired. At block 720, based on the first image, a first object associated with the user task is determined. At block 730, based on a second image of the first object, a set of steps for manipulating the first object to perform the user task is determined. At block 740, the robot device performs the set of steps to perform the user task.
[0096] According to some implementations of this disclosure, determining the first object includes: obtaining a first prompt word based on a first image and a user task, the first prompt word being used to determine an object for performing the user task from the first image; and receiving a first response from a machine learning model to the first prompt word in order to determine the first object.
[0097] According to some implementations of this disclosure, determining the first object includes: extracting keywords from the user task; and searching for the first object corresponding to the keywords in the first image.
[0098] According to some implementations of this disclosure, determining the first object includes: in response to determining that multiple first objects exist in the physical space, providing a first message to the user to indicate that multiple first objects exist in the physical space; and receiving a first response from the user to the first message to determine the first object.
[0099] According to some implementations of this disclosure, determining a set of steps for manipulating a first object to perform a user task includes: obtaining a second prompt word based on a second image and the user task, the second prompt word being used to determine a set of steps from the second image; and receiving a second response from a machine learning model to the second prompt word in order to determine a set of steps.
[0100] According to some implementations of this disclosure, a robotic device performs a set of steps including: in response to determining that a first step in the set of steps is to provide a second object associated with a user task to a first object, the robotic device obtains the second object; and the robotic device provides the second object to the first object.
[0101] According to some implementations of this disclosure, a robotic device performs a set of steps including: in response to determining that a second step in the set of steps is to operate a candidate control element in a first object, locating a candidate control element in a second image; and the robotic device operating the candidate control element.
[0102] According to some implementations of this disclosure, the operation of candidate control elements by the robot device includes: in response to determining a plurality of candidate control elements in a second image, obtaining an operation sequence associated with the plurality of candidate control elements; and the robot device operating the plurality of candidate control elements in the operation sequence.
[0103] According to some implementations of this disclosure, operating the first object further includes: providing a user with a second message associated with a plurality of candidate control elements; and operating the first object based on a second response from the user to the second message.
[0104] According to some implementations of this disclosure, operating the first object based on the second response includes at least one of the following: operating a plurality of candidate control elements based on the second response; or generating operation instructions for operating the first object based on the second response, so as to operate the first object.
[0105] Example devices and equipment
[0106] Figure 8 shows a block diagram of an apparatus 800 for performing a user task according to some implementations of the present disclosure. The apparatus 800 includes: an acquisition module 810 configured to acquire a first image of the physical space where a robot device is located; an object determination module 820 configured to determine a first object associated with the user task based on the first image; a step determination module 830 configured to determine a set of steps for manipulating the first object to perform the user task based on a second image of the first object; and an execution module 840 configured to cause the robot device to perform a set of steps to perform the user task.
[0107] According to some implementations of this disclosure, the object determination module 820 is further configured to: obtain a first prompt word based on a first image and a user task, the first prompt word being used to determine an object for performing the user task from the first image; and receive a first response from a machine learning model to the first prompt word in order to determine the first object.
[0108] According to some implementations of this disclosure, the object determination module 820 is further configured to: extract keywords from user tasks; and search for a first object corresponding to the keywords in the first image.
[0109] According to some implementations of this disclosure, the object determination module 820 is further configured to: provide a first message to a user in response to determining that multiple first objects exist in the physical space to indicate that multiple first objects exist in the physical space; and receive a first response from the user to the first message to determine the first object.
[0110] According to some implementations of this disclosure, the step determination module 830 is further configured to: obtain a second prompt word based on a second image and a user task, the second prompt word being used to determine a set of steps from the second image; and receive a second response from a machine learning model to the second prompt word in order to determine a set of steps.
[0111] According to some implementations of this disclosure, the execution module 840 is further configured to: in response to determining that the first step in a set of steps is to provide a second object associated with a user task to a first object, cause the robot device to acquire the second object; and cause the robot device to provide the second object to the first object.
[0112] According to some implementations of this disclosure, the execution module 840 is further configured to: locate the candidate control element in the second image in response to determining that the second step in a set of steps is to operate the candidate control element in the first object; and cause the robot device to operate the candidate control element.
[0113] According to some implementations of this disclosure, the execution module 840 is further configured to: in response to determining a plurality of candidate control elements in a second image, obtain an operation sequence associated with the plurality of candidate control elements; and cause the robot device to operate the plurality of candidate control elements in the operation sequence.
[0114] According to some implementations of this disclosure, the execution module 840 is further configured to: provide a user with a second message associated with a plurality of candidate control elements; and operate the first object based on a second response from the user to the second message.
[0115] According to some implementations of this disclosure, the execution module 840 is further configured to: operate a plurality of candidate control elements based on the second answer; or generate operation instructions for operating a first object based on the second answer, so as to operate the first object.
[0116] Figure 9 shows a block diagram of a device 900 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 900 shown in Figure 9 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 900 shown in Figure 9 can be used to implement the methods described above.
[0117] As shown in Figure 9, the computing device 900 is in the form of a general-purpose computing device. Components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage devices 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 900.
[0118] Computing device 900 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 900, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 930 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible within computing device 900.
[0119] The computing device 900 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0120] The communication unit 940 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 900 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0121] Input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 960 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 900 can also communicate as needed with one or more external devices (not shown) via communication unit 940. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 900, or with any device (e.g., network card, modem, etc.) that enables computing device 900 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0122] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0123] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0124] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0125] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0127] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for performing a user task, comprising: obtaining a first image of a physical space in which a robotic device is located; determining, based on the first image, a first object associated with the user task; determining, based on a second image of the first object, a set of steps for operating the first object in order to perform the user task; and the robotic device performing the set of steps in order to perform the user task.
2. The method of claim 1, wherein determining the first object comprises: obtaining, based on the first image and the user task, a first prompt word for determining, from the first image, an object for performing the user task; and receiving a first response of a machine learning model to the first prompt word in order to determine the first object.
3. The method of claim 1, wherein determining the first object comprises: extracting a keyword from the user task; and searching, in the first image, for the first object that corresponds to the keyword.
4. The method of claim 1, wherein determining the first object comprises: in response to determining that there are multiple first objects in the physical space, providing a first message to the user in order to indicate that there are multiple first objects in the physical space; and receiving a first answer from the user to the first message, determining the first object.
5. The method of claim 1, wherein determining the set of steps for operating the first object in order to perform the user task comprises: obtaining, based on the second image and the user task, a second prompt word for determining, from the second image, the set of steps; and receiving a second response of a machine learning model to the second prompt word in order to determine the set of steps.
6. The method of claim 1, wherein the robotic device performing the set of steps comprises: in response to determining that a first step in the set of steps is to provide, to the first object, a second object associated with the user task, the robotic device obtaining the second object; and the robotic device providing the second object to the first object.
7. The method of claim 6, wherein the robotic device performing the set of steps comprises: in response to determining that a second step in the set of steps is to operate a candidate control element in the first object, the robotic device locating the candidate control element in the second image; and the robotic device operating the candidate control element.
8. The method of claim 7, wherein the robotic device operating the candidate control element comprises: in response to determining, in the second image, multiple candidate control elements, obtaining an operation order associated with the multiple candidate control elements; and the robotic device operating the multiple candidate control elements in the operation order.
9. The method of claim 8, wherein operating the first object further comprises: providing, to the user, a second message associated with the multiple candidate control elements; and receiving a second answer from the user to the second message, determining the operation order. operating the first object based on a second answer of the user to the second message. 10.The method of claim 9, wherein operating the first object based on the second answer comprises at least any of: operating the plurality of candidate control elements based on the second answer; or generating an operation instruction for operating the first object based on the second answer so as to operate the first object. 11.An apparatus for performing a user task, comprising: an obtaining module configured to obtain a first image of a physical space in which a robotic device is located; an object determining module configured to determine, based on the first image, a first object associated with the user task; a step determining module configured to determine, based on a second image of the first object, a set of steps for operating the first object so as to perform the user task; and an execution module configured to cause the robotic device to perform the set of steps so as to perform the user task. 12.An electronic device, comprising: at least one processing unit; and at least one memory that is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 10. 13.A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to any one of claims 1 to 10. 14.A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Full process automatic restaurant service system
CN107180285A
Mobile home robot and controlling method of the mobile home robot
CN111542420A
Methods and systems for food preparation in a robotic cooking kitchen
CN112068526A
Structured data generation method and device and storage medium
CN112925802A
Method, device, equipment and medium for processing visual task by using universal model
CN115830330A