Method and apparatus for executing user task, and device and medium
By acquiring physical space images and using machine learning models to locate and manipulate objects, robotic devices can perform user tasks in complex environments. This solves the problem that existing robotic devices cannot perform complex tasks according to user needs, and improves the flexibility and efficiency of task execution.
Patent Information
- Application Number
- PCT/CN2024/106251
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-22
AI Technical Summary
Existing robotic devices struggle to perform complex tasks based on user needs, especially in complex physical spaces where user requirements cannot be determined and corresponding tasks cannot be executed.
By acquiring images of the physical space where the robot is located, machine learning models are used to locate objects associated with the user's task, and the robot is controlled to operate these objects to complete the user's task, including identifying objects, determining processing methods, and executing corresponding operations.
It improves the flexibility and accuracy of robotic devices when performing tasks, enabling them to associate tasks with user needs in complex environments and optimize task execution efficiency.
Smart Images

Figure CN2024106251_22012026_PF_FP_ABST
Abstract
Description
Methods, apparatuses, devices, and media for performing user tasks TECHNICAL FIELD
[0001] Exemplary implementations of the present disclosure generally relate to the field of robotics, and in particular, to methods, apparatuses, devices, and computer-readable storage media for performing user tasks using a robot. BACKGROUND
[0002] Robotics technology has been rapidly developed and has been widely used in multiple technical fields. Currently, various special-purpose robotic devices have been developed, for example, in an industrial environment, robots can be used to perform various tasks such as machining, grasping, sorting, packaging, etc. For another example, in a home environment, a sweeping robot, a glass wiping robot, etc. have been developed. However, robots can usually only perform pre-set fixed tasks and cannot perform different user tasks according to user needs.
[0003] SUMMARY
[0004] In a first aspect of the present disclosure, a method for performing a user task is provided. In the method, a first image of a physical space in which a robotic device is located is obtained; a first object associated with the user task is located based on the first image; in response to determining that the first image includes a second object for processing the first object, the robotic device operates the second object to process the first object; and the robotic device obtains the processed first object.
[0005] In a second aspect of the present disclosure, an apparatus for performing a user task is provided. The apparatus includes a first image obtaining module configured to obtain a first image of a physical space in which a robotic device is located; a first object locating module configured to locate a first object associated with the user task based on the first image; an executing module configured to cause the robotic device to operate a second object to process the first object in response to determining that the first image includes the second object for processing the first object; and a first object obtaining module configured to obtain the processed first object.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.
[0008] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to the first aspect of the present disclosure.
[0009] It is to be understood that the details set forth herein do not limit the key or critical features of the subject application, which are described in connection with the attached drawings and disclosed claims. Other features of the subject application will be apparent from consideration of the description and drawings in accordance with the principles set forth herein. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, aspects and advantages of various implementations of the present disclosure will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings, in which like reference numerals represent like elements, wherein:
[0011] FIG. 1 illustrates a block diagram of an application environment according to one example implementation of the present disclosure;
[0012] FIG. 2 illustrates a block diagram for performing a user task according to some implementations of the present disclosure;
[0013] FIG. 3 illustrates a block diagram of an image collection process according to some implementations of the present disclosure;
[0014] FIG. 4 illustrates a block diagram of a process for invoking a language model according to some implementations of the present disclosure;
[0015] FIG. 5 illustrates a block diagram of a process for invoking an action model according to some implementations of the present disclosure;
[0016] FIG. 6 illustrates a flowchart of a method for performing a user task according to some implementations of the present disclosure;
[0017] FIG. 7 illustrates a block diagram of an apparatus for performing a user task according to some implementations of the present disclosure; and
[0018] FIG. 8 illustrates a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION
[0019] Implementations of the present disclosure will now be described with reference to the attached drawings. While the present disclosure is susceptible to various modifications and alternative forms, structure and function described herein is shown by way of example in the drawings and is designated by reference number, and will be described in detail. It should be understood that the drawings and detailed description thereto are not intended to limit the scope of the present disclosure, but are intended to be exemplary.
[0020] In the description of the implementations of the present disclosure, the term "comprising" and its similar terms are understood to be open-ended, i.e., "including but not limited to". The term "based on" is understood to mean "based, at least in part, on". The term "one implementation" or "an implementation" is understood to mean "at least one implementation". The term "some implementations" is understood to mean "at least some implementations". Other explicit or implicit definitions can also be included below. As used herein, the term "model" can represent the association relationship between various data. For example, the above association relationship can be obtained based on various technical solutions that are currently known and / or will be developed in the future.
[0021] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.
[0022] It can be understood that, before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the use range, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to the relevant laws and regulations.
[0023] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as the electronic device, the application program, the server or the storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0024] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending the prompt information to the user may, for example, be the way of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0025] It can be understood that the above notification and the process of obtaining the authorization of the user are only illustrative, and do not limit the implementations of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementations of the present disclosure.
[0026] The term "in response to" used herein represents the state that the corresponding event occurs or the condition is met. It will be understood that the timing of the execution of the subsequent action performed in response to the event or the condition is not necessarily strongly associated with the time when the event occurs or the condition is established. For example, in some cases, the subsequent action can be performed immediately when the event occurs or the condition is established; while in other cases, the subsequent action can be performed after a period of time after the event occurs or the condition is established.
[0027] Example Environment
[0028] In recent years, robotic technology and machine learning technology have been widely applied in multiple application scenarios. However, robots are usually only able to perform pre-set fixed tasks and are not able to perform different user tasks according to user needs. Especially in complex application environments, it is difficult for a robot device to determine user needs and then perform corresponding tasks.
[0029] Simple robot devices that perform specific tasks have been developed, however, such simple robot devices are not able to understand complex user instructions and are not able to perform desired tasks according to user instructions in a complex physical space 160. At this time, it is desirable to control the operation of the robot in an effective manner and then perform the desired task.
[0030] According to one example implementation of the present disclosure, a method for performing a user task is proposed. Referring to FIG. 1, an application environment according to one example implementation of the present disclosure is described, and FIG. 1 shows a block diagram 100 of an application environment according to one example implementation of the present disclosure. As shown in FIG. 1, a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform multiple tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, a bedroom, a study, a kitchen, a bathroom, etc., or a combination of one or more of the above.
[0031] As shown in FIG. 1, the robot device 110 can include multiple parts. For example, a control unit 111 can serve as the control center of the robot device 110, and an application program can be loaded into the control unit 111 in order to control various parts of the robot device. The user 120 can use an interaction unit 112 to interact with the robot device 110, for example, input control instructions to the robot device 110 in order to perform a desired task using the robot device 110. The robot device 110 can include an arm 113 for performing actions such as grabbing, releasing, etc. For example, the arm 113 can grab an object and move the object to a desired location, etc.
[0032] Alternatively and / or additionally, the robotic device 110 can further include a collection unit 114. Here, the collection unit 114 can include various types, e.g., an image collection unit, a sound collection unit, etc. Alternatively and / or additionally, the robotic device 110 can further include a sensing unit for detecting surrounding objects, e.g., can detect the distance of the robotic device from surrounding objects based on laser, etc. The robotic device 110 can further include a driving unit 115, e.g., the robotic device 110 can be deployed on a movable base, and the driving unit 115 can drive the wheels of the base to move along a desired path.
[0033] The physical environment 160 can include one or more collection units 130, …, and 132, e.g., one or more image collection devices can be deployed in a room to collect images of the room from various angles. The physical environment 160 can include a control device 140, which can control the one or more collection units 130, …, and 132, etc. via a network (not shown). Alternatively and / or additionally, in a smart home environment, the control device 140 can control various electrical devices in the physical space 160.
[0034] Alternatively and / or additionally, a machine learning model (e.g., the model 150) can be provided to manage the physical space 160. It should be appreciated that although FIG. 1 shows the model 150 located inside the physical space 160, alternatively and / or additionally, the model 150 can be located at a remote device outside the physical space 160, and the control device 140, the robotic device 110, or other devices can access the remote model 150 via a network.
[0035] The model 150 can include one or more models. If the model 150 includes multiple models, the multiple models can include multiple types of models. The model 150 can include at least a language model (LM) and an action model, for example. The language model can have the capability of answering questions by learning from a large amount of corpus. The action model can control the robotic device 110 to perform various actions. The model 150 can further include an image recognition model, a text recognition model, etc., for example.
[0036] As shown in FIG. 1, the user 120 can instruct the robotic device 110 to operate various objects in the physical space 110. Here, the objects can be various items in a home environment, e.g., the user 120 can instruct the robotic device 110 to find a certain object in the physical space 160; as another example, the user 120 can instruct the robotic device 110 to place the found object to a designated location, etc.
[0037] Summary of performing a task
[0038] To at least partially address the deficiencies in the prior art, according to one example implementation of the present disclosure, a method for performing a user task is proposed. An overview of one example implementation of the present disclosure is described with reference to FIG. 2, which illustrates a block diagram 200 for performing a user task according to some implementations of the present disclosure.
[0039] As shown in FIG. 2, the robotic device 110 in the physical space 160 can receive a user task 210 from the user 120. At this time, the user task 210 can instruct the robotic device 110 to obtain a first object 220. For example, the user 120 can speak in natural language “get me two slices of bread”, which corresponds to the first object 220 being “slices of bread” in the example of FIG. 2.
[0040] A first image of the physical space in which the robotic device 110 is located can be obtained. For example, the first image can be obtained via the acquisition unit 114 disposed on the robotic device. Alternatively and / or additionally, the first image can also be obtained via at least any one of the acquisition units 130, …, and 132 disposed in the physical space.
[0041] The first object 220 associated with the user task is located based on the first image. For example, the user issues a user task “get me two slices of bread” to the robotic device. After the robotic device 110 receives the user task, the user task and the first image are sent to the machine learning model, which can determine the slices of bread in the physical space based on the user task and the first image.
[0042] The machine learning model can also determine a second object 230 related to the first object 220 from the first object 220, and the machine learning model can determine from the first image whether the second object 230 is included in the current physical space. For example, after the machine learning model receives the user task including the slices of bread sent by the robotic device, the machine learning model can also determine the second object 230 associated with the slices of bread based on the slices of bread. As an example, the second model can include a toaster, a bread knife, jam, and other tools or food that can be used with the slices of bread.
[0043] In response to determining that the first image includes the second object 230 for processing the first object 220, instructing the robotic device to operate the second object 230 to process the first object 220, the machine learning model determines from the first image whether there is a second object 230 related to the first object 220 in the current physical space. If the machine learning model determines from the first image that there is a second object 230, the machine learning model controls the robotic device to operate the second object 230 to process the first object 220. For example, the machine learning model determines from the first image that there is a toaster in the current physical space, at which time the machine learning model controls the toaster and places the bread slice (i.e., the first object 220) into the toaster for heating.
[0044] In some implementations, the robotic device can operate the bread via the arm 113, for example, the robotic device can trigger a key of the toaster via the arm 113 to cause the toaster to operate. In some other implementations, the robotic device and the toaster can be connected to a smart gateway in the physical space respectively and communicate via a network established by the smart gateway. In this way, the robotic device can directly control the operation of the toaster via the network.
[0045] In some other implementations, the machine learning model determines from the first image that there is jam in the physical space, and the machine learning model can control the robotic device to spread the jam on the bread.
[0046] The robotic device obtains the processed first object 220 and moves the first object 220 according to the user task. For example, the robotic device can move the heated bread slice to a location where the user is located for the user to take.
[0047] It should be understood that the above example in which the robotic device obtains the bread slice and heats the bread slice is only illustrative. In fact, the robotic device can obtain and process any suitable form of object according to the user task, for example, the user can issue a user task instructing the robotic device to obtain a can, and the robotic device can open the can while obtaining the can. The user issues a user task instructing the robotic device to put clothes into a washing machine, and the robotic device can also add laundry detergent into the washing machine while putting the clothes into the washing machine, and so on.
[0048] According to some implementations of the present disclosure, the above-described method can be executed at any computing device having computing capability. For example, the above-described method can be executed by an application deployed at the robotic device 110. Alternatively and / or additionally, the application can be deployed at the control device 140 so as to execute the above-described method.
[0049] With the example implementation of the present disclosure, the robotic device can associate with the task related to the user task while performing the user task, and thereby optimize the user task. In this way, the robotic device can find the desired object in the physical space, and determine in advance the processing manner that the user can need to process the object. In this way, the robotic device can improve the association and expansion of the task while performing the task, to further perform the task related to the user task, and thereby improve the efficiency of the user task delivery.
[0050] Detailed process of performing the task
[0051] Having described the overview of some implementations according to the present disclosure, in the following, more details about performing the user task will be described. For ease of description, in the following, more details about performing the user task will be described only by way of example of controlling the robotic device 110 to retrieve the slice of bread.
[0052] According to some implementations of the present disclosure, the first image can be from at least any of the following: a capturing device at the robotic device, a capturing device in the first physical space, a capturing device in the second physical space. More details about image capturing will be described with reference to FIG. 3, which shows a block diagram 300 of a process of image capturing according to some implementations of the present disclosure. As shown in FIG. 3, the first image (e.g., one or more images 310) of the physical space 160 can be acquired from the capturing unit 114 at the robotic device 110. Since the robotic device 110 can move freely in the physical space 160, the capturing unit 114 can capture images of various locations in the physical space, and thereby facilitate finding the target object.
[0053] Alternatively and / or additionally, the first image of the physical space 160 can be acquired from the capturing units 130, …, and 132. Here, the capturing units 130, …, and 132 can be pre-deployed at designated locations within the physical space 160, e.g., at the corner of the ceiling, etc. In this way, images of the physical space 160 taken from a top-down perspective can be obtained, and thereby facilitate understanding the layout of the physical space 160 as a whole, and thereby facilitate locating the target object.
[0054] According to some implementations of the present disclosure, whether the first image includes the first object 220 can be determined based on various manners. For example, the first object 220 can be recognized from the first image based on image recognition techniques. Alternatively and / or additionally, a prompt can be constructed and input to the model, so as to invoke the processing capability of the model to recognize the first object 220 from the first image. The prompt can be expressed as, for example, “please recognize ‘slice of bread’ from the following image”, and the captured image and the prompt are submitted to the model.
[0055] The machine learning model can process the image, and in the case that the image includes the first object 220, the model can output the location where the object is located (e.g., the region coordinates of the object in the image, and / or directly output the image of the region where the object is located, etc.). If the image does not include the target object, the model can output an answer such as “not found”. With some implementations of the present disclosure, whether the image includes the target object can be detected based on multiple manners, thereby improving the performance of the robotic device in acquiring the target object.
[0056] According to some implementations of the present disclosure, a first hint word is acquired based on the first image and the first object 220, and the first hint word is used to determine the second object 230 for processing the first object 220 from the first image. For example, the first hint word can be “identify the object for processing the bread slice from the image”, and the machine learning model can determine the object for processing the bread slice in the first image according to the first image and the first hint word. In some implementations, the machine learning model can determine multiple candidate objects for processing the bread slice according to the objects associated with the bread slice stored in the knowledge base, and then the machine learning model can determine the object for processing the bread slice in the first image according to the matching degree of the determined multiple candidate objects with the first image. In some implementations, the machine learning model can also identify multiple objects from the first image first, and then identify the association of the multiple objects respectively with the bread slice.
[0057] The first response of the machine learning model to the first hint word is received in order to determine the second object 230. After the machine learning model determines the second object 230 according to the first hint word, a second response is sent to the robotic device. In this way, the robotic device can determine the second object 230 for processing the first object 220 by means of the machine learning model, thereby improving the response speed of the robotic device while improving the accuracy of determining the second object 230.
[0058] According to some implementations of the present disclosure, a keyword can be extracted from the user task. A second object 230 corresponding to the keyword can be searched in the first image. For example, the user task issued by the user includes "help me get two slices of bread", the machine learning model can identify the keyword "slices of bread" from the user task, and determine the second object 230 (such as a toaster, an oven, etc.) related to "slices of bread" according to the keyword "slices of bread". In this way, the second object 230 related to the task can be more accurately determined according to the user task. In some implementations, the keyword can be the first object 220 of the user task. In some other implementations, the keyword can also be other content in the user task. For example, the user issues a user task "help me wash this piece of clothing" to the robot device, in which the first object 220 is the piece of clothing. The robot device can also determine the keyword "wash" from the user task, and further determine the laundry detergent and / or the washing machine as the second object 230 according to the keyword "wash".
[0059] According to some implementations of the present disclosure, in response to determining that there are multiple second objects 230 in the physical space, a first message is provided to the user to indicate that there are multiple second objects 230 in the physical space. Based on a first response to the first message from the user, the first object 220 is determined. The machine learning model determines from the first image that there are multiple second objects 230 in the physical space, and sends information to the robot device to make the robot device send the first message to the user in the form of sound or image.
[0060] For example, when the machine learning model determines from the first image that there are a toaster and an oven in the current physical space, the robot device can send a first message to the user. For example, the first message can include "find the toaster and the oven, which device to use to process the slices of bread?" The robot device can wait to receive a first response to the first message from the user after sending the first message to the user, to determine from the first response to use a second object 230 of the multiple second objects 230 to process the first object 220. For example, if the user replies to the robot device "use the toaster to process the slices of bread" in response to the first message. The machine learning model determines the toaster as the second object 230 according to the first response. In this way, the robot device can send the candidate objects obtained in relation to the first object 220 to the user for the user to select, thereby improving the accuracy of the robot device in completing the user task.
[0061] According to some implementations of this disclosure, a robotic device can be instructed to provide a first object 220 to a second object 230. The second object 230 is then used to process the first object 220. For example, the robotic device can place bread slices into a bread maker and start the bread maker to heat the bread slices. In some implementations, the robotic device can operate the bread using an arm 113; for example, the robotic device can trigger buttons on the bread maker via the arm 113 to drive the bread maker to operate. In some other implementations, the robotic device and the bread maker can be connected to a smart gateway in physical space and communicate via a network established by the smart gateway. In this way, the robotic device can directly control the operation of the bread maker through the network. In this way, the robotic device can control the second object 230 in multiple ways, improving the convenience of the robotic device in completing user tasks and reducing the precision requirements for controlling multiple buttons.
[0062] According to some implementations of this disclosure, the operation mode for operating the second object 230 can be determined based on the second image and the user task. The robot device is instructed to operate the second object 230 according to the operation mode. During operation of the second object 230, the robot device can acquire a second image related to the second object 230 using one or more acquisition units 130, ..., and 132 in the physical space and / or the acquisition unit 114 of the robot device. For example, the second image may include the operation buttons, operation panel, and / or display screen of the second object 230.
[0063] The machine learning model can determine the operation mode for the second object 230 based on the second image, and then the robot device can control the second object 230 based on the operation mode. For example, the robot device can move to the location of the bread maker, determine the bread slice placement slot of the bread maker through the second image, and use the arm 113 to put the acquired bread slice into the bread slice placement slot of the bread maker. Subsequently, the robot device can also use the arm 113 to operate the bread maker to control the bread maker to toast the bread slices. Through the above operations, the robot device can flexibly operate the second object 230, especially in physical spaces without a smart home network, the robot device can also accurately operate the second object 230 to complete user tasks.
[0064] In some other implementations, the machine learning model can also obtain operation instructions for the second object 230 stored in a knowledge base, and control the robotic device to operate the second object 230 according to these operation instructions. For example, the machine learning model can obtain operation instructions from the knowledge base such as "adjust the bread machine to the high setting, work for 30 seconds, and complete the toasting of bread slices," and the machine learning model can operate the bread machine according to the operation instructions for the bread machine in the knowledge base.
[0065] According to some implementations of this disclosure, based on a second image and a user task, a second prompt word is obtained. The second prompt word is used to determine the operation method for the second object 230 in the second image. A second response from a machine learning model to the second prompt word is received to determine the operation method. The second prompt word may include "How do I use the toaster in the picture?" The machine learning model can determine the operation method for the second object 230 based on the second image and the second prompt word. A second response regarding the second prompt word is then sent to the robot device. The robot device can determine the operation method for the second object 230 based on the second response, and then operate the second object 230 according to that operation method.
[0066] According to some implementations of this disclosure, instructions can be sent to a control device for managing the second object 230, instructing the control device to activate the second object 230 to process the first object 220. The robot and the bread maker can each connect to a smart gateway in the physical space and communicate via a network established by the smart gateway. In this way, the robot can directly control the bread maker's operation via the network. Controlling the second object 230 via the network reduces the debugging difficulty between the robot and different second objects 230, thereby improving the robot's versatility.
[0067] According to some implementations of this disclosure, a second message can be provided to the user, inquiring about processing parameters for processing the first object 220 using the second object 230. Based on the user's second response to the second message, the robot device can be instructed to operate the second object 230 to process the first object 220. While the robot device is processing the first object 220 using the second object 230, the robot device can send a second message to the user. Through this second message, the robot device can inquire with the user about specific parameter details for processing the first object 220 using the second object 230, i.e., the user's preference for processing the first object 220 using the second object 230. In this way, the robot device can more accurately complete the user's task, and the robot device's use of the second object 230 to process the first object 220 better meets the user's expectations.
[0068] For example, the second message sent by the robotic device to the user could include "How many slices of bread?", "How long to toast?", etc. The user provides a second response based on this second message, and the machine learning model can determine the appropriate action for the second object 230 based on that response. For instance, if the user answers "Toast two slices, toast for 30 seconds," the machine learning model can use this second response to control the robotic device and adjust its response parameters.
[0069] In some implementations, the user can also provide a second, ambiguous response, such as "bake less" or "bake on a higher heat." The machine learning model can then control the robotic device to perform corresponding operations based on these second responses (e.g., toasting a slice of bread, selecting a high heat setting). Furthermore, the robotic device can record the user's usage habits, allowing the machine learning model to determine how the second object 230 should handle the first object 220.
[0070] According to some implementations of this disclosure, in response to determining that the first object 220 has been processed, the robot device is instructed to provide the processed first object 220 to the user. After the first object 220 has been processed, the robot device can retrieve the first object 220 and move it to the location specified by the user. For example, after a bread machine has finished toasting bread slices, the robot device can remove the bread slices from the bread machine and transport them to the user's location. The robot device can retrieve the first object 220 promptly after the second object 230 has finished processing it, improving the efficiency of the robot device in completing the user's task.
[0071] In some implementations, the robot device can determine whether the bread maker has finished working based on the image acquired by the acquisition unit 114. For example, if a bread slice is detected popping out in the image acquired by the robot device, or a completion interface is displayed on the operation panel, the machine learning model determines that the bread maker has finished working, and then controls the robot device to retrieve the bread slice from the bread maker.
[0072] In some other implementations, if the bread maker has finished its work, it can send a completion signal to the smart home system indicating that the work is complete. Upon receiving the completion signal, the robot device can be controlled by a machine learning model to remove the bread slices from the bread maker.
[0073] According to some implementations of this disclosure, in response to determining that the first image does not include a second object 230 for processing the first object 220, the first object 220 is acquired. If the machine learning model does not identify a second object 230 related to the first object 220 in the first image, the robot learning model can control the robot to directly acquire the first object 220 and move it to the user-specified location. This method can quickly complete the user-specified task, improving the efficiency of user task processing. For example, if a user requests "Give me two slices of bread," after receiving the user task, if the machine learning model does not identify a second object 230 related to the bread slices in the first image, the machine learning model can control the robot to directly acquire the bread slices and transport them to the location where they are needed.
[0074] According to some implementations of this disclosure, a machine learning model may include one or more models. If the machine learning model includes multiple models, these multiple models may include multiple types of models. For example, a machine learning model may include at least a language model (LM) and an action model. The language model, by learning from a large corpus, is capable of question-answering. The action model can control the robotic device 110 to perform various actions. The language model and action model will be described in more detail below; it should be understood that the machine learning model may also include models such as image recognition models and text recognition models.
[0075] Referring to Figure 4 for further details regarding the determination of the second physical space, Figure 4 illustrates a block diagram 400 of the process of invoking a language model according to some implementations of this disclosure. As shown in Figure 4, the first object 220 to be acquired can be determined from the user task 210 as a "bread slice". At this point, a corresponding prompt 410 can be generated based on the first image 310 and the first object 220. The prompt 410 can be expressed, for example, as: "Please identify the bread slice from the following images", or as: "Where is the 'bread slice' placed in the following images", and so on. The prompt 410 and the image 310 can be input into the language model 420 so that the language model 420 can determine the position of the bread slice from the image 310.
[0076] According to some implementations of this disclosure, the language model 420 is a trained and fine-tuned model with rich knowledge of performing tasks across multiple domains. The language model 420 can determine that image 310 includes a second object 230 and determine the location of the second object 230. Figure 4 is merely illustrative; the language model 420 can process one or more images from different acquisition devices. For example, suppose another image includes a toaster, and the toaster can be identified as the second object 230.
[0077] By utilizing some implementation methods of this disclosure, the powerful processing capabilities and rich knowledge of the model can be leveraged to determine the first object 220 and the second object 230 that may be present in the physical space. Subsequently, the robotic device can be instructed to acquire the first object 220, proceed to the location of the second object 230, and process the first object 220 using the second object 230. Compared to existing technical solutions that can only locate target objects in physical space through image recognition, the technical solution of this disclosure can process the first object 220 based on the second object 230 within the physical space after finding the first object 220. In this way, the robotic device can further optimize tasks based on the completed tasks after completing the task, thereby improving the overall efficiency of performing user tasks.
[0078] According to some implementations of this disclosure, the second object 230 can be determined based on image recognition. Specifically, in determining the second object 230 in physical space, multiple candidate objects can be identified using the first image. The function of the candidate objects and their association with the first object 220 are determined based on a knowledge base, and the candidate objects with association are selected as the second object 230. Here, the knowledge base can be predefined and includes the association between physical space and objects. For example, the knowledge base can include: (bread slice, bread machine), (fruit, fruit knife), etc.
[0079] According to some implementations of this disclosure, a motion model can be used to determine the specific actions to be performed by the robot device. See Figure 5 for further details, which shows a block diagram 500 illustrating the process of invoking a motion model according to some implementations of this disclosure. As shown in Figure 5, a motion model 530 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device. This motion model can be a pre-trained and fine-tuned model.
[0080] It should be understood that the current state may include data from multiple aspects, such as an image of the robot device, an image of the robot device's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm. For example, 0 can be used to represent the gripper's closed state, and 1 can be used to represent the gripper's open state. Instructions and the current state can be input into the motion model 530, which then uses the motion model to determine the action to be performed by the robot device based on the instructions and the current state. Here, the action can represent the difference between the robot device's current pose and the next pose, and the difference between the tool's current state and the next state, etc.
[0081] A command 510 (e.g., "take a slice of bread") can be input into the motion model 530. Here, the command 510 can be expressed in natural language, and the command 510 can be determined from the response of the language model. Furthermore, the current state of the robot device can be obtained, and the motion model 530 can determine the corresponding action 540 based on the input data. For example, the orientation, position, speed, acceleration, etc., of each joint in the arm, and / or the wheels and / or other movable devices of the robot device at the next time point can be determined. Furthermore, the determined action 540 can be used to control the state of the robot device at the next time point.
[0082] Using some implementation methods disclosed herein, a relationship can be established between the language model and the action model, and the user's initial input, expressed in natural language, can be converted into specific actions that can be performed by the robotic device. In this way, the actions of the robotic device can be precisely controlled, thereby executing the user task with higher efficiency.
[0083] According to some implementations of this disclosure, after identifying the second object 230, the robot device can acquire images related to the second object 230 via acquisition unit 114 or acquisition units 130, ... 132, etc. The images may include the operation interface of the second object 230, and the machine learning model can control the robot device to operate the second object 230 based on the images.
[0084] According to some implementations of this disclosure, a target position associated with the second object 230 can be determined, and the robotic device can be instructed to move the first object 220 to the target position. For example, the robotic device needs to place a slice of bread into the heating port of a toaster. New instructions and the current state can be input into the motion model 530. The instruction can include "turn on the toaster," and the current state can indicate that the bread slice has been placed into the toaster. Furthermore, the motion model 530 can generate new actions to adjust the toaster's operating time and heating level, etc. Using some implementations of this disclosure, new instructions and states can be continuously input into the motion model to determine subsequent actions. In this way, the robotic device can be supported in handling complex problems in complex environments, thereby performing user tasks more accurately.
[0085] According to some implementations of this disclosure, after the second object 230 processes the first object 220, the robot device can be instructed to retrieve the first object 220 again and proceed to the location specified by the user (e.g., the user's current location). Specifically, images can be acquired in real time, and the user's location can be located within the images. Further, corresponding instructions can be determined based on the robot device's current location (e.g., location A) and the user's location (e.g., location B). The instruction can then be expressed as: move from location A to location B. At this point, the motion model 530 will generate a corresponding action that controls the robot device 110 to move from location A to location B along a determined trajectory. In this way, the robot device 110 completes the task of "getting a slice of bread".
[0086] According to some implementations of this disclosure, in response to determining that no second object 230 is found in the first image, the robot device 110 can be instructed to retrieve the first object 220 and move it to the user's location. Specifically, if no second object 230 is found in the first image, the robot device 110 can be directly instructed to go to the location of the first object 220 and retrieve it. In this case, the robot device 110 directly retrieves the bread slice.
[0087] It should be understood that although the foregoing description uses a Chinese language environment as an example to illustrate an exemplary implementation of this disclosure, alternatively and / or additionally, the technical solution of the exemplary implementation of this disclosure can be executed in multiple language environments. For example, the robot can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, the multilingual capabilities provided by machine learning technology can be used to control the robot in application environments of different languages. Furthermore, although the foregoing description uses retrieving bottled water as an example to illustrate the process of using a robotic device to perform user tasks, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, such as searching for other items in a room, placing an item in a designated location, etc.
[0088] According to some implementations of this disclosure, users can interact with the robot device through language, actions, gestures, etc. For example, users can speak the user task they wish to perform, predefine a certain action to specify the user task, and so on.
[0089] Alternatively and / or additionally, a user can interact with the robotic device via the interaction unit 112. For example, the user inputs a task represented by text and / or images and controls the robotic device to perform the task. Alternatively and / or additionally, the user can specify the execution conditions of the task, such as executing the task immediately, executing the task after a predetermined time, or executing the task when predetermined conditions are determined to be met (e.g., after the user wakes up), etc.
[0090] According to some implementations of this disclosure, various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment. For example, a Global Positioning System (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise position of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network. Alternatively and / or additionally, a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.
[0091] Alternate and / or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the locations of each object can be marked on this map. The robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, along with positioning algorithms to determine the location. Alternate and / or additional locations can also be used to deploy tracking units at important objects in the physical space; for example, tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.
[0092] According to some implementations of this disclosure, the robot's initial position and desired destination can be determined based on the methods described above. The robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
[0093] According to some implementations of this disclosure, after reaching the destination location, the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location. Constraints, i.e., the constraints that should be followed during task execution, can be determined using a language model and / or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model. For example, prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX," or "Please determine the precautions during the movement of object XXX," etc.
[0094] At this point, it can be determined that during the movement of an object (e.g., bottled water, plate, bowl, etc.), the object's original posture should be maintained (e.g., remaining vertical and not tilted). Furthermore, constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met. Using some implementation methods of this disclosure, safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
[0095] Using the exemplary implementation of this disclosure, a robotic device can perform user tasks in a complex physical space. In this way, after acquiring a first object 220, the robotic device can process the first object 220 based on a second object 230 within the physical space. In this way, after completing a task, the robotic device can further perform optimized tasks based on the completed task, thereby completing the expected user task.
[0096] Example process
[0097] Figure 6 illustrates a flowchart of a method 600 for performing a user task according to some implementations of this disclosure. At block 610, a first image of the physical space where the robot device is located is acquired. At block 620, a first object associated with the user task is located based on the first image. At block 630, in response to determining that the first image includes a second object for processing the first object, the robot device manipulates the second object to process the first object. At block 640, the robot device acquires the processed first object.
[0098] According to some implementations of this disclosure, determining the second object includes: obtaining a first prompt word based on a first image and a first object, the first prompt word being used to determine a second object for processing the first object from the first image; and receiving a first response from a machine learning model to the first prompt word in order to determine the second object.
[0099] According to some implementations of this disclosure, determining the second object includes: extracting keywords from the user task; and searching for a second object corresponding to the keywords in the first image.
[0100] According to some implementations of this disclosure, determining the second object includes: in response to determining that multiple second objects exist in the physical space, providing a first message to the user to indicate that multiple second objects exist in the physical space; and determining the first object based on a first response from the user to the first message.
[0101] According to some implementations of this disclosure, operating a second object to process a first object includes: a robotic device providing the first object to the second object; and using the second object to process the first object.
[0102] According to some implementations of this disclosure, using a second object to process a first object includes: determining an operation mode for operating the second object based on a second image and a user task; and a robotic device operating the second object according to the operation mode.
[0103] According to some implementations of this disclosure, determining the operation method for operating the second object includes: obtaining a second prompt word based on the second image and the user task, the second prompt word being used to determine the operation method for the second object in the second image; and receiving a second response from a machine learning model to the second prompt word in order to determine the operation method.
[0104] According to some implementations of this disclosure, using a second object to process a first object includes: sending an instruction to a control device for managing the second object, the instruction being used to instruct the control device to activate the second object to process the first object.
[0105] According to some implementations of this disclosure, the robot device operates the second object to process the first object, including: providing a second message to a user, the second message querying the user for processing parameters of the first object using the second object; and based on the user's second response to the second message, the robot device operates the second object to process the first object.
[0106] According to some implementations of this disclosure, obtaining the processed first object includes: in response to determining that the first object has been processed, the robotic device provides the processed first object to the user.
[0107] According to some implementations of this disclosure, the method 600 further includes: in response to determining that the first image does not include a second object for processing the first object, obtaining the first object.
[0108] Example devices and equipment
[0109] Figure 7 shows a block diagram of an apparatus 700 for performing a user task according to some implementations of the present disclosure. The apparatus 700 includes: a first image acquisition module 710 configured to acquire a first image of the physical space where a robot device is located; a first object localization module 720 configured to localize a first object associated with a user task based on the first image; an execution module 730 configured to, in response to determining that the first image includes a second object for processing the first object, cause the robot device to operate the second object to process the first object; and a first object acquisition module 740 configured to acquire the processed first object.
[0110] According to some implementations of this disclosure, the execution module 730 is further configured to: obtain a first prompt word based on a first image and a first object, the first prompt word being used to determine a second object for processing the first object from the first image; and receive a first response from a machine learning model to the first prompt word in order to determine the second object.
[0111] According to some implementations of this disclosure, the execution module 730 is further configured to: extract keywords from user tasks; and search for a second object corresponding to the keywords in the first image.
[0112] According to some implementations of this disclosure, the execution module 730 is further configured to: provide a first message to a user in response to determining that multiple second objects exist in the physical space to indicate that multiple second objects exist in the physical space; and determine a first object based on a first response from the user to the first message.
[0113] According to some implementations of this disclosure, the execution module 730 is further configured to: show the robot device providing the first object to the second object; and use the second object to process the first object.
[0114] According to some implementations of this disclosure, the execution module 730 is further configured to: determine an operation mode for operating the second object based on the second image and the user task; and cause the robot device to operate the second object according to the operation mode.
[0115] According to some implementations of this disclosure, the execution module 730 is further configured to: obtain a second prompt word based on the second image and the user task, the second prompt word being used to determine the operation mode for the second object in the second image; and receive a second response from a machine learning model to the second prompt word in order to determine the operation mode.
[0116] According to some implementations of this disclosure, the execution module 730 is further configured to: send instructions to a control device for managing a second object, the instructions being used to instruct the control device to start the second object to process the first object.
[0117] According to some implementations of this disclosure, the execution module 730 is further configured to: provide a second message to a user, the second message querying the user for processing parameters of the first object using the second object; and, based on the user's second response to the second message, cause the robot device to operate the second object to process the first object.
[0118] According to some implementations of this disclosure, the first object acquisition module 740 is further configured to: in response to determining that the first object has been processed, cause the robotic device to provide the processed first object to the user.
[0119] According to some implementations of this disclosure, the first object acquisition module 740 is further configured to: acquire the first object in response to determining that the first image does not include a second object for processing the first object.
[0120] Figure 8 shows a block diagram of a device 800 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 800 shown in Figure 8 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 800 shown in Figure 8 can be used to implement the methods described above.
[0121] As shown in Figure 8, the computing device 800 is in the form of a general-purpose computing device. Components of the computing device 800 may include, but are not limited to, one or more processors or processing units 810, memory 820, storage devices 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 800.
[0122] Computing device 800 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 820 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 830 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 800.
[0123] The computing device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG8, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0124] The communication unit 840 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 800 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.
[0125] Input device 850 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 860 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 800 can also communicate as needed with one or more external devices (not shown) via communication unit 840. These external devices, such as storage devices, display devices, etc., can communicate with one or more devices that enable user interaction with computing device 800, or with any device (e.g., network card, modem, etc.) that enables computing device 800 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interfaces (not shown).
[0126] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
[0127] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0128] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0129] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0131] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
A method for performing a user task, comprising: obtaining a first image of a physical space in which a robotic device is located; locating a first object associated with the user task based on the first image; in response to determining that the first image includes a second object for processing the first object, operating the second object by the robotic device to process the first object; and obtaining the processed first object by the robotic device. The method of claim 1, wherein determining the second object comprises: obtaining a first cue based on the first image and the first object, the first cue being used to determine the second object for processing the first object from the first image; and receiving a first response of a machine learning model to the first cue so as to determine the second object. The method of claim 1, wherein determining the second object comprises: extracting a keyword from the user task; and searching for the second object corresponding to the keyword in the first image. The method of claim 1, wherein determining the second object comprises: in response to determining that there are multiple second objects in the physical space, providing a first message to the user so as to indicate that there are multiple second objects in the physical space; and determining the first object based on a first reply from the user to the first message. The method of claim 1, wherein operating the second object to process the first object comprises: providing the first object to the second object by the robotic device; and processing the first object with the second object. The method of claim 5, wherein processing the first object with the second object comprises: determining an operation manner for operating the second object based on a second image and the user task; and operating the second object by the robotic device in accordance with the operation manner. The method of claim 6, wherein determining the operation manner for operating the second object comprises: obtaining a second cue based on the second image and the user task, the second cue being used to determine the operation manner for the second object in the second image; and receiving a second response of a machine learning model to the second cue so as to determine the operation manner. sending an instruction to a control device for managing the second object, the instruction being used to instruct the control device to initiate the second object to process the first object. The method of claim 1, wherein operating the second object to process the first object comprises: providing a second message to the user, the second message asking the user for a processing parameter for processing the first object with the second object; and operating the second object by the robotic device to process the first object based on a second reply from the user to the second message. in response to determining that the first object has been processed, providing the processed first object to the user by the robotic device. in response to determining that the first image does not include the second object for processing the first object, obtaining the first object. The method of claim 1, wherein processing the first object with the second object comprises: An apparatus for performing a user task, comprising: a first image obtaining module configured to obtain a first image of a physical space in which a robotic device is located; a first object locating module configured to locate a first object associated with the user task based on the first image; an executing module configured to, in response to determining that the first image includes a second object for processing the first object, cause the robotic device to operate the second object to process the first object; and an obtaining module configured to obtain the processed first object by the robotic device. The method according to claim 1, wherein obtaining the processed first object comprises: The method according to claim 1, further comprising: a first object obtaining module configured to cause the robotic device to obtain the processed first object. An electronic device comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform a method according to any of claims 1 to 11. A computer-readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement a method according to any of claims 1 to 11. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements a method according to any of claims 1 to 11.
Citation Information
Patent Citations
Full process automatic restaurant service system
CN107180285A
Mobile home robot and controlling method of the mobile home robot
CN111542420A
Methods and systems for food preparation in a robotic cooking kitchen
CN112068526A
Bionic housekeeping robot and control method
CN117325191A
Mechanical arm grabbing method driven by natural language
CN117773920A