Method and apparatus for executing user task, device, and medium

By receiving user tasks and using machine learning models to determine candidate methods, and combining image recognition and motion models to precisely control the operation of robot equipment, the problem of existing robot equipment being unable to flexibly perform tasks has been solved, achieving more efficient and diversified task completion.

WO2026016142A1PCT designated stage Publication Date: 2026-01-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/106257
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing robotic devices struggle to flexibly perform different tasks according to user needs, especially in complex physical spaces where they have difficulty understanding complex user instructions and executing the expected tasks.

Method used

By receiving user tasks, the robot uses machine learning models to determine candidate methods for processing objects, and acquires objects based on user selections, combining image recognition and motion models to precisely control the operation of the robot equipment.

Benefits of technology

It improves the flexibility and accuracy of robot equipment in performing tasks under different needs, enabling it to complete the tasks expected by users in a variety of ways, and enhancing the diversity of user interaction with robot equipment and task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024106257_22012026_PF_FP_ABST
    Figure CN2024106257_22012026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for executing a user task, a device, and a medium. The method for executing a user task comprises the following steps: receiving a user task from a user, the user task instructing a robot device to acquire a first object; determining at least one candidate mode for processing the first object; providing to the user a first message indicating the at least one candidate mode; and on the basis of the user's selection of a candidate mode among the at least one candidate mode, the robot device acquiring the first object. By means of the method, users can select different candidate modes for task execution, thereby completing intended user tasks in more diverse forms.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatuses, devices, and media for performing user tasks TECHNICAL FIELD

[0001] Exemplary implementations of the present disclosure generally relate to the field of robotics, and in particular, to methods, apparatuses, devices, and computer-readable storage media for performing user tasks using a robot. BACKGROUND

[0002] Robotics technology has been rapidly developed and has been widely used in multiple technical fields. Currently, various special-purpose robotic devices have been developed, for example, in an industrial environment, robots can be used to perform various tasks such as processing, grabbing, sorting, packaging, etc. For another example, in a home environment, a sweeping robot, a glass wiping robot, etc. have been developed. However, robots can usually only perform pre-set fixed tasks and cannot perform different user tasks according to user needs.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a method for performing a user task is provided. In the method, a user task is received from a user, the user task instructing a robotic device to acquire a first object. At least one candidate manner for processing the first object is determined. A first message indicating the at least one candidate manner is provided to the user. Based on a selection of a candidate manner in the at least one candidate manner by the user, the robotic device acquires the first object.

[0005] In a second aspect of the present disclosure, an apparatus for performing a user task is provided. The apparatus comprises: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; a determining module configured to determine at least one candidate manner for processing the first object; a message providing module configured to provide a first message indicating the at least one candidate manner to the user; and an acquiring module configured to cause the robotic device to acquire the first object based on a selection of a candidate manner in the at least one candidate manner by the user.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method according to the first aspect of the present disclosure when executed by a processor.

[0009] It should be appreciated that the contents described in this section are not intended to limit the key features or important features of the implementations of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages, and aspects of the implementations of the present disclosure will become more apparent by describing in detail some implementations thereof with reference to the attached drawings in which:

[0011] FIG. 1 shows a block diagram of an application environment according to one example implementation of the present disclosure;

[0012] FIG. 2 shows a block diagram for performing a user task according to some implementations of the present disclosure;

[0013] FIG. 3 shows a block diagram of an image acquisition process according to some implementations of the present disclosure;

[0014] FIG. 4 shows a block diagram of a process of invoking a language model according to some implementations of the present disclosure;

[0015] FIG. 5 shows a block diagram of a process of identifying an object from an image according to some implementations of the present disclosure;

[0016] FIG. 6 shows a block diagram of a process of invoking an action model according to some implementations of the present disclosure;

[0017] FIG. 7 shows a schematic diagram of an example interface according to some implementations of the present disclosure;

[0018] FIG. 8 shows a flowchart of a method for performing a user task according to some implementations of the present disclosure;

[0019] FIG. 9 shows a block diagram of an apparatus for performing a user task according to some implementations of the present disclosure; and

[0020] FIG. 10 shows a block diagram of a device capable of implementing multiple implementations of the present disclosure. DETAILED DESCRIPTION

[0021] Implementations of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several implementations of the present disclosure are described, it should be understood that the present disclosure can be embodied in many other forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and implementations described are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0022] In the description of implementations of the present disclosure, the term "includes" and its derivatives mean "including but not limited to". The term "based on" means "based at least in part on". The term "one implementation" or "the implementation" means "at least one implementation". The term "some implementations" means "at least some implementations". Other explicit or implicit definitions can also be included below. As used herein, the term "model" can represent the relationship between various data. For example, the above-mentioned relationship can be obtained based on various technical solutions known at present and / or to be developed in the future.

[0023] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.

[0024] It can be understood that before using the technical solutions disclosed by the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0025] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will require the acquisition and use of the personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that executes the operation of the technical solutions of the present disclosure according to the prompt information.

[0026] As an optional but non-limiting implementation, in response to receiving the active request of the user, the way of sending prompt information to the user may, for example, be the way of pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0027] It can be understood that the above notification and user authorization process is only illustrative and does not limit the implementations of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementations of the present disclosure.

[0028] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of the execution of a subsequent action performed in response to the event or condition is not necessarily strongly correlated with the time at which the event occurs or the condition is established. For example, in some cases, the subsequent action can be performed immediately upon the occurrence of the event or the establishment of the condition, while in other cases, the subsequent action can be performed after a period of time has elapsed since the occurrence of the event or the establishment of the condition.

[0029] Example Environment

[0030] In recent years, robotic technology and machine learning technology have been widely applied to multiple application scenarios. However, robots are generally only capable of performing pre-set fixed tasks and are not able to perform different user tasks according to user needs. In particular, it is difficult for a robotic device to perform a corresponding task in a manner that is currently needed by a user.

[0031] Simple robotic devices that perform specific tasks have been developed, however, such simple robotic devices are not able to understand complex user instructions or perform desired tasks in a manner desired by a user in a complex physical space 160. At this time, it is desirable to control the operation of the robot in an effective manner and thereby perform a desired task.

[0032] According to one example implementation of the present disclosure, a method for performing a user task is proposed. An application environment according to one example implementation of the present disclosure is described with reference to FIG. 1, which shows a block diagram 100 of an application environment according to one example implementation of the present disclosure. As shown in FIG. 1, a robotic device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robotic device 110 to perform multiple tasks. The physical space 160 can include, but is not limited to, one or more rooms. For example, in a home environment, the physical space 160 can include, but is not limited to, a living room, a bedroom, a study, a kitchen, a bathroom, etc., or a combination of one or more of the above.

[0033] As shown in FIG. 1, the robotic device 110 can include multiple parts. For example, a control unit 111 can serve as a control center of the robotic device 110, and an application program can be loaded into the control unit 111 in order to control various parts of the robotic device. The user 120 can use an interaction unit 112 to interact with the robotic device 110, for example, to input control instructions to the robotic device 110 in order to perform a desired task using the robotic device 110. The robotic device 110 can include an arm 113 for performing actions such as grasping, releasing, etc. For example, the arm 113 can grasp an object and move the object to a desired location, etc.

[0034] Alternatively and / or additionally, the robotic device 110 can further include a collection unit 114. Here, the collection unit 114 can include various types, e.g., an image collection unit, a sound collection unit, etc. Alternatively and / or additionally, the robotic device 110 can further include a sensing unit for detecting surrounding objects, e.g., can detect the distance of the robotic device from surrounding objects based on laser, etc. The robotic device 110 can further include a driving unit 115, e.g., the robotic device 110 can be deployed on a movable base, and the driving unit 115 can drive the wheels of the base to move along a desired path.

[0035] The physical environment 160 can include one or more collection units 130, …, and 132, e.g., one or more image collection devices can be deployed in a room to collect images of the room from various angles. The physical environment 160 can include a control device 140, which can control the one or more collection units 130, …, and 132, etc. via a network (not shown). Alternatively and / or additionally, in a smart home environment, the control device 140 can control various electrical devices in the physical space 160.

[0036] Alternatively and / or additionally, a machine learning model (e.g., the model 150) can be provided to manage the physical space 160. It should be appreciated that although FIG. 1 shows the model 150 located inside the physical space 160, alternatively and / or additionally, the model 150 can be located at a remote device outside the physical space 160, and the control device 140, the robotic device 110, or other devices can access the remote model 150 via a network.

[0037] The model 150 can include one or more models. If the model 150 includes multiple models, the multiple models can include multiple types of models. The model 150 can include at least a language model (LM) and an action model, for example. The language model can have the ability to answer questions by learning from a large amount of corpus. The action model can control the robotic device 110 to perform various actions. The model 150 can further include an image recognition model, a text recognition model, etc., for example.

[0038] As shown in FIG. 1, the user 120 can instruct the robotic device 110 to operate various objects in the physical space 110. Here, the objects can be various items in a home environment, e.g., the user 120 can instruct the robotic device 110 to find a certain object in the physical space 160; as another example, the user 120 can instruct the robotic device 110 to place the found object to a designated location, etc.

[0039] Summary of performing a task

[0040] To at least partially address the deficiencies of the prior art, according to one example implementation of the present disclosure, a method for performing a user task is proposed. A summary according to one example implementation of the present disclosure is described with reference to FIG. 2, which illustrates a block diagram 200 for performing a user task according to some implementations of the present disclosure.

[0041] As shown in FIG. 2, the robotic device 110 in the physical space 160 can receive a user task 210 from the user 120. At this time, the user task 210 can instruct the robotic device 110 to obtain a first object (for ease of description, the first object can be referred to as a target object). For example, the user 120 can speak in natural language "I want to eat an apple", and the first object is "apple" (e.g., the first object 220) in the example of FIG. 2. The robotic device 110 can obtain a space image of the physical space 160 in which the robotic device is located. For example, the space image can be obtained via at least any one of the acquisition units 114, 130, …, and 132.

[0042] In turn, at least one candidate way for processing the first object can be determined. For example, in a real-world scenario, the user can have multiple eating methods for the apple, such as eating directly, eating after peeling, eating after juicing, etc., and thus the eating method that the user can want can be determined. As shown in FIG. 2, assuming that the first object 220 is included in the room, the robotic device 110 can be instructed to move and obtain the first object 220.

[0043] According to some implementations of the present disclosure, the above-described method can be executed at any computing device with computing capability. For example, the above-described method can be executed with an application deployed at the robotic device 110. Alternatively and / or additionally, the application can be deployed at the control device 140 in order to execute the above-described method. Specifically, the powerful processing capability of the model 150 can be invoked in order to determine at least one candidate way for processing the first object 220. In turn, the robotic device 110 can provide a first message indicating the at least one candidate way to the user 120. The first message can be presented to the user 120 via the interaction unit 112. Based on the selection of a candidate way among the at least one candidate way by the user 120, the robotic device 110 can obtain the first object. For example, the user can be directly provided with the apple, can be provided with the apple and a fruit knife, or can be provided with apple juice after juicing, etc.

[0044] With the example implementation of the present disclosure, the robotic device can perform the user task in a plurality of different operation modes. In this way, the robotic device can provide the user with at least one candidate mode, and process the target object in the user-desired mode based on the user's selection. In this way, the flexibility and accuracy of the robotic device in performing the task under different user requirements can be improved, thereby completing the intended user task.

[0045] Detailed procedure of performing the task

[0046] Having described the overview of some implementations according to the present disclosure, in the following, more details about performing the user task will be described. For ease of description, in the following, more details about performing the user task will be described by way of example only with the robotic device 110 processing the apple.

[0047] According to some implementations of the present disclosure, the spatial image can be from at least any of the following: a capturing device at the robotic device and a capturing device in the physical space. More details about image capturing are described with reference to FIG. 3, which shows a block diagram 300 of an image capturing procedure according to some implementations of the present disclosure. As shown in FIG. 3, the spatial image (e.g., one or more images 310) of the physical space 160 can be acquired from the capturing unit 114 at the robotic device 110. Since the robotic device 110 can move freely in the physical space 160, the capturing unit 114 can capture images of various locations in the physical space, thereby facilitating the search for the target object.

[0048] Alternatively and / or additionally, the spatial image of the physical space 160 can be acquired from the capturing units 130, …, and 132. Here, the capturing units 130, …, and 132 can be pre-deployed at designated locations within the physical space 160, e.g., at the corner of the ceiling, etc. In this way, images of the physical space 160 taken from a top-down perspective can be obtained, thereby facilitating an overall understanding of the layout of the physical space 160, thereby facilitating the localization of the target object.

[0049] According to some implementations of the present disclosure, whether the spatial image includes the first object can be determined based on a plurality of ways. For example, the first object can be recognized from the spatial image based on image recognition techniques. Alternatively and / or additionally, a prompt can be constructed and input to the model, so as to invoke the processing capability of the model to recognize the first object from the spatial image. The prompt can be expressed as, for example, "please recognize 'apple' from the following image", and the captured image and the prompt are submitted to the model.

[0050] The model can process the image, and in a case where the image includes the first object, the model can output a location where the object is located (e.g., a region coordinate of the object in the image, and / or directly output an image of the region where the object is located, etc.). With some implementations of the present disclosure, the location of the target object in the image can be detected based on a variety of manners so that the robot device can subsequently acquire the target object.

[0051] According to some implementations of the present disclosure, at least one candidate manner for processing the first object is determined. For example, in a real scenario, a user can have multiple eating methods for an apple, such as eating directly, eating after peeling, and juicing, etc. Specifically, in the process of determining the at least one candidate manner, the first prompt word can be obtained based on the user task and the user information of the user, and the at least one candidate manner is determined based on a first response of the machine learning model to the first prompt word.

[0052] More details about determining the at least one candidate manner are described with reference to FIG. 4, which shows a block diagram 400 of a process of invoking a language model according to some implementations of the present disclosure. As shown in FIG. 4, the first object 220 to be acquired can be determined as “apple” from the user task 210, at which time the first prompt word 420 can be obtained based on the user task 210 and the user information 410. The first prompt word 420 can be expressed as, for example, “please determine the eating method of ‘apple’ according to the user information”, or as, for example, “according to the user information, how can ‘apple’ be eaten”, etc.

[0053] According to some implementations of the present disclosure, the user information 410 can include a manner used by the user in the past to process the first object or a manner set in advance by the user to process the first object, such as that the user is more inclined to “eat directly” in the past. The first prompt word 420 can be input to the language model 430 to determine multiple eating methods of the apple. Alternatively and / or additionally, if it is found based on the usage history that the user often “eats after peeling”, the candidate manner can be determined as “eating after peeling”.

[0054] According to some implementations of the present disclosure, the language model 430 is a model that is trained and fine-tuned, and has rich knowledge of performing tasks in multiple domains. The language model 430 can determine the first object 220 as “apple” from the user task 210, and determine the eating method of “apple”. FIG. 4 is merely illustrative, and the language model 430 can process one or more user information at the same time, and determine one or at least one candidate way for processing the first object. For example, assuming that the user likes to eat apples both by peeling and by juicing, “eat by peeling” and “eat by juicing” can be determined as candidate ways. Alternatively and / or additionally, assuming that the user task is “I want to eat banana”, the model can determine “banana” as the first object, and can determine multiple eating methods of banana.

[0055] With some implementations of the present disclosure, at least one candidate way for processing the target object can be determined by using the powerful processing capability and rich knowledge of the model. In turn, the robot device can be instructed to obtain the target object based on the candidate way selected by the user. Compared with the prior art solution that can only process the target object by a preset way, the technical solution of the present disclosure can process the target object by the way currently desired by the user. In this way, the convenience and accuracy of task execution can be improved, and the intended user task can be completed in a more diverse form.

[0056] More details are described with reference to FIG. 5, which shows a block diagram 500 of a process of recognizing objects from an image according to some implementations of the present disclosure. As shown in FIG. 5, the object 510 (apple) and the object 520 (fruit knife) can be recognized from the image 310. According to some implementations of the present disclosure, an image of a physical space where the robot device is located can be obtained, and in response to the image indicating that the physical space includes a second object associated with a candidate way in the at least one candidate way, the candidate way is determined.

[0057] In addition, the first prompt word for determining the at least one candidate way can also be obtained based on the image. Specifically, in the process of determining the first object, multiple second objects associated with the processing of the first object can be recognized from the image, and the at least one candidate way can be determined based on the multiple second objects respectively. The powerful processing capability and rich knowledge of the model can be used to obtain the first prompt word for determining the at least one candidate way based on the image. Here, the knowledge base can be predefined and include the association relationship between the first object and the second object. For example, the knowledge base can include that the fruit knife can be used for peeling operation on the apple, the juicer can be used for juicing operation on the apple, and the like.

[0058] With some implementations of the present disclosure, various objects in a physical space can be determined using image recognition techniques and / or machine learning techniques, and at least one candidate way of processing a first object can be determined. In this way, the diversity of methods of processing the first object can be improved.

[0059] According to some implementations of the present disclosure, a second object corresponding to the candidate way can be determined based on an image of a physical space where the robotic device is located, and the robotic device can be instructed to acquire the first object and the second object in response to determining that the second object is a movable object. As shown in FIG. 5, in the image 310, the object 510 “apple” can be determined as the first object, and the object 520 “fruit knife” can be determined as the second object, both of which are movable objects. At this time, the robotic device can use the arm to grab the apple and the fruit knife to perform peeling. Alternatively and / or additionally, the robotic device can provide the apple and the fruit knife to the user, so that the user determines whether to peel based on his / her own needs.

[0060] On the contrary, in response to determining that the second object is an immovable object, the robotic device can be instructed to use the second object to process the first object, and the robotic device can be instructed to acquire the processed first object. For example, in a case where the object 530 “juicer” in the image 310 is determined as the second object, the robotic device can move to the position of the second object, and directly use the second object to process the first object (e.g., put the apple into the juicer to juice), to obtain the processed first object (e.g., apple juice). With some implementations of the present disclosure, different processing ways can be determined by recognizing movable or immovable second objects, and the flexibility of task execution is improved.

[0061] According to some implementations of the present disclosure, in the process of using the second object to process the first object, an operation way of operating the second object can be determined based on the image and the user task, and the robotic device can be instructed to operate the second object according to the operation way in order to process the first object. Specifically, the handle of the fruit knife can be recognized from the image, and at this time it can be determined that the handle needs to be held and the blade needs to be moved to peel the apple.

[0062] Alternatively and / or additionally, a second prompt can be constructed based on the image and the user task, and the model can be asked to perform the user task in a manner of operation by a second object in the image. The second prompt can be expressed as, for example, “determine the way to use the fruit knife from the following image”, and the second prompt and the corresponding image can be sent to the model. At this time, the model can return: hold the handle and move the blade. In turn, the robot device can be instructed to hold the handle and move the blade to peel the apple. Alternatively and / or additionally, a series of instructions for performing the peeling action can be transmitted to the robot device in order to control the robot device. With some implementations of the present disclosure, the powerful processing capability of the model can be invoked to solve unknown problems in complex environments, and in turn determine the action that needs to be performed by the robot device. In this way, the ability of the robot device to handle complex tasks can be improved, and the user task can be performed in a more accurate manner.

[0063] According to some implementations of the present disclosure, an action model can be utilized to determine the specific action to be performed by the robot device. More details are described with reference to FIG. 6, which shows a block diagram 600 of a process of invoking an action model according to some implementations of the present disclosure. As shown in FIG. 6, an action model 630 can be provided, which can determine the specific action to be performed by the robot device based on the current state of the robot device and the instruction, and which can be a pre-trained and fine-tuned model.

[0064] It should be understood that the current state can include a plurality of aspects of data, such as an image of the robot device, an image of the environment of the robot device, pose data of the robot arm (e.g., positions of various joints of the robot arm (POS1, …)), and a state of a tool (e.g., a gripper, a knife, etc.) fixed at the end of the robot arm. For example, 0 can be used to represent a closed state of the gripper, and 1 can be used to represent an open state of the gripper. The instruction and the current state can be input to the action model 630, and in turn the action model can be utilized to determine the action to be performed by the robot device based on the instruction and the current state. Here, the action can represent a difference between the current pose of the robot device and the next pose, and a difference between the current state of the tool and the next state, etc.

[0065] The instruction 610 (e.g., "grab the apple") can be input to the action model 630, where the instruction 610 can be expressed in natural language and can be determined from a response of the language model. Further, a current state of the robotic device can be acquired, and the action model 630 can determine a corresponding action 640 based on the input data. For example, the orientations, positions, velocities, accelerations, etc. of various joints in the arm, and / or wheels and / or other movable devices of the robotic device at a next point in time can be determined. Further, the determined action 640 can be utilized to control the state of the robotic device at the next point in time.

[0066] Although the above describes a specific procedure of determining an action with "grab the apple" as an example, alternatively and / or additionally, the instruction 610 can be used to perform other complex operations, e.g., "peel the apple", "operate the juicer", etc. At this time, the current image and the current pose of the robotic device can be continuously acquired in order to drive the robotic device to reach a next pose.

[0067] With some implementations of the present disclosure, a correlation can be established between the language model and the action model, and a user task initially input by a user in natural language can be converted into a specific action executable by a robotic device. In this way, the action of the robotic device can be precisely controlled, and the user task can be performed with higher efficiency.

[0068] It should be appreciated that although the above describes one example implementation according to the present disclosure with a Chinese language environment as an example. Alternatively and / or additionally, the technical solution of one example implementation according to the present disclosure can be performed in a plurality of language environments. For example, a robot can be controlled in a Chinese, English, Japanese, French, etc. environment. Specifically, the robot can be controlled in application environments of different languages based on the multi-language capability provided by the machine learning technology. Further, although the above describes a procedure of performing a user task with a robotic device with eating an apple as an example, alternatively and / or additionally, the robotic device can be controlled to perform other user tasks, e.g., to acquire other items in a room, to process a certain item in other various ways, etc. Assuming that a user instructs the robotic device to retrieve a slice of bread, at least one candidate way can include eating it directly, eating it after being toasted with a toaster, etc.

[0069] According to some implementations of the present disclosure, the at least one candidate manner is determined based on the first response to the first prompt word of the machine learning model. More details are described with reference to FIG. 7, which shows a schematic diagram 700 of an example interface according to some implementations of the present disclosure. As shown in FIG. 7, the selection of the candidate manner can be implemented through the interaction unit 112 on the robotic device 110. The interaction unit 112 can be a display screen with interaction function, and the content displayed on the display screen can include a plurality of controls, such as a first control 710, a second control 720, and a third control 730. Each of the plurality of controls can correspond to a candidate manner, for example, the first control 710 corresponds to “eat directly”, the second control 720 corresponds to “eat after peeling”, and the third control 730 corresponds to “eat after juicing”. The user can select the processing manner for the first object by triggering the control. With some implementations of the present disclosure, the interaction content between the user and the robotic device can be greatly enriched, and the diversity of task execution is expanded.

[0070] Alternatively and / or additionally, the user can interact with the robotic device via the interaction unit 112, for example, the user inputs a task represented in words and / or images, and controls the robotic device to execute the task. Alternatively and / or additionally, the user can specify the execution condition of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when it is determined that a predetermined condition is met (for example, when the user makes an action of eating an apple), and the like.

[0071] According to some implementations of the present disclosure, the robotic device can provide a variety of messages to the user, for example, assuming that the robotic device does not find other objects that can be used to process the apple, the robotic device can ask the user whether it is necessary to eat directly. For another example, assuming that the robotic device does not find the apple and only finds the tool for processing the apple, the robotic device can inform the user that the apple is not found, and the like. Alternatively and / or additionally, the robotic device can ask the user where the desired object can be found, and go to the location specified by the user to find the desired object. Alternatively and / or additionally, if the desired object cannot be found, the robotic device can ask the user whether it is necessary to purchase, and the like.

[0072] According to some implementations of the present disclosure, a variety of positioning algorithms can be utilized to determine the location of the robotic device, as well as individual objects, in the physical environment. For example, a global positioning system (GPS) can be deployed at the robotic device, and satellite signals can be used to determine the precise location of the robotic device. Alternatively and / or additionally, a communication unit can be deployed at the robotic device, by means of which signals between the communication unit and a base station, and utilizing a communication network, the location of the robotic device can be determined. Alternatively and / or additionally, Wi-Fi access points can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspots in order to determine the location via Wi-Fi signal strength and the location of known Wi-Fi access points. Alternatively and / or additionally, the communication unit at the robotic device can support Bluetooth functionality, at which time Bluetooth signals and known Bluetooth device locations can be used to determine the location of nearby devices. An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, and in turn determine the location of the robotic device.

[0073] Alternatively and / or additionally, a visual positioning system can be used to determine the location of the robotic device and / or individual objects. A map of the physical space can be pre-acquired, and the location of individual objects can be annotated in the map. The robotic device can utilize a backscatter detection unit to detect the distance from surrounding objects, and in conjunction with the captured images and the map of the physical space, determine the specific location of individual objects. Specifically, computer-aided design (CAD) and geographic information systems (GIS) can be used, and positioning algorithms can be utilized to determine the location. Alternatively and / or additionally, tracking units can be deployed at important objects in the physical space, for example, a tracking unit can be added at the remote control of a household appliance (e.g., a television remote control, an air conditioner remote control), so that the robotic device can timely acquire the precise location of important objects, and the like.

[0074] According to some implementations of the present disclosure, the original location of the robotic device itself and the destination location to which it is expected to go can be determined based on the methods described above. The robotic device can determine a path from the original location to the destination location. For example, the surrounding environment images can be constantly acquired, and the path can be constantly updated while ensuring to avoid obstacles, and the robotic device is caused to move along the path to the destination location.

[0075] According to some implementations of the present disclosure, after reaching the destination location, the robotic device can perform the specified task. For example, a specified object can be fetched and moved to a corresponding location. The constraints, i.e., constraints that should be followed during the performance of the task, can be determined with the language model and / or the knowledge base. For example, an image and a corresponding prompt word can be obtained, and the image and the prompt word can be input to the language model, and in turn, the constraints can be received from the language model. For example, the prompt word can be determined as: “please determine the constraints that should be followed during the movement of the XXX object based on the following image,” or “please determine the precautions during the movement of the XXX object,” and the like.

[0076] At this time, it can be determined that the original posture of the object (e.g., a bottle of water, a plate, a bowl, etc.) should be maintained (e.g., the vertical direction should be maintained and the object should not be tilted) during the movement of the object. Further, the constraints can be input to the action model, and in turn, a series of actions output by the action model will perform the corresponding task while ensuring the constraints. With some implementations of the present disclosure, the safety during the operation of the robotic device can be ensured, and in turn, the accidental damage to an object can be avoided, and the like.

[0077] With exemplary implementations of the present disclosure, a user can select different candidate manners for performing a task, and the robotic device can complete the intended user task in a more diverse form. In this way, the flexibility and accuracy of the robotic device in performing a task under different requirements can be improved, and in turn, the intended user task can be completed.

[0078] Example process

[0079] FIG. 8 illustrates a flowchart of a method 800 for performing a user task according to some implementations of the present disclosure. At block 810, a user task is received from a user, the user task instructing a robotic device to fetch a first object. At block 820, at least one candidate manner for handling the first object is determined. At block 830, a first message indicating the at least one candidate manner is provided to the user. At block 840, based on a selection of a candidate manner from the at least one candidate manner by the user, the robotic device fetches the first object.

[0080] According to some implementations of the present disclosure, determining the at least one candidate manner comprises: based on the user task and user information of the user, obtaining a first prompt word, the first prompt word being used to determine the at least one candidate manner; and based on a first response of a machine learning model to the first prompt word, determining the at least one candidate manner.

[0081] According to some implementations of the present disclosure, the method 800 further includes obtaining an image of a physical space in which the robotic device is located, and determining the candidate approach in response to the image indicating that the physical space includes a second object associated with the candidate approach of the at least one candidate approach.

[0082] According to some implementations of the present disclosure, obtaining the first prompt further includes obtaining an image of a physical space in which the robotic device is located, and obtaining the first prompt based on the image.

[0083] According to some implementations of the present disclosure, obtaining the first object includes determining a second object corresponding to the candidate approach based on an image of a physical space in which the robotic device is located, and in response to determining that the second object is a movable object, the robotic device obtaining the first object and the second object.

[0084] According to some implementations of the present disclosure, obtaining the first object includes, in response to determining that the second object is an immovable object, the robotic device processing the first object with the second object, and the robotic device obtaining the processed first object.

[0085] According to some implementations of the present disclosure, processing the first object with the second object includes determining an operation approach for operating the second object based on the image and the user task, and the robotic device operating the second object in accordance with the operation approach to process the first object.

[0086] According to some implementations of the present disclosure, determining the operation approach for operating the second object includes obtaining a second prompt based on the image and the user task, the second prompt being used to determine the operation approach for performing the user task by the second object in the image, and determining the at least one candidate approach based on a first response to the first prompt by the machine learning model.

[0087] Example apparatuses and devices

[0088] FIG. 9 illustrates a block diagram of an apparatus 900 for performing a user task, according to some implementations of the present disclosure. The apparatus 900 includes a receiving module 910 configured to receive a user task from a user, the user task indicating a robotic device to obtain a first object, a determining module 920 configured to determine at least one candidate approach for processing the first object, a message providing module 930 configured to provide a first message to the user indicating the at least one candidate approach, and an obtaining module 940 configured to cause the robotic device to obtain the first object based on a selection of a candidate approach of the at least one candidate approach by the user.

[0089] According to some implementations of the present disclosure, the determining module 920 is further configured to: obtain, based on the user task and the user information of the user, a first prompt word, the first prompt word being used to determine the at least one candidate manner; and determine the at least one candidate manner based on a first response of the machine learning model to the first prompt word.

[0090] According to some implementations of the present disclosure, the obtaining module 940 is further configured to: obtain an image of a physical space in which the robotic device is located; and the determining module 920 is further configured to: determine the candidate manner in response to the image indicating that the physical space includes a second object associated with the candidate manner in the at least one candidate manner.

[0091] According to some implementations of the present disclosure, the obtaining module 940 is further configured to: obtain an image of a physical space in which the robotic device is located; and the determining module 920 is further configured to: obtain the first prompt word based on the image.

[0092] According to some implementations of the present disclosure, the obtaining module 940 is further configured to: determine, based on an image of a physical space in which the robotic device is located, a second object corresponding to the candidate manner; and cause the robotic device to obtain the first object and the second object in response to determining that the second object is a movable object.

[0093] According to some implementations of the present disclosure, the obtaining module 940 is further configured to: cause the robotic device to process the first object with the second object in response to determining that the second object is an immovable object; and cause the robotic device to obtain the processed first object.

[0094] According to some implementations of the present disclosure, the obtaining module 940 is further configured to: determine, based on the image and the user task, an operation manner for operating the second object; and cause the robotic device to operate the second object according to the operation manner so as to process the first object.

[0095] According to some implementations of the present disclosure, the determining module 920 is further configured to: obtain, based on the image and the user task, a second prompt word, the second prompt word being used to determine an operation manner for performing the user task by the second object in the image; and determine the at least one candidate manner based on a first response of the machine learning model to the first prompt word.

[0096] FIG. 10 illustrates a block diagram of a device 1000 that can implement a number of implementations of the present disclosure. It should be understood that the computing device 1000 illustrated in FIG. 10 is merely an example and should not be construed as any limitation of the functionality and scope of the implementations described herein. The computing device 1000 illustrated in FIG. 10 can be used to implement the methods described above.

[0097] As shown in FIG. 10, computing device 1000 is in the form of a general-purpose computing device. Components of computing device 1000 can include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processing unit(s) 1010 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 1020. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of computing device 1000.

[0098] Computing device 1000 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Memory 1020 can be volatile (such as, for example, registers, cache, RAM), non-volatile (such as, for example, ROM, EEPROM, flash memory), or some combination of the two. Storage 1030 can be removable or non-removable and can include, but is not limited to, magnetic disks, optical disks, or tape. Storage 1030 can also comprise solid-state drives (SSDs) or any other suitable type of storage devices. Storage 1030 can be used to store data (e.g., training data for training) and / or instructions that enable the functioning of or at least part of the functionality of computing device 1000.

[0099] Computing device 1000 can further include additional removable / non-removable, volatile / nonvolatile storage media. Although not shown in FIG. 10, a disk drive or other memory expansion card interface for reading from and / or writing to a removable, non-removable, volatile, and / or non-volatile computer storage medium can be provided. In these instances, each drive can be connected to the bus by one or more data media interfaces. Memory 1020 can include a computer program product 1025 that has one or more program modules configured to carry out the various methods or actions of the implementations of the present disclosure.

[0100] Communication unit(s) 1040 enable communication over communication media to other computing devices. Additionally, the functionality of components of computing device 1000 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in a distributed environment.

[0101] Input device 1050 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 1060 can be one or more output devices, such as a display, a speaker, a printer, etc. Computing device 1000 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through communication unit 1040, and with one or more devices that enable a user to interact with computing device 1000, and / or any devices (e.g., a network card, a modem, etc.) that enable computing device 1000 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface (not shown).

[0102] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.

[0103] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0104] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium. The instructions stored on the computer readable storage medium can be used to program a computer, a programmable data processing apparatus, and / or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0105] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0106] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0107] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although being shown and described in terms of certain implementations and overall functions, the implementations are not intended to exclude other implementations or technologies. Modifications and changes can be made in arrangement, operation, and details of the methods and apparatus described. Many modifications and variations of the described implementations are possible and will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. None, individually or taken in any combination, is to be treated as limiting the described implementations unless otherwise indicated.

Claims

1. A method for performing a user task, comprising: receiving a user task from a user, the user task instructing a robotic device to acquire a first object; determining at least one candidate way for processing the first object; providing a first message to the user indicating the at least one candidate way; and based on a selection of a candidate way of the at least one candidate way by the user, the robotic device acquires the first object.

2. The method of claim 1, wherein determining the at least one candidate way comprises: based on the user task and user information of the user, acquiring a first cue, the first cue being used to determine the at least one candidate way; and based on a first response of a machine learning model to the first cue, determining the at least one candidate way.

3. The method of claim 2, further comprising: acquiring an image of a physical space where the robotic device is located; and in response to the image indicating that the physical space includes a second object associated with a candidate way of the at least one candidate way, determining the candidate way.

4. The method of claim 2, wherein acquiring the first cue further comprises: acquiring an image of a physical space where the robotic device is located; and based on the image, acquiring the first cue.

5. The method of claim 1, wherein acquiring the first object comprises: based on an image of a physical space where the robotic device is located, determining a second object corresponding to the candidate way; and in response to determining that the second object is a movable object, the robotic device acquires the first object and the second object.

6. The method of claim 5, wherein acquiring the first object comprises: in response to determining that the second object is an immovable object, the robotic device processes the first object with the second object; and the robotic device acquires the first object processed.

7. The method of claim 6, wherein processing the first object with the second object comprises: based on the image and the user task, determining an operation way for operating the second object; and the robotic device operates the second object in accordance with the operation way to process the first object.

8. The method of claim 7, wherein determining the operation way for operating the second object comprises: based on the image and the user task, acquiring a second cue, the second cue being used to determine an operation way for performing the user task by the second object in the image; and based on a first response of a machine learning model to the first cue, determining the at least one candidate way.

9. An apparatus for performing a user task, comprising: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; a determining module configured to determine at least one candidate way for processing the first object; ​ ​ ​ ​ ​ a message providing module configured to provide a first message indicative of the at least one candidate modality to the user; and an obtaining module configured to obtain, by the robotic device, the first object based on a selection of a candidate modality among the at least one candidate modality by the user. 10.An electronic device comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 8. 11.A computer readable storage medium having stored thereon a computer program, the computer program, when executed by a processor, causing the processor to implement the method according to any one of claims 1 to 8. 12.A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Robot control method and device, electronic equipment and storage medium

    CN111975772A

  • Robot control method and robot

    CN114488879A

  • Robot task planning method and device and terminal equipment

    CN116810770A

  • Providing diet assistance in a session

    US20200242964A1